The Web Is Your Oyster - Knowledge-Intensive NLP against a Very Large Web Corpus

Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Dmytro Okhonko, Samuel Broscheit, Gautier Izacard, Patrick Lewis, Barlas Oğuz, Edouard Grave, Wen-tau Yih, Sebastian Riedel

Introduction

The ability to access and manipulate knowledge has become one of the core features of modern NLP systems. Knowledge-intensive NLP (KI-NLP) tasks such as fact checking, open-domain question answering (ODQA) and entity linking typically specify a source of factual knowledge necessary to provide an explainable solution. Often, this source is Wikipedia Dinan et al. (2019); Christodoulopoulos et al. (2020); Kwiatkowski et al. (2019), for obvious reasons: it tends to be highly accurate, it is well-structured and small enough to test computationally demanding architectures. Still, there exist many reasons to look beyond Wikipedia. First, it covers a lot of ground, but certainly not everything, and in practice many information needs cannot be fulfilled based on Wikipedia alone Redi et al. (2021). Second, even for topics it does cover, there might be biases that cannot be resolved without looking at a broader context Wagner et al. (2016); Graells-Garrido et al. (2015). Unsurprisingly, Wikipedia has also never gained much traction in common sense NLP. Contrary to the factual knowledge, common sense is believed to be universally accepted by humans while stated more implicitly Xie and Pu (2021). There have been many attempts to capture such implicit knowledge vie common sense knowledge bases Speer and Havasi (2012); Sap et al. (2019a); Bhakthavatsalam et al. (2020). While yielding great results, their supervised nature makes them hard to generalize and expand.

By virtue of its sheer scale, the web promises access to knowledge both broader and more in-depth than Wikipedia, providing not only sheer facts but also context useful in inferring rules of common sense reasoning. Along with the benefits, however, come new challenges—lack of structure, inconsistent document quality and noisy or harmful content on one hand Luccioni and Viviano (2021), increasing infrastructural demands on the other. Today, the impact of these challenges on knowledge tasks is not clear—while work investigating the use of web in KI-NLP exists, it usually relies on commercial, black-box search engines, focuses on individual tasks, primarily ODQA Joshi et al. (2017); Bajaj et al. (2018); Talmor and Berant (2019); Nakano et al. (2021), or only uses general web content at pre-training time Guu et al. (2020); Borgeaud et al. (2022); Lewis et al. (2020a).

We propose to use a web corpus as a universal, uncurated and unstructured knowledge source for multiple KI-NLP tasks at once. We aim to answer the following question: what impact does replacing Wikipedia with a large-scale web corpus have on the performance of knowledge-intensive systems? Specifically, should we expect them to improve, since for a given fact, there is more potential evidence on the web, or degrade, due to uncurated nature of the data? Multiple factors such as the scope of knowledge covered by the corpus, the ability of retrievers to generalize across downstream tasks and the scalability of the solution may contribute to the answer. We propose a unified retrieval infrastructure and analyze these aspects of our setup in depth. We then explore if our web index can serve as a knowledge source in common sense NLP.

We leverage an open web corpus coupled with strong retrieval baselines instead of a black-box, commercial search engine—an approach which facilitates transparent and reproducible research and opens up a path for future studies comparing search engines optimised for humans with retrieval solutions designed for neural networks. We use a subset of CCNet Wenzek et al. (2020) covering 134M documents split into 906M passages as the web corpus which we call Sphere. While far from the full web scale, Sphere is orders of magnitude larger than previously studied knowledge sources (cf. Table 2). We consider two retrieval architectures—BM25 Robertson (2009) and DPR Karpukhin et al. (2020), and combine them with a FiD reader component Izacard and Grave (2020). To facilitate large-scale dense retrieval, we open-source distributed-faiss—a wrapper around the FAISS similarity search library Johnson et al. (2017), simplifying the distribution of indices across machines. We use KILT Petroni et al. (2021), a standard KI-NLP benchmark, as well as a range of common sense tasks listed in Table 3 to evaluate our work.

Despite inconsistent quality of the web and the fact that KILT was specifically designed to query knowledge from Wikipedia, we find that Sphere-based models can match or outperform baselines grounded in Wikipedia on a subset of KILT tasks. In some cases this holds even when we aggressively filter Sphere by removing not just Wikipedia itself but also content that looks like it. Moreover, we find that retrieval from Sphere can improve the common sense abilities of FiD, despite the lack of explicit common sense knowledge in retrieved passages. To our knowledge, this is the first time a general purpose search index improves language models on common sense tasks.

We also find ample room for future work: while dense retrieval outperforms sparse methods in most prior work, in our case the opposite is true. How to develop large scale and universal dense indices supporting a multitude of tasks hence remains an open question. To summarize, we make the following contributions.

We replace Wikipedia with Sphere as the knowledge source for a selection of KILT tasks, achieving state of the art results on two.

We carry out an in-depth analysis of our web corpus and retrievers, identifying potential reasons for both gains and losses in end-to-end performance on respective tasks.

We show that general purpose web retrieval from Sphere can improve common sense reasoning on 8 tasks when compared to a comparable, fully parametric language model.

We release sparse and dense indices of Sphere and open-source distributed-faiss.

Background

We typically call an NLP task knowledge-intensive if a human would not be reasonably expected to solve it without access to an external knowledge source. KI-NLP tasks are usually solved with retriever-reader systems: first, a retriever surfaces a small set of relevant documents from the knowledge source, then a reader uses the context to generate an answer Chen et al. (2017); Lewis et al. (2020b); Guu et al. (2020). For the purpose of this work we expand our definition of knowledge-intensive to any NLP task which, beyond core language capabilities, also requires knowledge—be it factual or common sense—to solve the problem at hand.

2 Retrieval models

We consider two retrieval architectures. BM25 Robertson (2009) is a popular sparse model, where queries and documents are represented as high-dimensional, sparse vectors, with dimensions corresponding to vocabulary terms and weights indicating their importance. DPR Karpukhin et al. (2020) is a dense model which embeds queries and documents into a latent, real-valued vector space of a much lower dimensionality—an idea originating from the Latent Semantic Analysis Deerwester et al. (1990). DPR is based on a neural bi-encoder architecture with passages and queries embedded with separate text encoders. Although both sparse and dense models use the distance in the vector space as the relevance function, they need different indexing schemes to support efficient retrieval.

3 Reader models

Reader models typically consume a set of context documents retrieved from a knowledge source and the task input, and return the output—either a class label or text. In this work, we use an abstractive Fusion-in-Decoder (FiD) reader from Izacard and Grave (2020) — an encoder-decoder architecture, where each context document is concatenated with the input and embedded by the encoder. In the decoder, attention is performed over encoded passages and then the output is generated.

Search Infrastructure

A question equally important to the choice of the knowledge corpus itself pertains to the feasibility of implementing a research-friendly search infrastructure on top of it. An index is a data structure which stores representations of corpus documents, built with the objective of optimizing the retrieval efficiency. For sparse methods, this goal is typically achieved with an inverted index—a space-efficient technique entertaining the support of multiple robust libraries such as Pyserini Lin et al. (2021b). Efficient dense retrieval is enabled by maximum inner product search algorithms Shrivastava and Li (2014); Guo et al. (2015) leveraged by tools like FAISS Johnson et al. (2017), a robust library for similarity search and clustering of dense vectors. As the size of the text corpus grows, a FAISS index may exceed typical, single-server hardware limits for both GPU and RAM. Two main approaches for handling scale emerge: compression of the document embeddings and distribution of the index over multiple servers. Good compression rates can be achieved with quantizers available in FAISS out-of-the-box or with more sophisticated bi-encoder training pipelines Yamada et al. (2021); Zhan et al. (2021)—this may help reduce the index size by a few times factor but does not solve the core scaling issue. The ability to distribute a FAISS index is what we address with our open-source release of distributed-faiss https://github.com/facebookresearch/distributed-faiss At indexing time, the distributed-faiss client receives batches of embeddings to be indexed and routes them to the index servers guaranteeing a balanced data distribution. At retrieval time, the client queries all servers and aggregates the results. The service is model-independent and operates with supplied embeddings and metadata. We also release indices of Sphere https://github.com/facebookresearch/Sphere both for the sparse retrieval baseline, compatible with Pyserini, and our best dense model compatible with distributed-faiss.

Experimental setup

We experiment with two universal (non task-specific) knowledge sources. First, the KILT knowledge source based on the 2019/08/01 Wikipedia snapshot, comprising 5.9M articles split into 22.2M passages of 100 tokens. We refer to this corpus as Wikipedia in the reminder of this paper. Second, we use CCNet Wenzek et al. (2020) to create our web corpus. CCNet processes Common Crawl by performing deduplication, language identification and quality filtering (articles are split into three quality tiers: head, middle and tail based on perplexity under a Wikipedia-based language model). We use the head tier of a single CCNet snapshot in English, additionally excluding all articles containing wikipedia.org in their URL. We pick the CCNet snapshot corresponding to the August 2019 Common Crawl snapshothttps://commoncrawl.org/2019/08/august-2019-crawl-archive-now-available/ as it is temporally the closest to the KILT knowledge source. It consists of 134M web articles and yields 906.3M passages of 100 tokens. We call this corpus Sphere in the remainder of this paper.

Downstream Tasks.

We use the KILT benchmark as a KI-NLP evaluation suite for our work. KILT contains 11 tasks, split into 5 categories: fact checking, entity linking, slot filling, ODQA and dialog. Since entity linking is intrinsically tied to the underlying Wikipedia corpus, as entity labels are equivalent to the titles of Wikipedia pages representing respective entities, we choose to exclude it from our work. We also experiment with a collection of common sense tasks—see Table 3 for the full list of datasets and shortcuts we use to refer to them. For each downstream tasks, we train a FiD reader model, using T5-base Raffel et al. (2019) as initialization for the KILT tasks, and T5-large for the common sense tasks. Otherwise, we follow the finetuning setup proposed by FiD authors. Unless otherwise stated, we train FiD with the top 100 passages retrieved from considered corpora.

Baseline retrievers.

We experiment with three retrieval baselines: BM25, DPRmulti, a variant of DPR pre-trained in a multi-task fashion on KILT by Maillard et al. (2021), and DPRweb, a new DPR model trained for the purpose of this work (details in the next paragraph). Both DPR models use 768-dim encoders. We use Pyserini Lin et al. (2021b) to build the BM25 index of Sphere and Wikipedia and distributed-faiss to build the dense indices. We build an HNSWSQ8 FAISS index of Sphere, with 16 physical and 32 logical nodes providing a good accuracy and latency of 100–200 ms. The index, which consists of the embeddings and metadata occupies 1.9TB of disk space. We use a flat FAISS index for Wikipedia. We follow Karpukhin et al. (2020) and index 100-token long passages along with article titles. We apply all three retrievers to Sphere. As DPR has been repeatedly shown to outperform BM25 in retrieval from Wikipedia on KI-NLP tasks Karpukhin et al. (2020); Maillard et al. (2021), we skip BM25 in those experiments. Due to the lack of retrieval supervision, we only consider BM25 in our common sense experiments.

Training DPRweb.

Our goal is to leverage Sphere in the training of a DPR web retriever. Given the lack of any explicit retrieval supervision over Sphere, we use two proxy metrics to track performance: answer-in-context@k (AIC@k), indicating the fraction of examples for which there exists a passage containing the gold answer among the top-kk retrieved ones; and answer+entity-in-context@k (AEIC@k), indicating the fraction of examples for which among the top-kk passages there exists one containing both the gold answer and the main entity of the datapoint—we use the Wikipedia title of the gold retrieval passage defined by KILT as the main entity. We train DPRweb by finetuning a PAQ-based Lewis et al. (2021) bi-encoder checkpoint Oğuz et al. (2021) for 40 epochs on 16 GPUs. We source finetuning data from KILT tasks compatible with the AIC metric—so those with short-form textual answers (T-REx, zsRE, NQ, HoPo and TQA) and apply the model in zero-shot fashion to the remaining tasks. We balance the number of datapoints per dataset by sampling with the same rates as in Maillard et al. (2021). For each datapoint we source context passages in 4 ways (with in-batch negatives in all cases): (1) gold Wikipedia passages as positives and BM25-based negatives (the same as in original DPR); (2) gold Wikipedia passages as positives and hard Wikipedia negatives (BM25 index based, the same as were used to finetune DPRmulti); (3) weakly supervised web positives from the BM25 Sphere index; (4) weakly supervised web positives from the DPRmulti Sphere index. We obtain the weakly supervised positive web passages by picking the top result returned by respective baseline retrievers containing the gold answer of a given datapoint. We use the batch size 32 and default DPR hyperparameters otherwise.

Results

We present our main results in Table 4. It is important to remember that the KILT benchmark was designed with a specific Wikipedia snapshot in mind and examples for which no evidence was found were removed. Thus, there is a strong bias towards Wikipedia as the knowledge source, and the performance of systems using it can be considered topline. We also note that ours is the first paper to report FiD results on KILT—our baseline FiD+DPRmulti model outperforms similar DPR-based architectures across the board. In order to factor out the impact of moving to a stronger reader, we mainly focus on comparing our Sphere-based models to our Wikipedia baselines, with FiD reader in both cases. Our Sphere-based FiD+BM25 architecture establishes a new state of the art (SOTA)We compare to current (Nov. 2021) leaderboard results at kiltbenchmark.com published on arxiv.org. on FEV and TQA (see Table 13 in Appending for examples). We also note that a FiD+DPRweb model beats SOTA on zsRE, NQ and HoPo with Wikipedia as the knowledge source.

With Sphere as the knowledge source, DPRweb outperforms DPRmulti on all KILT tasks but ELI5 (see Table 12 in the Appendix for retrieval evaluation), yielding notable gains downstream: +8 points on zsRE, +6 on TQA, +5 on T-REx. Interestingly, when used to retrieve from Wikipedia, DPRweb also helps. We see gains on all tasks used in DPR finetuning except T-REx, with new SOTAs on zsRE, NQ and HoPo. By contrast, results on tasks excluded from DPR finetuning are not consistent. Unlike with FEV, where DPRweb yields SOTA both against Wikipedia and Sphere, we don’t observe downstream gains for the long-form QA and dialog, which highlights the challenge of zero-shot transfer in dense retrieval.

2 Universal web retrieval

Given that we use Sphere without any explicit alignment to the datasets we consider, the question of whether the corpus actually contains information necessary to solve the task at hand becomes of major importance. First, we note that the popularity of the topic impacts how well represented it is on the web. This inadvertently leads to a limited knowledge coverage of rare topics when working with incomplete snapshots rather than an exhaustive index of the web. As a consequence, both slot-filling tasks suffer a large drop in performance when switching from Wikipedia to Sphere (see Section A.1 of the Appendix for more details). Subsequently, we observe that on other tasks, Sphere is competitive with Wikipedia. The SOTA that the FiD+BM25 architecture achieves on TQA is our most salient result, outperforming our best model grounded in Wikipedia by over 6 points. TQA can be considered one of the least Wikipedia-dependent of all KILT tasks—an encouraging evidence that web knowledge may be particularly useful in satisfying diverse information needs, especially those going beyond Wikipedia. Finally, in Table 6, we report results on KILT dev sets for the best systems using Wikipedia and Sphere respectively, and a hypothetical, hybrid, oracle system which is correct if either of them is correct. The oracle outperforms both baselines, suggesting that evidence provided by Sphere adds value on top of Wikipedia.

Wikipedia on the web.

Based on a simple heuristic (details in Section A.2 of the Appendix), we estimate that over 5% of Sphere passages are likely a copy from Wikipedia, with 47% of Wikipedia passages having an equivalent in our web corpus. Following this observation we note that all considered retrieval methods have a bias towards Wikipedia, surfacing a disproportionally high number of Wikipedia-based passages, with BM25 being the least biased. In the Appendix, we analyze the impact of Wikipedia passages on the Sphere downstream results further.

Sparse vs. dense models.

The AIC gains we observe when moving from DPRmulti to DPRweb on Sphere correlate well with downstream performance. However, we don’t see a similarly strong dependency between DPRweb and BM25 (Figures 2(a) and 2(c))—though the former often achieves better AIC scores, it lags downstream for all datasets but NQ. To explain this, we ablate on the number of retrieved passages (see Figure 3 in the Appendix). The fewer contexts we consider, the smaller the BM25 advantage—if we use only the top one, the DPRweb-based model is better across the board, correlating well with AIC@1. This suggests that while DPRweb is able to find a good top document, the quality of the larger result set is worse—possibly because of false positives introduced when using AIC as retrieval supervision. We investigate result set quality further by looking at the AEIC metric. Here, BM25 achieves the best results on all datasets except NQ (Figure 2(d)), correlating better with downstream performance. We check how often the main entity is present in the input itself—NQ is an outlier in this regard, with the lowest fraction of datapoints containing the main entity (Table 5). It has been shown previously that BM25 is better at lexical exact-match on the salient spans in the query Chen et al. (2021). In our experiments BM25 can leverage this advantage—however, if the queries are more challenging in this regard as it is in the case with NQ, DPR becomes competitive.

Conclusions.

Even though neural retrievers such as DPR beat BM25 by a large margin on Wikipedia, we haven’t been able to apply them to Sphere with a similar success. It was suggested before that bi-encoders may be inherently not expressive enough for the purpose of large scale retrieval Luan et al. (2021). Still, we do see avenues for improvement. AIC may be too weak of a signal for retrieval supervision, with AEIC emerging as a potential alternative. In addition, our DPR models display a bias towards Wikipedia-based results which could be mitigated by picking better positive samples for finetuning from the web. In line with previous research Maillard et al. (2021); Oğuz et al. (2021) we also note that zero shot transfer of DPR models doesn’t yield good results, leaving the challenge of building universal, neural web retrievers open.

3 Common Sense Tasks

Rather than competing with the state of the art, which, for many common sense tasks, can be achieved with billion-parameter-scale, closed-book language models Lourie et al. (2021), we propose a proof-of-concept experimental setup. Our goal is to validate a hypothesis that knowledge augmentation with a general-purpose web index can positively impact the performance of an end-to-end system on common sense tasks. Table 7 contains downstream results for Wikipedia and Sphere-augmented FiD models, as well as a comparable in size, closed-book, T5-large baseline. We observe that retrieval brings consistent gains, with Sphere providing a clearly better improvement than Wikipedia on COPA, PIQA, HellaSWAG and NumerSense, indicating that it can serve as a broader source of common sense knowledge. When investigating the passages retrieved from Sphere, we find that they rarely surface explicit rules or generic statements expressing common sense knowledge. Rather, the retriever finds instances of real world situations that serve to build an on-the-fly, common sense understanding of the problem at hand (see examples in Table 14 in the Appendix). This poses an interesting challenge to the reader which needs to infer general rules based on specific illustrations of their application - like in the CommonsenseQA example, where the model should predict that people like to have coffee in the office based on a description of a dream office with a coffee table in it.

Related Work

Most existing research into factual KI-NLP uses Wikipedia as the source of knowledge Kwiatkowski et al. (2019); Joshi et al. (2017); Thorne et al. (2018); Yang et al. (2018); Dinan et al. (2019); Petroni et al. (2021). In this paper, we instead study our ability to solve KI-NLP tasks with web as the background corpus. Previous works that operate on web Joshi et al. (2017); Bajaj et al. (2018); Talmor and Berant (2018) typically rely on results from black-box search engines to create a corpus. A CCNet snapshot has been considered as a knowledge source in dialog research by Komeili et al. (2021), where authors use it together with a Wikipedia snapshot. As far as we know, our work is the first to consider an uncurated snapshot of the web without Wikipedia as a knowledge source for multiple KI-NLP tasks at once. Moreover, our scale is significantly larger than previously attempted (see Table 2). There are other large scale resources that could be considered to tackle KI tasks, such as large collections of question-answer pairs Lewis et al. (2021); Huber et al. (2021), structured knowledge sources Berant et al. (2013); Levy et al. (2017); Elsahar et al. (2018) or domain specific collections Tsatsaronis et al. (2015); Saikh et al. (2021). As for the common sense NLP tasks, while large pretrained models have achieved remarkable performance Raffel et al. (2019); Brown et al. (2020); Lourie et al. (2021), researchers have been seeking external repositories of common sense to boost performance further Mitra et al. (2019); Lin et al. (2021a); Xu et al. (2021). Existing common sense resources include both structured knowledge bases Fellbaum (2010); Speer and Havasi (2012); Sap et al. (2019a) and natural language statements Bhakthavatsalam et al. (2020). To the best of our knowledge, ours is the first work exploring common sense retrieval from a web corpus at this scale.

Discussion and Future Work

Harnessing the vast textual resources available online today through white-box retrieval may be the source of the next big break in NLP. In our current work, we propose to use a web snapshot as a universal, uncurated and unstructured knowledge source for multiple factual and common sense knowledge tasks at once. We see encouraging results even in the experimental setup with a strong pro-Wikipedia bias, which suggests that Sphere is a competitive knowledge source with the potential of pushing the state of the art—especially for tasks with diverse information needs. At the same time, while remaining closer to the needs of real-world applications, our setup exposes limitations of existing retrievers, providing a challenging test bed for future innovations. One of the key problems, which we aim to address in the future, regards the quality of retrieved information. Using Wikipedia as the knowledge source allows researchers to assume the high quality of the corpus documents. When transitioning to a web corpus, we no longer have the certainty that any document is good, truthful or unique, or that a certain gold document containing all the necessary information even exists. Future work should focus on the ability of the models to assess the quality of the retrieved documents, handle duplicates, detect potential false claims and contradictions, prioritize more trustworthy sources and refrain from providing the answer if no sufficiently good evidence exists in the corpus.

References

Appendix A Appendix

Slot-filling tasks suffer the biggest drop in downstream performance (see Table 4 in the main paper) when moving from Wikipedia to Sphere, which we investigate further by exploring their per-predicate accuracy (see Table 8). We observe a high variance in accuracy for the most common predicates, with those referring to more general concepts (e.g. crosses, country) scoring higher than more specific ones (e.g occupant, performer). We further note that all non-slot filling tasks incorporate a notion of input popularity in the data collection process. In FEV, claims were collected for \sayapproximately 50,000 popular [Wikipedia] pages. These consisted of 5,000 from a Wikipedia \saymost accessed pages list and the pages hyperlinked from them, NQ contains aggregated, real-world search engine questions, in HoPo, authors \saymanually curate 591 categories from the lists of popular pages by WikiProject to source their questions from. Finally, TQA questions were collected from trivia-related websites and further filtered to only include questions with high-quality search results. On the contrary, both T-REx and zsRE triplets were sourced from unfiltered WikiData snapshots, and further sampled uniformly to match KILT size limits. This suggests the knowledge coverage on the web is not balanced, with popular topics receiving more representation than rare ones.

TriviaQA.

A Sphere-grounded FiD model achieves a SOTA performance, beating our best Wikipedia-based model by over 6 points. We note that TQA is an outlier among other datasets and it can be considered one of the least Wikipedia-dependent of all KILT tasks. Questions and answers in TQA were created independently by trivia enthusiasts and only distant supervision was applied to collect Wikipedia evidence. We test a hypothesis that the Sphere advantage over Wikipedia might result from the fact that it would contain trivia websites with questions from the dataset. We find this not to be the case though: filtering out passages which contain input questions verbatim from the result sets of respective samples does not meaningfully impact downstream performance (see Table 9 for more context).

A.2 Wikipedia vs. the web

Excluding Wikipedia URLs from Sphere was an early design decision. However, Wikipedia text dissemination on the web goes beyond Wikipedia itself. We apply a simple ngram filtering heuristic testing if a web passage has at least one 8-gram overlap with a Wikipedia passage to establish if it was based on Wikipedia (a method inspired by Radford et al. (2019)). We will refer to such a passage as wiki-based. First, we note that as much as 5% of passages in our web corpus are wiki-based, adding up to almost 46M passages in total while the original Wikipedia corpus contains only 22M passages. This surprisingly high number can be partly explained by how Sphere was constructed - the head CCNet tier we used contains the documents with the lowest perplexity under a Wikipedia-based language model, favoring the inclusion of wiki-based passages into the corpus. We further note that almost 47% of the passages present in the KILT knowledge source inspire at least one web passage in Sphere, suggesting that big chunk of Wikipedia has been copied somewhere on the web.

Wikipedia bias in retrieval.

We then look at the median number of wiki-based passages retrieved from Sphere for respective datasets in Table 10. It turns out that all retrieval methods have a bias towards Wikipedia - the average median number of wiki-based results retrieved by the BM25 retriever is 12.1 and it increases sharply for the DPR-based methods, with 22.6 for DPRmulti and 24.2 for DPRweb. DPRmulti is a Wikipedia retriever so it is not surprising that it is biased towards wiki-based passages. However, it is unexpected that fine-tuning leads to a retriever yielding even more wiki-based results than the original. The analysis from the previous paragraph may shed some light here: we estimate that as many as 34% of DPR-based and 22% of the BM25-based web training samples include wiki-based passages, so the fine-tuning process will reinforce the Wikipedia bias present in the baseline retrievers.

Impact of Wikipedia on Sphere results.

Finally, we seek to establish how much of the Sphere performance is thanks to the wiki-basesd passages contained in the corpus. In Table 11, we present the relative change in downstream results for Sphere-based models if we use the more aggressive ngram filtering strategy. We generally see worse downstream results, however, the drop is not as dramatic as we would have expected. In particular on TQA, even though suffering a small drop, the BM25-based architecture still obtains a SOTA performance with 77.46 points of exact match. These observations leave us optimistic about usefulness of the web as a knowledge source for KI-NLP tasks.

How to treat wiki-based passages when comparing web-based with pure Wikipedia-based solutions remains an open question. One may argue that aggressive filtering of wiki-based passages from the web would be the right course of action. At the same time, the quality of wiki-based copies is often poorer than the original and may degrade over time as Wikipedia gets updated. In a small subset of cases, we may also be facing a situation where the web inspires Wikipedia, not vice-versa. Ideally, we would want to aim for a retriever which would be able to recognize these situations and favor more reliable sources.