Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, Jimmy Lin
Introduction
Information access is a fundamental human right. Specifically, the Universal Declaration of Human Rights by the United Nations articulates that “everyone has the right to freedom of opinion and expression”, which includes the right “to seek, receive and impart information and ideas through any media and regardless of frontiers” (Article 19). Information access capabilities such as search, question answering, and recommendation are important technologies for safeguarding these ideals.
With the advent and dominance of deep learning and approaches based on neural networks (particularly transformer-based large language models) in natural language processing, information retrieval, and beyond, the importance of large datasets as drivers of progress is well understood (Lin et al., 2021b). For retrieval models in English, the MS MARCO datasets (Bajaj et al., 2018; Craswell et al., 2021; Lin et al., 2022) have had a transformative impact in advancing the field. Similarly, for question answering (including the so-called “open-domain” retrieval-based variant), there exist many resources in English, such as SQuAD (Rajpurkar et al., 2016), TriviaQA (Joshi et al., 2017), and Natural Questions (Kwiatkowski et al., 2019). We have recently witnessed efforts in building resources for non-English languages, for example, CLIRMatrix (Sun and Duh, 2020), XTREME (Hu et al., 2020), MKQA (Longpre et al., 2021), mMARCO (Bonifacio et al., 2021), TyDi QA (Clark et al., 2020), XOR-TyDi (Asai et al., 2021), and Mr. TyDi (Zhang et al., 2021). These initiatives complement cross-lingual retrieval evaluations from TREC, CLEF, NTCIR, and FIRE that date back many years, largely focused on specific language pairs. Nevertheless, there remains a paucity of resources for languages beyond English. Existing datasets are far from sufficient to fully develop information access capabilities for the 7000+ languages spoken on our planet (Joshi et al., 2020). Our goal is to take a small step towards addressing these issues.
To stimulate further advances in multilingual retrieval, we have built the MIRACL dataset, comprising human-annotated passage-level relevance judgments on Wikipedia for 18 languages, totaling over 700k query–passage pairs on 77k queries. These languages—Arabic (ar), Bengali (bn), English (en), Spanish (es), Farsi (fa), Finnish (fi), French (fr), Hindi (hi), Indonesian (id), Japanese (ja), Korean (ko), Russian (ru), Swahili (sw), Telugu (te), Thai (th), Chinese (zh), and two “surprise” languages to be revealed later—are written using 11 distinct scripts, originate from 10 different language families, and collectively encompass over three billion native speakers around the world. They include what the research community would typically characterize as high-resource languages as well as low-resource languages.
Figure 1 shows an example of a query from MIRACL in Thai along with a relevant and a non-relevant passage. Along with the MIRACL dataset, our broader efforts (i.e., the “MIRACL project”) include organizing a WSDM 2023 Cup challengehttps://www.wsdm-conference.org/2023/program/wsdm-cup that provides a common evaluation methodology, a leaderboard, and a venue for a competition-style event with prizes. To provide starting points that the community can rapidly build on, we also share reproducible BM25, mDPR, and hybrid baselines as part of our Pyserini (Lin et al., 2021a) and Anserini (Yang et al., 2018) toolkits.
This paper focuses on providing a descriptive overview of the MIRACL dataset to coincide with our initial data release. It is our intention to periodically update this document with additional details about the WSDM 2023 Cup challenge as well as the broader MIRACL project over time.
Dataset Overview
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that spans 18 different languages (see Table 1). More precisely, the task we model is the standard ad hoc retrieval task as defined by the information retrieval community, where given a corpus , the system’s task is to return for a given an ordered list of top- documents from that maximizes some standard metric of quality such as nDCG. In our formulation, a query is a well-formed natural language question in some language (one of 18) and the documents draw from a snapshot of Wikipedia in the same language that has been pre-segmented into passages (and thus each passage has a fixed unique identifier).
Thus, our focus is monolingual retrieval across diverse languages, where the queries and the corpora are in the same language (e.g., Thai queries searching Thai documents), as opposed to cross-lingual retrieval, where the queries and the corpora are in different languages (e.g., searching a Swahili corpus with Arabic queries). As a terminological note, consistent with the parlance in information retrieval, we use the term “document” to refer generically to the unit of retrieval, even though in this case the “documents” are in actuality passages from Wikipedia.
In total, we have gathered over 700k manual relevance judgments (i.e., query–passage pairs) for around 77k queries across Wikipedia in these 18 languages, where all assessments have been performed by native speakers. Section 3 describes the annotation process in more detail. For each language, these relevance judgments are divided into a training set, a development set, and two test sets (that we call test-A and test-B); more details below. The MIRACL dataset is released under an Apache 2.0 License. To evaluate the quality of the system output, we use standard information retrieval metrics such as nDCG at a fixed cutoff and recall at a fixed cutoff.
The MIRACL dataset is built on the Mr. TyDi multilingual retrieval dataset (Zhang et al., 2021), which is in turn built on the TyDi QA dataset (Clark et al., 2020). Beyond the 11 languages originally covered by Mr. TyDi and TyDi QA, we included 7 additional languages. Details below:
Existing (Known) Languages: Mr. TyDi and TyDi QA cover 11 languages: Arabic (ar), Bengali (bn), English (en), Finnish (fi), Indonesian (id), Japanese (ja), Korean (ko), Russian (ru), Swahili (sw), Telugu (te), and Thai (th). We take advantage of existing queries in these languages as a starting point and provide “denser” annotations with respect to Wikipedia passages. That is, each query in Mr. TyDi has on average only a single positive (relevant) passage (Zhang et al., 2021). For existing Mr. TyDi queries, MIRACL provides richer judgments (i.e., more positive as well as negative labels). Beyond more comprehensive annotations of existing Mr. TyDi queries, MIRACL also includes new queries (i.e., created from scratch) that have been annotated with the same methodology.
New (Known) Languages: To the languages in Mr. TyDi and TyDi QA, we add 5 new (known) languages: Hindi (hi), Spanish (es), French (fr), Farsi (fa), and Chinese (zh). For these languages, there are no existing resources to build on, and thus all data are generated from scratch.
New (Surprise) Languages: Finally, there are two new (surprise) languages whose identities are presently hidden, but will be revealed at a later date as part of the WSDM 2023 Cup challenge. The construction of data for the surprise languages follows the same procedure as the new (known) languages.
At a high-level, MIRACL contains two classes of queries: those inherited from Mr. TyDi and those that were created from scratch. For the first class of queries, the splits in MIRACL align with the splits from Mr. TyDi. In more detail:
Training sets: For the 11 existing Mr. TyDi languages, the training sets comprise subsets of the queries from the Mr. TyDi training sets (originating from TyDi QA). The main difference is that MIRACL provides richer annotations for more passages. However, our annotators were not able to find relevant passages for some queries from Mr. TyDi; we call these “invalid” queries, and they were removed from the MIRACL training sets (see Section 3.4). For this reason, there are fewer queries in the MIRACL training sets than in their corresponding Mr. TyDi splits. For the new languages (both known and surprise), the training data comprise entirely of newly generated queries.
Development sets: Similar to the training sets, the MIRACL development sets align with Mr. TyDi for existing languages, but with the “invalid” queries discarded (see above). For the new languages, the development sets comprise entirely of newly generated queries.
Test-A sets: The test-A sets exist only for the existing Mr. TyDi languages and align with the test sets in Mr. TyDi (once again, with “invalid” queries discarded).
Test-B sets: For all languages, the test-B sets are comprised entirely of new queries that have never been released before (compared to test-A sets, which ultimately draw from TyDi QA and thus have been publicly available for quite some time now). These queries will be used in the final evaluation for the WSDM 2023 Cup challenge, and can be viewed as a true held-out test set.
Detailed statistics for MIRACL are shown in Table 1. For each language, split combination described above, we report the number of queries (# Q) and the number of relevance judgments (# J), which includes both relevant as well as non-relevant passages. The size of each Wikipedia corpus is reported in terms of the number of passages (# Passages) and the number of articles (# Articles) from which they were derived. As the relevance judgements were provided on query–passage pairs, the total number of articles is provided primarily as a reference. The final column indicates whether the language is part of Mr. TyDi and TyDi QA, or is one of the new languages.
It is our expectation that there are sufficient examples in the training and development sets of MIRACL to train (specifically, fine-tune) transformer-based retrieval models. We believe that MIRACL will provide a valuable dataset for building and evaluating dense retrieval models such as DPR (Karpukhin et al., 2020), late-interaction models such as ColBERT (Khattab and Zaharia, 2020), as well as reranking models such as monoBERT (Nogueira and Cho, 2019) and monoT5 (Nogueira et al., 2020). For these classes of retrieval models, there is an emerging thread of research (Shi et al., 2020; MacAvaney et al., 2020; Nair et al., 2022; Zhang et al., 2022) we hope that MIRACL will further catalyze.
Dataset Construction
To build MIRACL, we hired native speakers of each language as annotators to provide high-quality queries and judgments. At a high level, our workflow comprised two phases: First, the annotators were asked to generate well-formed natural language questions based on “prompt” paragraphs. Then, they were asked to assess the relevance of the top- query—passage pairs produced by an ensemble baseline retrieval system.
An important feature of MIRACL is that our dataset was not constructed via crowd-sourced workers, unlike other previous efforts such as SQuAD (Rajpurkar et al., 2016). Instead, we hired 31 annotators (both part-time and full-time) across all languages. Each annotator was interviewed prior to being hired and was verified to be a native speaker of the language they were working in. Our team created a consistent onboarding process that included training the annotators on exactly the tasks they were asked to perform. Throughout the annotation process, we checked randomly sampled data to monitor annotation quality. We believe that this design decision yielded a higher quality dataset than could have been obtained by crowd-sourcing means. Our team began interviewing annotators in mid-April 2022; dataset construction began in earnest in late April 2022 and continued until the end of September 2022.
For each MIRACL language, we prepared a pre-segmented passage corpus from a raw Wikipedia dump. For the existing languages in Mr. TyDi, we used exactly the same raw Wikipedia dump as Mr. TyDi and TyDi QA (from January 1, 2019 for Thai and from February 1, 2019 for the others). For the new languages, we used the versions released on March 1, 2022. We parsed the Wikipedia articles using WikiExtractorhttps://github.com/attardi/wikiextractor and segmented them into passages based on natural discourse units using two consecutive newlines in the wiki markup as the delimiter.
Each passage is given a unique identifier based on an X#Y schema, where X refers to a unique Wikipedia article and Y refers to the passage (numbered sequentially) within that article. Thus, it is possible for a system to reconstruct the passage/article relationships. To be clear, however, our task is to retrieve relevant passages, and thus query–passage pairs form the basic annotation unit. The corpus for each language is distributed in JSON lines format.
A key difference between MIRACL and Mr. TyDi is the manner in which the corpus was prepared. Since Mr. TyDi derived passage-level relevance annotations from TyDi QA, it retained exactly those same passages. However, since TyDi QA was not originally designed for retrieval, it did not provide consistent passage segmentation for all of Wikipedia. Thus, the Mr. TyDi corpora comprised a mix of TyDi QA passages and custom segments that were heuristically adjusted to “cover” the entire raw Wikipedia dump. This inconsistent segmentation is a weakness of Mr. TyDi that we rectified in the design of MIRACL. The downside, unfortunately, is that annotated passages from Mr. TyDi may no longer exist in the MIRACL corpora. Thus, to take advantage of existing annotations, we “projected” relevant passages from Mr. TyDi onto our newly prepared corpora. This process is described below.
2. Annotation Workflow
MIRACL was created using a two-phase process adapted from TyDi QA, which was in turn built on best practices derived from previous work; see discussion by Clark et al. (2020). The two phases are query generation and relevance assessment, detailed below:
Query Generation. In the first phase, annotators were shown “prompts” from Wikipedia that provide contexts to elicit queries. The prompts were extracted from the first 100 words of randomly selected Wikipedia articles. To generate high-quality queries, we asked that the annotators avoid generating queries that are answerable by the prompts themselves. Annotators were asked to generate well-formed natural language questions that are likely (in their opinion) to have precise, unambiguous answers. They could skip prompts that they did not find to be “inspiring”.
Note that in this phase, annotators were asked to generate questions in batch based on the prompts. As a result of this workflow, at this point, the annotators have not yet examined any retrieval results (i.e., passages from Wikipedia), and it could be the case that their questions can not be readily answered by information contained in the corpus. We discuss this case in Section 3.4.
Relevance Assessment. In the second phase, for each query from the previous phase, we asked the annotators to judge the binary relevance of the top- candidate passages () from a baseline ensemble retrieval system that combined three separate models:
Lexical matching with BM25: We used the Anserini (Yang et al., 2018) implementation based on the open-source Lucene search library with default parameters and the corresponding language-specific analyzer provided by Lucene.
A single-representation bi-encoder model, mDPR (Karpukhin et al., 2020; Zhang et al., 2021): We trained an mDPR model using the Tevatron toolkit (Gao et al., 2022), where the model was initialized from an mBERT checkpointbert-case-multilingual-base on HuggingFace. and then fine-tuned using the training set of MS MARCO Passage (Zhang et al., 2022). Retrieval was performed in a zero-shot manner.
A late-interaction model, mColBERT (Khattab and Zaharia, 2020; Bonifacio et al., 2021): We trained mColBERT using the authors’ official repository.https://github.com/stanford-futuredata/ColBERT#colbertv1 Similar to mDPR, the model was initialized from the same mBERT checkpoint as mDPR and then fine-tuned on MS MARCO Passage. As with mDPR, retrieval was performed in a zero-shot manner.
For each query, we retrieved the top- passages () using each model. We then performed ensemble fusion by first normalizing all retrieval scores to the range $$ and then averaging the scores. A final ranked list was then generated from these new scores. Based on initial experiments, we found that annotating the top 10 passages for each query yielded a good balance in terms of obtaining diverse passages and efficiently utilizing annotator effort.
For queries from the 11 existing languages in Mr. TyDi, we augmented the set of passages to be assessed with “projected” relevant passages from Mr. TyDi. This allowed us to take advantage of existing annotations transparently in our workflow. To accomplish this, we used relevant passages from Mr. TyDi as queries to search the corresponding MIRACL corpus using BM25. As the passages in MIRACL and Mr. TyDi differ in terms of segmentation but not content, the top retrieved passages are likely to have substantial overlap with the annotated passage in Mr. TyDi (and hence are also likely to be relevant).We cannot simply assume that highly similar passages from MIRACL are also relevant.
We used a simple heuristic to determine how many of these results to re-assess. If the score of the top retrieved passage from MIRACL is 50% higher than the score of the passage ranked second, we are more confident that the top passage is a good match for the original relevant passage. In this case, we only add the top passage to the set of candidates that the assessor considers. Otherwise, we add the top 5 passages. Note that these “projected” passages are not specially identified from the perspective of the annotator. In all cases, they simply receive a set of passages to judge per query, without any explicit knowledge where the passages came from.
3. Fold Creation and Data Release
During the annotation process, all queries followed the same workflow. After the annotation process concluded, we divided MIRACL into training sets, development sets, test-A sets, and test-B sets, as described in Section 2.
For existing languages, the training, development, and test-A sets align with the training, development, and test sets in Mr. TyDi, and the test-B sets are formed by the newly generated questions. For the new (known) languages, all generated queries were split into training, development, and test-B sets with a split ratio of , , and . Note that there are no test-A sets for these languages. The identity of the surprise languages and associated relevance judgments are currently kept secret and will be released as part of the WSDM 2023 Cup challenge.
Presently, all training and development sets have been released to the public, including both queries and relevance judgments. There are no current plans to release the relevance judgments for the test-A and test-B sets.
4. Discussion
Since MIRACL inherits from Mr. TyDi, which was in turn built on TyDi QA, it is worthwhile to discuss important differences in the annotation workflow.
At the core, TyDi QA (Clark et al., 2020) is not a dataset to evaluate retrieval. Rather, it is best characterized as a dataset for the so-called machine reading comprehension task, where the system is given (query, text) pairs and asked to identify the answer within the text. In TyDi QA, candidate passages for annotation are selected from only the top Wikipedia article based on a Google search. In contrast, in MIRACL we draw candidate passages from across all Wikipedia articles. This makes relevant passages in MIRACL more diverse.
Since we asked annotators to assess the top- candidate passages per query, MIRACL annotations are richer and more diverse than Mr. TyDi judgments (which were mechanistically generated from TyDi QA as no new annotations were performed). Furthermore, we believe that explicitly judged negative examples are quite valuable (compared to, for example, implicit negatives in MS MARCO sampled from BM25 results), as recent work has demonstrated the importance of so-called “hard negative” mining (Xiong et al., 2021). Since the candidate passages already come from an ensemble comprising baseline neural models, these are particularly suitable examples for various contrastive learning techniques.
One important detail worth discussing: Since the query generation and relevance assessment phases were decoupled, for some fraction of queries in each language, annotators were not able to identify any relevant passage in the pairs provided to them. For MIRACL, we aimed to have at least one relevant passage per query, and therefore these queries were discarded from the final dataset. We refer to these queries as “invalid” for convenience.
We note that just because no relevant passage was found in the top-10 candidates presented to the annotator, it is not necessarily the case that no relevant passage exists in the corpus. However, we did spot check a few of these “invalid” queries using a variety of tools, including interactive searching on the web. At least based on our limited efforts, we were not able to find relevant passages. While retrieval models that are able to self-assign confidence to its output (including the special case where the model may believe that no answer exists in the corpus), assessing such a capability is challenging in a retrieval-based setting. Short of exhaustively evaluating the corpus (obviously impractical), we cannot conclude with certainty that no relevant passage exists.
Baselines
As neural retrieval models have gained in sophistication in recent years, the “software stack” for end-to-end systems has grown more complex. This has increased the barrier to entry for “newcomers” who wish to start working on multilingual retrieval. We believe that the growth of diversity of languages introduced in MIRACL should be accompanied by an increase in the diversity of participants.
The need to provide “easy to use” starting points and the ongoing quest to make research more reproducible are, we believe, the same challenge. To this end, our research group has devoted substantial effort in build two toolkits to support reproducible research: Anserini (Yang et al., 2018) in Java and Pyserini (Lin et al., 2021a) in Python.
For MIRACL, we make available in Pyserini three baselines to serve as foundations that others can build on. Baselines scores for these retrieval models are shown in Table 2 in terms of two standard retrieval metrics, nDCG@10 and Recall@100. In more detail:
BM25 (Robertson and Zaragoza, 2009), one of the retrieval models used in the ensemble system to generate candidate passages for annotation. Thakur et al. (2021) and Zhang et al. (2021) both show BM25 to be a robust baseline when evaluated zero-shot across domain and languages.
mDPR (Karpukhin et al., 2020; Zhang et al., 2021), another one of the retrieval models used in the ensemble system to generate candidate passages for annotation. DPR is representative of the family of dense retrieval methods that has proven to be effective for many retrieval tasks.
Hybrid combines the scores of BM25 and mDPR results. For each (query, document) pair, the hybrid score is computed as , where we set without tuning. Scores of BM25 and mDPR ( and ) are first normalized to $$.
Code, documentation, and instructions for reproducing these baselines are available at the MIRACL repository, https://github.com/project-miracl/miracl. It is our intention to release more baselines as the evaluation progresses.
Ongoing Work
The broader MIRACL project has only begun. The first step on our journey is the release of the training and development sets in the 16 known languages of MIRACL. These resources are now publicly available on the MIRACL website at http://miracl.ai/.
The next phase of our efforts will focus on finalizing the details of the WSDM 2023 Cup challenge, which will provide a common evaluation methodology, a leaderboard, and a venue for a competition-style event with prizes. It is our plan to support two tasks: retrieval on the known languages as well as the surprise languages. In the first case, data have already been released, while in the second case, participants will have only a very limited amount of time to rapidly develop retrieval capabilities. We are finalizing the “rules of the game” as well as the timing of the test set releases.
We invite the community to come join us and help make more MIRACLs!
Acknowledgments
This research was supported in part by the Natural Sciences and Engineering Research Council (NSERC) of Canada and a gift from Huawei. Computational resources were provided in part by Compute Ontario and Compute Canada.