From Little Things Big Things Grow: A Collection with Seed Studies for Medical Systematic Review Literature Search

Shuai Wang, Harrisen Scells, Justin Clark, Bevan Koopman, Guido Zuccon

Introduction

Systematic reviews are comprehensive reviews about a particular body of literature. Systematic reviews are used extensively in medicine for various reasons: by clinicians to make health decisions and by governmental and institutional policy and practice decisions about health topics. To ensure comprehensiveness, creators of systematic reviews employ a standardised approach to searching for literature in the form of Boolean queries. However, creating Boolean queries is often so complex that the authors of systematic reviews often engage with information specialists: highly specialised librarians skilled in searching for medical literature for systematic reviews. Boolean query could be considered one of the most important aspects of a systematic review: it not only controls how many potentially relevant studies are retrieved but, perhaps more importantly, the total number of studies to screen. In other words, because systematic reviews aim to be comprehensive, the researchers will screen (i.e., assess for relevance to be included in the final systematic review) every study retrieved by the Boolean query. To ensure that no relevant studies are left out in the screening phase and to account for assessor biases, it is common for at least two assessors to screen studies. Naturally, since there may be thousands, if not tens of thousands of studies to screen, and since the screening process is duplicated by independent assessors, the Boolean query must be as effective as possible.

One technique that information specialists may use to assist in creating effective queries is to use ‘seed studies’: exemplar documents known to the information specialists a priori and often provided by the researchers of the systematic review. Information specialists use seed studies in various ways: identifying potential terms or phrases to use in the Boolean query; validating their search by analysing the precision and recall of their search against seed studies. Although, as we will demonstrate in the following sections, seed studies may not be a good indicator for query effectiveness.

The techniques that information specialists use to design and validate their search using seed studies have not gone unnoticed by the Information Retrieval community. There have been several studies that have exploited seed studies for automatic query formulation (Scells et al., 2021) and screening prioritisation (i.e., ranking the set of retrieved studies) (Lee and Sun, 2018; Wang et al., 2022). However, one major limitation of these previous studies is that the collections that they used did not contain any seed studies. Instead, a small portion of the final relevant studies were held out as a set of ‘pseudo seed studies’. Given that these works use studies that are known to be relevant and included in a systematic review, we believe that the effectiveness of these methods is overestimated. However, methods that use pseudo seed studies are relying of documents that are known to be relevant and included in a systematic review; we posit and show that the effectiveness is thus overestimated. To address this issue, this resource paper presents a new Information Retrieval test collection for systematic review literature search; our novel contribution is that our collection includes real seed studies for each systematic review topic.

Our test collection enables the development of new Information Retrieval methods for systematic review literature search that depend on or utilise seed studies. To this end, we also investigate the impact of seed studies versus pseudo seed studies by reproducing two prominent methods from the literature (automatic query formulation and seed-driven document ranking). In addition, we demonstrate a new technique that can use seed studies which is often actually utilised when constructing a systematic review but overlooked in Information Retrieval methods: snowballing (otherwise known as citation chasing). This method is discussed in more detail in the sections that follow. We make our test collection and the results of all of our experiments and analysis available at github.com/ielab/sysrev-seed-collection.

Background & Related Work

Creating a systematic review is a highly time-consuming, costly and complex task, which involves numerous stages (Bullers et al., 2018). One of the initial stages in the process is searching and screening medical literature. This stage contributes the most to the costs of a systematic review (McGowan and Sampson, 2005), and it is at this stage that our collection is intended to be used. Figure 1 provides a high-level overview of the early stages of the systematic review process, and the possible uses cases or tasks that our collection enables. The figure shows that information specialists first use seed studies to create a Boolean query. The most common use for seed studies is identifying and extracting terms that are topically related to the systematic review. The Boolean query is then executed to retrieve a set of studies for manually screening by the researchers of the systematic review. In reality, there are two phases of screening: one at the abstract level, which we have described here; and one at the full text level, where the full-text of the abstract level studies are assessed for inclusion in the systematic review. In our collection, the data that we were provided did not contain the abstract level relevance information of studies, nor the studies that were marked as not-relevant at an abstract level and full text level. We will discuss the implications of this in more detail in the sections that follow. The stages that follow in the systematic review creation process (e.g., data extraction or study synthesis) are not relevant to this collection, but are impacted by the methods that are created using our methods. For example, one of the use-cases of our collection, screening prioritisation (i.e., ranking studies), can be used to begin the following stages earlier

2. Seed Studies

Seed studies are used extensively by information specialists for query formulation and validation in systematic review literature search (Clark, 2013; Hausner et al., 2012). The use of seed studies is somewhat related to various explicit relevance feedback mechanisms, e.g. Rocchio (Rocchio and Salton, 1965). In the Information Retrieval domain, explicit relevance feedback has been used in applications such as query expansion (Lavrenko and Croft, 2017; Kaptein et al., 2008; Colace et al., 2011), active learning (Pereira et al., 2020), and user modelling (Jayarathna et al., 2015).

Active learning using explicit relevance feedback has been extensively studied in the area of technology assisted reviews (Cormack and Grossman, 2016; Zhang et al., 2018; Cormack and Grossman, 2014, 2015; Lagopoulos et al., 2018; Lee and Sun, 2018; Zou et al., 2018; Abualsaud et al., 2018): a catch-all phrase for high-recall tasks such as systematic review literature search, legal search, patent search, among others. The difference between explicit relevance feedback and seed studies is subtle but important. Seed studies cannot be considered relevant because they (1) have not necessarily been assessed as such when provided to an information specialist; and (2) may not be studies suitable for a systematic review (i.e., systematic review typically synthesise the results of several randomised controlled trials).Randomised controlled trials with many patients are considered a higher form of evidence compared to studies like case reports which may pertain to a single patient. Therefore, although seed studies may be used in the place of explicit relevance feedback documents for certain applications such as query expansion, they should not be considered of equal weight in terms of relevance to a systematic review.

3. IR Systematic Review Collections

There are several test collections for research in systematic review literature search. These include i) the CLEF Technology Assisted Review collections (Kanoulas et al., 2017, 2018, 2019); ii) a similar test collection from Scells et al. (2017); iii) a test collection about systematic review updates (Alharbi and Stevenson, 2019); and iv) collections for tasks such as data extraction and summarisation (Norman et al., 2018). However, as we have raised earlier, none of these collections contain seed studies. This means that the true effectiveness is unknown of any existing methods that claim to use seed studies. Our collection enables the proper evaluation of existing and future methods that use seed studies.

4. Snowballing

In order to maximise comprehensiveness, one additional step in the systematic review creation process that is currently unexplored by current Information Retrieval methods is that of snowballing or citation chasing (Miew Keen Choong, 2014). Snowballing identifies new studies to screen from the references or citations of other relevant studies. Snowballing is often performed ‘backwards’ and ‘forwards’ over the references to identify all studies that cite a given study and all studies cited by a given study. Although commonly performed in practise, there are very few studies in the literature that explore the topic of snowballing (Prasad et al., 2018). Nevertheless, snowballing has been shown to contribute 51% of the included studies for a systematic review (Greenhalgh and Peacock, 2005). The reason we investigate snowballing is to provide a more complete set of relevance assessments. Our preliminary experiments showed that not all included studies were retrieved by the Boolean query. We obtained total recall for many of our topics only when snowballing. Therefore, we perform snowballing to collect a more realistic set of studies that would be screened by researchers of a systematic review.

Collection Details

The basis of our collection comes from our co-author at the Bond Institute for Evidence-Based Healthcare, who is a senior information specialist. Each topic is a systematic reviews created using Boolean queries developed by our co-author over the past five years. We were provided unstructured data that we organised into 40 topics.

Each topic in our collection has several attributes, as shown in Table 1. Firstly, we assign a unique ID to each systematic review in our collection. As our collection contains completed and published systematic reviews, we also include the title and the URL of the published review. We also include a short description of the search provided by our information specialist co-author.

The next set of attributes corresponds to the relevance assessments. The included studies represent studies that are relevant at the full-text level. This means that these studies were assessed as sufficient to be included in the final review. As such, it should be noted that evaluation of an Information Retrieval system using our collection is more ‘strict’ (i.e., it is naturally more challenging to identify those studies that will be included in the final systematic review than those that are potentially relevant at an abstract level). Additionally, we provide the set of retrieved studies that were retrieved by the Boolean query.

The final attribute of each topic corresponds to snowballed studies. We include two sets of snowballed studies for each topic: one corresponding to snowballing the seed studies and another corresponding to snowballing the retrieved included studies. The second set of snowballed studies corresponds to simulating the process researchers of a systematic would do in order to identify additional relevant studies after assessing the set of retrieved studies. Note that this is an overlooked process in Information Retrieval methods in this space. We take this opportunity to do an initial investigation of the impact snowballing has on retrieval effectiveness and integrate the comparison of snowballing seed studies into the investigation.

2. Data Processing

Several attributes of topics required data processing beyond extraction. Specifically, the included studies, the retrieved studies, and the snowballed studies required additional processing to generate.

The raw data that we were provided did not include any relevance assessment information. As such, we were required to create the relevance assessments for each topic. To do this we manually extracted the list of included studies that were used in the analysis of the systematic review. Note that this was not as simple as scraping the references of a systematic review, but involved manually matching the citations used in the analysis to references of published studies. Sometimes an included study is not available in PubMed. Here, we do not include these studies in our relevance assessments as we assume it is not retrievable.

2.2. Retrieved studies

The raw data that we were provided also did not include the studies that were retrieved by the Boolean query. As such, we reproduced the search for all topics to obtain a set of retrieved studies. We automated this step by using the Entrez API (Maglott et al., 2005). The date restrictions were applied to each Boolean query, therefore others should be capable of reproducing the set of retrieved documents for each topic if necessary.

2.3. Snowballed studies

Rather than manually snowballing studies, we use two tools: Citationchaser (Haddaway et al., 2021) and SpiderCite (Clark et al., 2020). We first use CitationChaser to format a list of studies into the specific format required by SpiderCite.The format is RIS; a standardised tag format that enables citation programs to exchange data, and the only format SpiderCite supports as input. Using SpiderCite, we obtain the list of cited and citing studies given the input set. Each study in SpiderCite has a DOI that we use to retrieve the study from PubMed.

3. Collection Statistics & Analysis

On average, 15 seed studies are used for query construction, with a median number of 12. The average number of included studies per systematic review is 29, with a median number of 11.5. When comparing the overlap of seed studies and included studies, we found that, on average, only 36.3% of seed studies are contained in the included studies, and make up 26.1% of all included studies. This finding suggests that even though seed studies are helpful for Boolean query construction, most are disregarded after the query construction phase. This demonstrates that treating included studies as pseudo seed studies as in prior work (Lee and Sun, 2018; Scells et al., 2021) may lead to inaccurate results.

3.2. Searching analysis

When using the Boolean queries and date restrictions to retrieve studies, the mean number of retrieved documents per query is approximately 1,326, with median 709. Across several topics, we also found that not all included studies could be found in the retrieved studies. We found that only 75.5% of all included studies across all topics can be found inside the retrieved studies. There are two possible reasons why some included studies can not be found in candidate documents. The first is that some included studies may have been identified through snowballing. The second is that some studies may have been identified by searching in a database other than PubMed (although such studies may still exist in PubMed).

However, topic #18 does not contain any included studies in the retrieved studies. This particular systematic review only has a single included study. For topic #18, the reason that only one study is included is because many studies screened as relevant at the abstract level encountered a high risk of bias in the full-text.

3.3. Snowballing analysis

Our collection includes two snowballed sets of studies for each topic. The first set corresponds to snowballed seed studies (seed-snowballing) and the other corresponds to simulated snowballing of the included studies in the retrieved set (screened-snowballing).

For the seed-snowballing set, we find that 35 topics retrieved at least one included study. The topics that did not retrieve any included studies using this snowballing technique were 46, 52, 53, 66, and 96. The average number of documents snowballed per query is 1142, generally smaller than that retrieved from the searched documents list, as the number of seed studies used are smaller.

For the screened-snowballing set, only 34 topics retrieve at least one additional included study. The six topics that did not retrieve any additional included studies already contained all included studies prior to snowballing. These topics were 7, 10, 17, 39, 64, and 66. These topics represent an interesting research problem: applying snowballing to these topics would have resulted in wasted time and effort by the researchers of the systematic review. We leave investigations into determining whether or not to apply snowballing for future work. After removing already screened studies, the screened-snowballing sets contain, on average, 1000 studies.

3.4. Effectiveness comparison of retrieval methods

Finally, we aim to investigate the effectiveness of different search methods when used alone and when combined. The results of this analysis are presented in Figure 2. These plots show that using the seed studies alone (i.e., no searching and no snowballing) achieves the lowest recall but the highest precision. Snowballing the seed studies does increase recall but dramatically lowers precision.

Many topics retrieve almost all included studies with the Boolean query. Yet, combining the retrieved studies with other methods further improves recall while lowering precision. Combining both snowballing methods with the retrieved studies obtains the highest recall and lowest precision. However, total recall is still not achieved. Thus, we analysed the recall for different queries in Figure 3.

We found that for some topics, the recall is unusually low. We investigated these topics and found that these systematic reviews used several medical literature databases other than PubMed to retrieve studies. This results in the inclusion of studies that exist in the PubMed database, but are not retrieved by the PubMed query. We also randomly chose some topics reaching total recall and found that even though some of them still use a combination of multiple medical literature databases, they all used PubMed as one of the search sources. Thus, it is possible to create a more effective query for these topics. We leave such a problem for future work.

Query Formulation

We begin our demonstration of the use-cases of our collection with a reproduction of two automatic query formulation methods that use seed studies, see 1 in Figure 1. We use the implementations of Scells et al. (2021). We compare both automatic query formulation methods when using seed studies and pseudo seed studies.

The query formulation experiments are based on two existing methods from the literature (Scells et al., 2020a, b). These methods are fully automated adaptations of manual or semi-automated procedures that information specialists use in practice. The first is called the conceptual method and is what the majority of information specialists use to formulate queries (Clark, 2013). The automated conceptual method (Scells et al., 2020a) takes as input a preliminary string for identifying salient terms (we use the title of the systematic review). The seed studies are then used to optimise the coverage of different combinations of terms expanded from those in the title. The second is called the objective method and is a more recent procedure that takes a statistical approach to query formulation (Hausner et al., 2012). At a high level, the automated objective method (Scells et al., 2020b) first identifies and ranks salient terms from seed studies using term frequency statistics of the seed studies and a background collection. Next, terms are filtered and added to Boolean clauses by tuning these statistics to a held-out portion of seed studies. Given that the objective method relies on a held-out portion and the conceptual does not, we run both methods for three iterations using different arrangements of seed studies so that both methods use the same set of seed studies each iteration. We run these three iterations twice: once for the real seed studies and once for the pseudo seed studies.

One other aspect of the automatic versions of the two query formulation methods is the notion of an instantiation. In other words, the inclusion or exclusion of different aspects of Boolean queries (e.g., including or excluding MeSH, or including or excluding phrases). To this end, we perform experiments for only the most effective instantiation of each method (Conceptual/Phrase, and Objective/Phrase/Recall/MeSH). We refer the reader to the original study (Scells et al., 2020b) for a comprehensive description of all experimental settings and implementation details that we have used.)

2. Results & Analysis

The results of our automatic query formulation reproduction study using our collection are presented in Table 4.2. We report precision, recall, and average number of studies retrieved.

We continue with another possible use-cases for this collection with a reproduction of a method that uses seed studies for screening prioritisation (i.e., ranking the set of retrieved studies), see 2 in Figure 1. The techniques was originally proposed by Lee and Sun (2018) and we use the reproduced implementation provided by Wang et al. (2022). We investigate the effectiveness of seed-driven document ranking (SDR) methods, again comparing seed studies with pseudo seed studies. Note that as per previous studies, SDR refers to the specific process of screening prioritisation and a method of screening prioritisation called the SDR method. We make this distinction clear by referring to ‘SDR’ as the task and the ‘SDR method’ as the ranking function.

The SDR method proposed by Lee and Sun (2018) is based on the observation that terms in relevant documents are more similar than terms of irrelevant documents. Thus, a ranking model is devised based on this observation which weights each term in the document based on the inter-study similarity. In the SDR method, a study’s relevance score is calculated by the sum of pre-computed term weighting multiplied by the likelihood of the term to appear (calculated using the query likelihood model — QLM). In previous research, two study representation models have been explored:

where clinical terms in a study are used.

Bag of words representations are more effective than the bag of clinical words representation (Wang et al., 2022). Therefore, we adopt the bag of words representation (as indicated by BOW) here. Adding to this we investigate a new aspect of SDR: seed studies versus pseudo seed studies.

Firstly, we assume a single included study as an available seed study for each systematic review topic, as per Lee and Sun (2018). As there are multiple included studies corresponding to every systematic review topic, the overall effectiveness of retrieval methods on each topic is calculated using the average effectiveness of utilising every included study. Using this leave-one-out cross-validation strategy, the results gathered tend to be more reliable and unbiased across all included studies in a systematic review topic (Sammut and Webb, 2010).Topic #18 only has one included study. We disregard this topic in the evaluation across all experiments.

1.2. Multiple pseudo seed studies

Using multiple pseudo seed studies is more effectiveness then using a single seed study for SDR (Wang et al., 2022). For our collection, we also perform SDR with multiple pseudo seed studies. We adopt the experimental settings of Wang et al. (2022) to evaluate the effectiveness of SDR when using multiple pseudo seed studies. We adopt the same seed study grouping strategy in which 20% of included studies are chosen using a sliding-window approach. The groups are then combined by concatenating their titles and abstracts to act as the input to the retrieval methods. The effectiveness on each topic is then calculated using the average effectiveness from all groups.

1.3. Seed studies

Using the seed studies in our collection, we can now realistically investigate the effectiveness of SDR. This experiment combines all the real seed studies by concatenating their titles and abstracts, as used in the multiple pseudo seed study experiments. The combined studies then act as an input for the SDR method and all of the included studies are used for evaluation.

1.4. Retrieval Methods

Apart from the original SDR method proposed by Lee and Sun (2018), we also perform experiments using several baselines: BM25, a query likelihood model (QLM), and a word embedding-based model (AES). As in the previous SDR papers, we used two pre-trained word embeddings for the AES method: one trained on PubMed and Wikipedia and one trained on only PubMed. Additionally, we also include fusion methods used in the original paper to interpolate the SDR method with AES using the same parameters in the original article (α\alpha = 0.3).

1.5. Evaluation Measures

We evaluate the different arrangements of seed studies and methods with rank-based measures. We use the same evaluation measures as in our reproducibility study (Wang et al., 2022). In addition to MAP, which measures the ranking effectiveness of the entire list of studies, we also measure precision, recall, and nDCG at different cut-offs — {10,100,1000}, and Last Relevant% (LR%), which reports the percentage of studies that must be screened in order to identify all included studies.

2. Results & Analysis

The results of using a single pseudo seed study, multiple pseudo seed studies, and seed studies are shown in Table 3.

For single pseudo seed studies, the SDR method improves effectiveness for deep evaluation metrics (i.e., precision@{100,1000}, recall@{100,1000}, and nDCG@{100,1000}). SDR-BOW-AES-P is the most effective, apart from shallow evaluation measures (i.e., precision @10, recall@10 and nDCG@10.). For shallow evaluation measures, QLM achieves the highest effectiveness. This may be because the word embeddings based measures help alleviate vocabulary mismatch, improving recall but at the expense of early precision.

For multiple pseudo seed studies, SDR is the most effective on all evaluation measures except precision@1000, recall@1000, and LR%. Although, it is interesting to observe that the fused method is not able to achieve higher effectiveness than using a single retrieval method alone. This may be due to the poor performance of the AES method, and remains an interesting future challenge for automatically determining when to apply fusion. Using multiple pseudo seed studies is always more effective than single seed studies.

All the aforementioned results are in line with our previous reproduction study Wang et al. (2022), with the exception of the finding here that result fusion did not increase effectiveness.

2.2. Comparison of pseudo-seed results and real-seed results

Finally, we found that using real seed studies verses multiple pseudo seed studies dramatically impacts effectiveness. Firstly, even though the effectiveness of using seed studies still outperforms the use of a single pseudo seed study in some methods, the effectiveness of seed studies is significantly worse than using multiple pseudo seed studies. One possible explanation is that seed studies are not always relevant to the systematic review topic (pseudo seed studies are by definition). Using a non-relevant seed study could degrade the quality of results. This effect is highlighted in Figure 4, where the effectiveness of using all seed studies is closer to using a single pseudo study than it is to using multiple pseudo seed studies.

Secondly, using multiple pseudo seed retrieval with the SDR and SDR fused methods consistently outperforms other retrieval methods. However, when using real seed studies, the QLM method is most effective. Our explanation is that while the SDR method still captures the semantic meaning of terms through term weighting, one common use of seed studies in systematic review creation is term extraction, which is better represented by methods like QLM.

In conclusion, pseudo seed studies are not representative of real seed studies and this impacts seed-drive retrieval (SDR). That is why it is important to have test collections with real seed studies such as the one provided in this paper.

3. Overall Findings

Using pseudo seed studies produces unrealistic results compared to using seed studies (i.e., the results are higher than if one used seed studies). This is likely a result of the fact that the pseudo seed studies are an example of explicit relevance feedback, whereas seed studies are more akin to pseudo relevance feedback.

The choice of ranking models in this task (e.g., BM25, QLM, SDR, AES) is dependent on the kind of seed studies used. This may be due to how terms are weighted (i.e., relevant terms are likely to appear in pseudo seed studies but are less likely to appear in seed studies).

Our last demonstration of the use-cases of our collection is a new technique that we have devised specifically for this paper: ranking with snowballing, see 3 in Figure 1. Our collection provides two snowballing sets for each topic. Using seed studies and the two snowballing sets, we investigate the impact of snowballing when using SDR. We investigate two use cases of snowballing within the context of SDR:

The effectiveness of SDR on a combined set of seed-snowballing studies and retrieved studies.

The effectiveness of SDR on the screened-snowballing set.

These two use cases demonstrate (1) the effectiveness of ranking combined seed-snowballed and retrieved studies sets (i.e., integrating the seed-snowballed set into the retrieved studies and then ranking); and (2) the effectiveness of ranking post-screening (i.e., simulating the ranking of the screened-snowballed set using the screened retrieved studies as input).

In Section 5 we investigated the effectiveness of SDR using different arrangements of seed studies. We found that when real seed studies are used, the QLM method outperforms the SDR method in almost all evaluation measures. In this experiment, we investigate if similar results arise when the seed-snowballing set is combined with retrieved studies. As the first experiment, we used the same experimental setting as in Section 5.1.3. We first combine the retrieved studies and the seed-snowballing set. Next, we concatenate titles and abstracts of seed studies as input to the ranking methods. We report the same evaluation measures as described in Section 5.1.5.

1.2. Screened snowballing document ranking

As the second experiment, we simulate the process of screening prioritisation for the screened-snowballed set. Included studies in the retrieved set plus the seed studies are used as input for SDR. Given the difference in performance of the two sets of studies (i.e., retrieved included studies and seed studies), one could consider different term weighting functions depending on the study. However, we leave such investigation for future work. We perform our evaluation on the included studies that did not appear in the retrieved set of studies. This means that the results for this experiment are not comparable to the results of the other screening prioritisation experiments that appear earlier in this paper. We report a subset of the evaluation measures as described in Section 5.1.5.

2. Results & Analysis

Results for this experiment are shown in Table 4. The most effective methods are QLM and SDR, with SDR more effectiveness on shallow measures and more effectiveness on deeper measures. These results are highlighted further in Figure 5. When comparing these combined results with the seed studies result from Table 3, we found that all methods can achieve higher recall{@100,@1000} and LR%, which suggests that adding the seed snowballing set can significantly increase the number of relevant studies retrieved.

Despite the fact that there are more studies overall when combining the seed-snowballing set with retrieved studies, it is worth it in terms of the overall improvement in effectiveness. In practise, it is beneficial to combine the results of seed-snowballing with retrieved studies for screening prioritisation.

2.2. Screened snowballing document ranking

Results for this experiment are shown in Table 5. The two best performing methods here were QLM and the AES methods. QLM is the most effective for MAP and nDCG@100. Meanwhile, fusion methods are the most effective for precision@100, recall@100, and LR%. These results are highlighted in Figure 6. These findings demonstrate that QLM is effective for the SDR task when ranking screened-snowballing studies.

3. Overall Findings

From the ranking with snowballing experiments, we found that:

Combining the seed-snowballing set and retrieved studies is beneficial for SDR.

QLM and AES are best for screened-snowballing document ranking. However, there is room for future work in determining the optimal combination of studies to use for SDR.

We present a new test collection to properly evaluate systematic review literature search methods which uses seed studies. In addition to our test collection that includes seed studies, we provided a detailed analysis of our collection. Here, we found that as a unique set of studies before query construction, only a small portion of seed studies will be included in the systematic review. We also investigated the impact of seed studies by reproducing two existing methods that use seed studies. Our experiments show that using pseudo seed studies overestimates the effectiveness.

The test collection also enables an investigation of the difference between snowballing seed studies and screened studies. Here, we found that ranking a combined lists of studies from seed-snowballing documents and searched candidate documents may further boost the effectiveness of ranking models.

The test collection makes two important contributions: (1) it enables a considerably more realistic evaluation of methods that use seed studies (i.e., one can use all included studies for relevance assessments, instead of requiring a held-out portion as pseudo seed studies); and (2) it provides realistic data to develop or train new methods that use seed studies. Seed studies are vital for effective query formulation for information specialists and are commonly used. Despite this, many existing Information Retrieval test collections for systematic review literature search do not contain seed studies. Our test collection will promote the development and realistic evaluation of methods that seed studies can be used to improve systematic review literature search. This includes methods we already explored in this paper like screening prioritisation (Lee and Sun, 2018; Wang et al., 2022) and query formulation (Scells et al., 2020a), and some we leave for future works such as active learning (Cormack and Grossman, 2015) or MeSH term suggestion (Wang et al., 2021), etc. Such methods can have considerable real-world impacts, as systematic reviews are highly time consuming and costly. Cheaper and faster systematic reviews can have dramatic implications for patient outcomes and institutional policy decisions that affect the health decisions of entire countries.

Shuai Wang is supported by a UQ Earmarked PhD Scholarship and this research is funded by the Australian Research Council Discovery Projects programme ARC DP DP210104043.