Automated MeSH Term Suggestion for Effective Query Formulation in Systematic Reviews Literature Search

Shuai Wang, Harrisen Scells, Bevan Koopman, Guido Zuccon

Introduction

A medical systematic review is a comprehensive review of literature for a highly focused research question. Systematic reviews are seen as the highest form of evidence and are used extensively in healthcare decision making and clinical medical practice. In order to synthesise literature into a systematic review, a search must be undertaken. A major component of this search is a Boolean query. The Boolean query is often developed by a trained expert (i.e., an information specialist), who works closely with the research team to develop the search, and usually has some knowledge of the domain been searched. The most commonly used database for searching medical literature is PubMed. Due to the increasing size and scope of these databases, and of PubMed in particular, the Medical Subject Headings (MeSH) thesaurus was developed to conceptually index studies . MeSH is a controlled vocabulary thesaurus arranged in a hierarchical tree structure (specificity increases with depth in a parent→\rightarrowchild relationship, e.g., Anatomy→\rightarrowBody Regions→\rightarrowHead→\rightarrowEye…etc.). Indexing and categorising studies with MeSH terms enables queries to be developed which incorporate both free-text keywords and MeSH terms — enabling more effective searches. The use of MeSH terms in queries has been shown to be more effective than free-text keywords alone , e.g., they increase precision and are far less ambiguous than free-text . However, it is still difficult even for expert information specialists to be familiar with the entire MeSH controlled vocabulary — at the time of writing, MeSH contains 29,640 unique headings.

PubMed has attempted to overcome this difficulty by developing a method called Automatic Term Mapping (ATM). ATM is an automatic query expansion method which attempts to seamlessly map free-text keywords in a query to one of the three categories (index tables): MeSH, journal name or author name . Although ATM is applied by default for all queries issued to PubMed, it has several semantic limitations: it is inaccurate when used to expand free-text acronyms into MeSH terms ; it produces different MeSH expansions even though synonymic free-text terms are used ; and has difficulty disambiguating between MeSH terms and journal names . Despite these limitations, the use of ATM for MeSH term suggestion has been shown to increase the precision of free-text searches in the genomic domain , and is the state-of-the-art method for the MeSH Term suggestion task. However, its use has, to the best of the authors’ knowledge, not been empirically evaluated in the context of improving the effectiveness of systematic review literature search queries.

Recent advances in the use of pre-trained language models (PLMs) such as BERT , T5 , and GPT-3 have delivered state-of-the-art performance in many natural language processing tasks. Typically, a pre-trained language model is trained on a large corpus using the transformer architecture to “get familiar” with language representations. Then the model is fine-tuned to downstream tasks to perform with high effectiveness across the target task. The transformer architecture is an encoder-decoder model training structure that does not use recurrence and convolutions . Prior work showed that using PLMs can significantly increase effectiveness in ad-hoc search as well as in professional search .

In this article, we introduce the task of MeSH term suggestion for Boolean queries used in systematic review literature search This article is an extension of our previous work published at the 2021 Australasian Document Computing Symposium .. We model this task within the context of an information specialist looking for MeSH terms to add to a query without MeSH terms currently present. We also propose a framework to evaluate the effectiveness of the suggestion of MeSH terms on established collections of systematic review literature search queries. This article adds to a recent stream of research that has focused on computational methods for the assisted formulation or refinement of Boolean queries for systematic review creation, and more generally to research on computational methods for technology-assisted reviews . Furthermore, we propose two categories of methods for the MeSH term suggestion task, including methods based on the BERT pre-trained language model and methods not based on BERT (lexical methods). We show that our methods suggest MeSH terms that outperform the effectiveness of the MeSH terms selected by the information specialists and included in the original queries. Our methods are readily integrable into tools for information specialists to help with the construction of systematic review Boolean queries.

The introduction of the new task of suggesting MeSH terms for systematic review literature search (Boolean queries), modelled within the context of an information specialist looking for MeSH terms to add to a query without MeSH terms present.

The formulation of MeSH term suggestion methods to help information specialists and researchers to construct Boolean queries for systematic review creation.

An empirical evaluation of the effectiveness of different MeSH terms suggestion methods

An understanding of how the MeSH terms suggested by the proposed automatic methods differ from those originally selected by information specialists formulating the query.

Material and methods

We start by outlining the task of MeSH term suggestion for Boolean queries that do not already contain MeSH terms. We assume the user has entered a Boolean query without MeSH terms. A Boolean query can be viewed as a tree where Boolean operators (e.g. AND, OR) represent the internal nodes of the tree, while free-text atomic clauses and MeSH Terms are the leaves. Free-text atomic clauses are one or more words that express a concept, e.g., a disease, a treatment or a population aspect. We call each of the first level nodes of the tree (i.e. the nodes at depth 1) a query fragment. Typically, a query fragment represents an individual aspect of an information need ; specifically, each query fragment corresponds to a different PICO element, i.e. population, intervention, control, and outcome . These concepts are shown in Figure 1. The task of MeSH term suggestion is to identify appropriate MeSH terms to be added as leaves to a query fragment. In this article, we suggest MeSH terms for each query fragment independently of each other. We leave the investigation of query fragment dependencies concerning MeSH term suggestions for future work.

Figure 2 gives an intuition for how we obtain query fragments from a Boolean query, how MeSH terms are suggested for a given query fragment, and how we perform defragmentation to construct a new Boolean query that includes MeSH terms. The figure shows that after fragmentation (i.e. the process of deriving query fragments), we remove all the MeSH terms from each query fragment. We then apply a MeSH term suggestion technique which adds new MeSH terms into a query fragment. The new query fragments that now contain suggested MeSH terms are then defragmented by combining all of the query fragments corresponding to the original query with the AND operator.

This work extends our existing line of research into MeSH term suggestion , where we previously developed several techniques that depend on pre-existing lexical matching systems. One limitation of these systems is their dependence on manually crafted rules that are expensive to create and have limitations in terms of how words are matched to MeSH terms (e.g., spelling variants, acronyms, misspellings). This article instead investigates the use of pre-trained language models, i.e., BERT, for the task of MeSH term suggestion. These neural models have been shown to be resilient to the shortcomings of lexical-based systems . However, neural models have their own limitations, particularly requiring large amounts of training data. The following sections first provide a brief overview of our existing lexical-based techniques and then describe our new neural techniques in detail, specifically addressing the need for ad-hoc training data.

2 Lexical MeSH Term Suggestion

Our lexical-based methods are formulated as a pipeline of three steps: retrieval, ranking, and refinement. The following sections provide a brief overview of each of these steps. For a more comprehensive discussion of the lexical-based methods, refer to our previous work .

The first step in our MeSH term suggestion pipeline is the retrieval of MeSH terms. The retrieval of MeSH terms is facilitated by three different methods:

The entire free-text only query fragment is submitted to the PubMed Entrez API for automatic term mapping (ATM). This is the default system used by PubMed for automatically adding MeSH terms to queries.

Each free-text atomic clause in a query fragment is submitted to MetaMap .Version 2018 with options set to default values. The results are filtered to only include those entities derived from the MeSH source. All of the mapped MeSH terms are recorded for each of the free-text terms in a query fragment. Additionally, the score is recorded for each MeSH term.

We index UMLS version 2019AB using the MRCONSO, MRDEF, MRREL, and MRSTY tables. into Elasticsearch v7.6. Each free-text atomic clause in the query fragment with MeSH terms removed is submitted to the Elasticsearch index. The results are filtered to only include synonyms of concepts derived from the MeSH source. Additionally, the BM25 score is recorded for each MeSH term.

For the MetaMap and UMLS approaches, the same MeSH term may be retrieved multiple times for a given free-text fragment. To overcome this issue, we re-score the MeSH terms using rank fusion (CombSUM) . The intuition for this re-scoring is that highly common MeSH terms that also obtain a high score from these retrieval methods should be scored highly overall (thus ranked higher than common MeSH terms and highly scoring MeSH terms).

Ranking

Once MeSH terms have been retrieved, they are ranked according to the approach for entity ranking described by Jimmy et al. by adapting features proposed by Balog . In total, we use eleven entity features. Positive instances correspond to MeSH terms in the original query fragment; negative instances correspond to MeSH terms not in the original query fragment (binary labels). With features and instance labels, we train a learning-to-rank (LTR) model for each retrieval method. In addition to LTR, we also investigate a rank fusion approach , where we combine the normalized MeSH term suggestion scores from each of the three methods to produce a new ranking that incorporates the highest ranking MeSH terms from each method. The intuition for investigating rank fusion in this context is that each method may retrieve different MeSH terms; and those terms may be ranked differently each time. Therefore, we boost MeSH terms that are retrieved and ranked highly by multiple methods.

Refinement

Finally, we seek to refine the suggested MeSH terms by estimating a rank cut-off. We do this using a score-based gain function. Formally, the cumulative gain CGCG for a MeSH term at rank pp is

where the score for a MeSH term is equal to 1−normalised score1-normalised\ score (i.e., min-max normalisation) for the MeSH term.

We tune a parameter, κ\kappa, for each retrieval method which controls the percentage of total CGCG allowed to be observed before the ranking is cut-off (i.e., a refinement of the ranking). We tune κ\kappa from 5% to 95% in increments of 5%. The intuition for re-scoring MeSH terms becomes apparent when used with the κ\kappa parameter: the highest-ranking MeSH term will receive a score of 0, resulting in at least one MeSH term suggested for every query fragment.

Note that MeSH terms may share the same score, i.e., they may be tied. We take a conservative approach to account for the problem of tied MeSH terms at the boundary of the cut-off specified by κ\kappa. Whenever we encounter ties, we treat all of the tied MeSH terms as a single accumulation of gain that equals the summed gain across the scores of the tied MeSH terms. This treatment has the effect that tied MeSH terms account for much larger accumulations of gain. Therefore, tied MeSH terms at the top of rankings are more likely to be included in the cut-off than tied MeSH terms at the bottom. In essence, either all tied MeSH terms are considered within the cut-off (i.e., ties at the top of the ranking), or no tied MeSH terms are considered (i.e., ties at the bottom of the ranking).

3 BERT MeSH Term Suggestion

Next, we extend our MeSH term suggestion methods using fine-tuned PLM models. Firstly, PLM models are typically chosen from the same domain in which the task is conducted.

We show the architecture of our fine-tuning and inference processes in Figure 4. We use BioBERT as the base PLM, as the context of this paper is medical systematic reviews. BioBERT is a PLM pre-trained on PubMed abstracts and PubMed Central (PMC)PubMed Central is the repository containing full-text articles of the open-access part of the PubMed database. full-text articles using the BERT training architecture . After fine-tuning, BioBERT has achieved state-of-the-art performance on many medical-related tasks, including biomedical named entity recognition, relation extraction and question answering .

Ideally, training data closely related to the target task should be used to fine-tune a PLM to achieve the highest effectiveness. Ideally, in our case, we would use professionally constructed medical systematic review Boolean queries to fine-tune our model. However, PLMs are typically data-hungry and require a large number of labelled training samples. In systematic review literature search, several public datasets are available with Boolean queries, such as the CLEF TAR collections , the collection of Wang et al. , and the collection of Scells et al. . Between these datasets, however, only 253 unique topics would be available to train the model: an insufficient amount to effectively fine-tune a BERT model.

Instead, we create training samples by approximating the target task using data obtained from PubMed. We use the publicly available PubMed baseline to obtain the metadata about all published articles up to the start of 2022. The metadata contains information such as the title and abstract, but importantly for this work, it also includes author-assigned keywords and the relevant MeSH terms for an article. We use the assigned keywords and MeSH terms for every article in the PubMed dataset to approximate the task of MeSH term suggestion. To maximise the amount of training data, we also extract keywords from the title (as not all PubMed articles contain keywords). To tokenise titles, we use the process described by Wang et al. . Firstly, we tokenise the title using Gensim , and then we remove stopwords too using NLTK . We use the toolkit proposed by Gao et al. to develop a dense retriever to suggest MeSH Terms. The model is fine-tuned with localized contrastive loss using triples of <ka,i,ma+,ma−><k_{a,i},m_{a}^{+},m_{a}^{-}> where aa is a PubMed article, ka,ik_{a,i} is the iith keyword in the PubMed article, ma+m_{a}^{+} are the MeSH terms for the PubMed article, and ma−m_{a}^{-} are ten randomly sampled MeSH terms from the MeSH thesaurus. Many MeSH terms contain spaces or punctuation. Our model considers each MeSH term a unique token in the model vocabulary. Once the model is fine-tuned, we obtain an encoding for all MeSH terms. At inference time, we create an encoding for a keyword to obtain a score using the [CLS] token for all MeSH terms. Thus, our method scores and ranks all MeSH terms given a keyword.

Ranking Suggestions

The goal of MeSH term suggestion is to suggest MeSH terms for each query fragment. However, the result from the BERT suggestion method consists of a ranked list of MeSH term suggestion for each free text atomic clause. We need to combine the rankings for each MeSH term. We formulate this combination task into two steps, (1) choosing how we represent a MeSH term ranking, and (2) choosing where to cut off the ranking. We present an overview of the combination task in Figure 3.

First, we choose the best way to represent a ranking, which means deciding if MeSH terms should be suggested individually for every free text atomic clauses, as a whole for every fragment, or using other heuristics to decide how the representation should be computed. We designed three ranking representation methods:

Atomic BERT: Firstly, we treat suggestions for each free text atomic clause individually, essentially applying no strategy to combine the suggestions.

Fragment BERT: Next, we study the combination of all MeSH term rankings for a given query fragment. We apply rank fusion (normalised CombSUM ) to all of the free text atomic clauses in a query fragment. For computational reasons, we only use the top 20 MeSH terms for each free text atomic clause.

Semantic BERT: Finally, we study semantically grouping free text atomic clauses and apply the same rank fusion technique as above, but this time to each group. We show an example of a semantic group in Table 1. To derive semantic groups, we first take all free text atomic clauses from the fragment and obtain word2vec embeddings for each free text atomic clause. Then we compute cosine similarities between all free text atomic clause to decide if they are semantically related. In our experiments, we apply a threshold of 0.7 on the similarity. We use a word2vec model pre-trained on PubMed and Wikipedia . There are two reasons we use word2vec rather than BERT for semantic groups. First, if we apply our proposed BERT model, we note that we fine-tuned using semantic pairs of free text atomic clauses and MeSH terms: thus, calculating the similarity between two free text atomic clauses can cause a model mismatch. Secondly, the use of an additional BERT model will increase the latency in producing suggestions at inference time, as each free text atomic clause needs to be encoded twice.

Second, we choose where to cut off the ranking of MeSH terms from the ranking representations. We propose four strategies to cut off MeSH term rankings:

First only (FO): The first MeSH term of the ranking is selected for each ranking representation.

Same as free text atomic clauses (SA): The number of MeSH terms selected equals the number of free text atomic clauses in each fragment (i.e, only applicable to Fragment BERT).

Same as original (SO): The MeSH terms selected equals the number of MeSH terms in the query fragment prior to removing MeSH terms (i.e, only applicable to Fragment BERT).

Linear (LN): The number of MeSH terms selected is learnt using a linear function with respect to the number of free text atomic clauses in the fragments (i.e, only applicable to Fragment BERT).

4 Evaluation

The end goal of a systematic review literature search is to find all of the relevant literature at the minimum cost. Thus, an effective Boolean query minimises the number of documents retrieved while maximising the retrieval of relevant documents. In our MeSH term suggestion task, we use the retrieval effectiveness of defragmented Boolean queries to evaluate MeSH term suggestion.

The MeSH terms included in the original query have been derived often after careful consideration by expert information specialists. We therefore consider how the MeSH terms included in the original queries differ from those suggested by the methods investigated in this work; specifically, we measure the overlap between the suggested MeSH terms and the MeSH terms included in the original query. We note that a MeSH term that is not in the original query may not necessarily be less effective of a search term than one included in the original query.

To evaluate the effectiveness of the suggested MeSH terms for the task of systematic review literature search, once query fragments are defragmented, the retrieval effectiveness is evaluated using typical systematic review literature search evaluation measures: precision, recall, and Fβ{F\beta}, with β={1,3}\beta=\{1,3\}. The PubMed Entrez API is used to directly issue defragmented Boolean queries to obtain retrieval results. As PubMed is constantly updated with new studies, we apply a date restriction to all queries for reproducibility purposes. We use the Jaccard index measure to evaluate the overlap of MeSH terms between those suggested by the investigated methods and those included in the original query.

For both evaluation settings (i.e., Boolean query retrieval and evaluation against original MeSH terms), we evaluate the lexical suggestion method in two settings: (i) all, where all retrieved MeSH terms are considered; and (i) cut, where the score-based cut-off is used. We also evaluate all BERT suggestion methods and compare their effectiveness with that of the original query and the lexical methods.

5 Experimental Setup

For our experiments, we use topics from the CLEF TAR task from 2017, 2018, and 2019 . 15 topics are discarded due to lack of MeSH termsDiscarded topics are: 2017: CD007427, CD010771, CD010772, CD010775, CD010783, CD010860, CD011145; 2018: CD007427, CD009263, CD009694; 2019: CD006715, CD007427, CD009263, CD009694, CD011768.. An additional topic is discarded because of retrieval issuesThe additional discard topic is 2017: CD010276., likely resulting from the fact that we translate queries automatically from one format (Ovid Medline) into another format (PubMed) . In total, we used 116 unique topics, as each year has partial overlap. For each topic, we automatically divide the Boolean query for that topic into query fragments . Each fragment contains at least one MeSH term. This results in a total of 311 unique query fragments (2.68 fragments per query on average). For each query fragment, we corrected any errors (e.g., spelling mistakes, syntactic errors), extracted MeSH terms, keywords, query fragments with MeSH terms, and query fragments without MeSH terms. For training the LTR model for each lexical method, we use the pre-split training and test portions from the CLEF datasets. The 2019 topics are also split on systematic review type (intervention and diagnostic test accuracy — indicated as ‘intervention’ and ‘dta’ respectively in the results), while those for 2017 and 2018 are all diagnostic test accuracy. We use the ‘quickrank’ library for LTR, instantiated with LambdaMART trained to maximise nDCG. We leave other settings as per default.

For learning the linear function of the BERT suggestion method to decide the cut-off value of the MeSH term ranking list, as described in Section 2.3, we use the training portions of the CLEF TAR datasets. First, we obtain all the fragments from the CLEF TAR training splits. We count the number of free text atomic clauses and MeSH terms in each fragment. We then perform linear regression on these numbers to determine a function for each CLEF TAR dataset. We show the linear regression in Figure 5.

Results

Results in this section are presented on the test splits of the CLEF TAR datasets (i.e., 2017, 2018, 2019-dta, 2019-intervention). We first analyse the search effectiveness of lexical methods versus our new BERT methods and then analyse the MeSH suggestion effectiveness compared to the MeSH terms originally used.

The results of the lexical methods presented in Table 4 are the same reported in our previous work . We discuss them briefly here for completeness. Unrefined methods generally contain higher recall than corresponding refined methods, with lower precision. This finding indicates that adding more MeSH terms in the query fragments can cause both more relevant and irrelevant studies to be retrieved. When compared using F1 and F3, the refined methods consistently outperform unrefined methods for each dataset. The U-CUT method achieves the highest effectiveness on CLEF 2017 and 2018, while A-CUT can achieve the highest effectiveness on CLEF 2019 dta; M-CUT for CLEF TAR 2019 intervention dataset in terms of F1 and F3. In terms of recall, the unrefined fusion method achieves the highest recall among all lexical suggestion methods. This gain in the recall is likely because unrefined fusion combines all of the MeSH terms suggested by the other three methods (ATM, MetaMap, and UMLS) using ‘OR’. This suggests that the unrefined fusion method is not beneficial for improving the precision of a Boolean query. However, suppose semi-automatic MeSH term suggestion can be used. Information specialists may be able to use the suggestion and apply their expertise to decide which MeSH terms may be included to achieve higher performance.

BERT Methods

We first compare the effectiveness of the BERT methods with the original Boolean query. The result shows that under all evaluation measures (Precision, F1, F3 and Recall), BERT methods can outperform the original query on CLEF TAR 2017, 2018 and 2019-intervention, while effectiveness is generally worse for CLEF TAR 2019-dta. Note that CLEF TAR 2019-dta only contains eight unique topics; the lower effectiveness is likely due to a handful of topics.

Next, we compare the effectiveness of BERT suggestions against lexical suggestions. When comparing with un-refined lexical methods, the effectiveness of BERT suggestions is comparable in terms of F1 and F3, showing substantial gains across all datasets. However, compared with refined lexical methods, BERT suggestions generally obtain comparable results to refined lexical suggestions, except in CLEF TAR 2019-dta, in which refined lexical suggestion methods achieve higher effectiveness. In terms of recall, BERT suggestions obtain slightly higher recall to un-refined lexical suggestions, but substantially higher recall than refined lexical suggestions.

As mentioned in Section 3.1, unrefined lexical methods are effective to achieve higher recall while refined lexical methods are effective to achieve higher Precision, F1 and F3. We find that the MeSH terms suggested by BERT can obtain similar recall effectiveness to un-refined lexical methods while F1 and F3 can be comparable to refined lexical methods. Therefore, compare with lexical methods, BERT methods may be preferred to suggest more effective MeSH Terms.

2 Impact of BERT Ranking Representations

We compare different ranking representations of BERT, including Atomic BERT, Semantic BERT and Fragment BERT. We use the same cut-off strategy to compare these three representations fairly. We find that the precision, F1 and F3 values of Fragment BERT are the highest among the three methods, while recall of Fragment BERT is the lowest. However, only one MeSH term is suggested for each fragment when Fragment BERT is used. This trade-off of precision and recall also suggests the same finding we described for lexical methods, where adding more MeSH terms can cause more studies to be retrieved. Between Semantic BERT and Atomic BERT, Semantic BERT is able to obtain higher precision while recall is lower than Atomic BERT. When comparing using F1 or F3, Semantic BERT always achieves higher effectiveness. Therefore, the use of Semantic BERT is preferred over Atomic BERT.

3 Impact of Cut-off Strategy

When comparing different cut-off strategies for BERT suggestionsOnly Fragment BERT is considered as SA, SO, LN are only applicable to Fragment BERT., we find that FO can consistently achieve the highest Precision, F1 and F3 compared to other cut-off methods. On the other hand, the recall value of using FO is the lowest among all other methods, indicating that the trade-off of precision and recall is again caused by the number of MeSH terms added to the query. For the other three cut-off strategies, including SA, SO, and LN, we find that SO and LN consistently outperform SA, suggesting that information specialists have an intuition for how many MeSH terms to add to a query.

4 Are Suggested MeSH Terms the Same as those in the Original Queries?

Next, we study the overlap between the MeSH term suggested by the considered methods and those included in the original query; this is reported in Table 2 and is measured with the Jaccard index. One immediate observation is that the overlap of all based methods is considerably higher than that of lexical methods. This observation is based on that the highest value of the Jaccard index in each dataset always appear in the BERT suggestion method. Moreover, when applying the SO cut-off strategy to Fragment BERT, the highest overlap is always obtained, which indicates that BERT suggestion methods also agree on the Terms chosen by systematic reviewers. .

The previous results reported in Table 4 highlighted that, in general, BERT methods were better than lexical methods in suggesting effective search terms; and these were more effective than those in the original queries, although differences were not statistically significant. These results, in conjunction with the findings in Table 2, indicate that the BERT methods identify very similar MeSH terms that present in the original queries – and the MeSH terms identified by BERT methods are more effective than those provided by other methods.

Intuitively, we further analyse whether search effectiveness and the suggestion of MeSH terms that are included in the original query correlate, meaning that Mesh Terms used in the original Boolean query may be of very high quality and should be used as gold standard. The Jaccard index measure is used once more to represent the similarity between suggested MeSH terms and those present in the original query, and F1 is used to represent the search effectiveness of the associated query (with suggested MeSH terms included). The results of this correlation analysis are reported in Fig 6. We find that, while for all lexical methods search effectiveness is weakly correlated with the overlap of MeSH terms, this is not the case for BERT methods. This indicates that MeSH term suggestion from the original query may not be the best MeSH term suggestion to suggest. In fact, it often is that MeSH terms that are suggested but not included in the original query provide higher search effectiveness than the original MeSH terms themselves.

5 Search Stability

We next analyse the search effectiveness stability of different MeSH term suggestion methods on a topic-by-topic basis. With search effectiveness stability we refer to the amount of variance across topics of the measured search effectiveness obtained when using queries with MeSH terms suggested by a specific MeSH term suggestion method. The larger the effectiveness, the lower the stability. We only analyse the best-performing lexical (U-CUT) and BERT (F-B-FO) methods.

Figure 7, which combines the topics of all of the CLEF TAR datasets test splits, shows that for most of the topics, both kinds of MeSH suggestion methods can outperform or match the effectiveness of the original queries. We also find that our MeSH term suggestion methods sometimes obtain lower effectiveness. It is unclear if these are difficult topics to suggest MeSH terms for, or if there are mistakes in these queries that cause the poor effectiveness (e.g., spelling mistakes in the free text atomic clauses that were not detected at the time of data cleaning).

6 Case Study

Given the findings above, we next seek to investigate the reasons for highly effective or ineffective results. We choose topic CD009642 and topic CD004414 from the CLEF TAR 2019-intervention dataset as they are representative topics where suggestion methods outperform the effectiveness (CD009642) or struggle to match the effectiveness (CD004414) of the original query. Query fragments corresponding to these topics for all the suggestion methods are shown in Tables 6 and 7. We also show their search effectiveness in Table 3.

Firstly, we find that both the suggested MeSH terms and the search effectiveness are similar for all lexical methods. One exception is UMLS in topic CD004414, which suggests more MeSH terms than the other two methods, causing a drop in effectiveness.

On the other hand, MeSH terms suggested by BERT methods appear to differ greatly from lexical methods. BERT methods have captured both lexically similar MeSH terms and terms semantically related to the input free text atomic clauses. One example is shown in the first fragment of topic CD009642. While all lexical methods suggest Lidocaine which is lexically equal to lidocain, BERT methods suggest similar drugs such as Procaine. Another example in topic CD004414 shows that BERT methods can use this semantic matching ability to suggest MeSH terms indicating the method of intervention, shown in suggesting Patch Tests. Therefore, BERT methods suggest MeSH terms that are not bound to the lexical semantics of a free text atomic clause. Another advantage of BERT methods is that they guarantee that at least one MeSH term will be suggested. For lexical methods, suggestions are based on pre-existing rule-based knowledge; thus, when free text atomic clauses can not be matched, no MeSH terms can be suggested (e.g., ATM does not suggest any MeSH term in Fragment 1 of CD009642). Overall, we believe that the semantic matching of BERT may sometimes be detrimental to MeSH suggestion. We leave the investigation into how to prevent BERT from suggesting MeSH terms that are not relevant to the information need of a query fragment (e.g., suggesting a MeSH term that is the intervention of an outcome for a query fragment) for future work.

Conclusion

In this article, we extend our previous line of research on suggesting MeSH terms for Boolean queries for systematic review literature search. This task adds to a recent stream of research that has focused on computational methods for the assisted creation or refinement of Boolean queries for systematic review creation. In addition to the lexical methods we proposed previously, in this new line of work, we introduced a new set of BERT based MeSH suggestion methods. We undertook a comprehensive evaluation and analysis of our new MeSH suggestion methods. We compared the effectiveness of the suggested MeSH terms from our new methods to both our existing lexical methods and the original queries formulated by information specialists. We found that the MeSH terms originally chosen by information specialists were often not the most effective choice and that more effective MeSH terms can be suggested automatically by our new methods.We also found that using BERT methods can generally achieve higher effectiveness than the Lexical method in MeSH Term suggestion: this may be due to the fact that BERT methods were often able to capture deeper semantic relationships. This finding motivates future work to combine lexical and BERT methods in order to reap the benefits of both approaches. Combining such sparse and dense approaches has seen much success in related areas of research, such as ad-hoc search .

In addition, we believe that the full potential of using MeSH entities in our suggestion method is unexplored. In our future work, we project three research directions using more information from MeSH entities to achieve more effective MeSH term suggestions, including (1) Use of MeSH tree hierarchy: MeSH entities are organised in a tree hierarchy. The parent-child relationship of entities may further restrict the number of MeSH terms suggested by MeSH Term suggestion methods (Exp: Use parent MeSH entity to restrict which child entities can appear in the suggestion list). (2) Use of MeSH categories: MeSH entities are categorised according to their natures (term, concept, descriptor and category). The nature of MeSH entities may be used in the fine-tuning process of the MeSH term suggestion methods to represent the MeSH entities. (3) Use of external MeSH definition: Each MeSH entity has a corresponding Wikipedia page to explain its content and uses. These comprehensive pages may be used to further fine-tune our MeSH term suggestion model to achieve effective MeSH term suggestions.

Identifying MeSH terms to add to a Boolean query for a systematic review literature search is a difficult task for information specialists. The findings of this article have implications for both the Information Retrieval and Systematic Review communities. Firstly, our methods can be used in automatic query formulation situations (see, e.g., work by Scells et al. ). Secondly, they can be integrated into existing tools to assist information specialists in formulating more effective queries .

Appendices

Acknowledgments

Shuai Wang is supported by a UQ Earmarked PhD Scholarship and this research is funded by the Australian Research Council Discovery Projects program ARC Discovery Project DP210104043.

References