CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domain
Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow
Introduction
Collecting and understanding key clinical information, such as disorders, symptoms, drugs, etc., from electronic health records (EHRs) has wide-ranging applications within clinical practice and research (Leaman et al., 2015; Wang et al., 2018). A better understanding of this information can, on the one hand, facilitate novel clinical studies, and, on the other hand, help practitioners to optimize clinical workflows. However, free text is ubiquitous in EHRs. This leads to great difficulties in harvesting knowledge from EHRs. Therefore, natural language processing (NLP) systems, especially information extraction components, play a critical role in extracting and encoding information of interest from clinical narratives, as this information can then be fed into downstream applications. For example, the extraction of structured information from clinical narratives can help in decision making or drug repurposing (Marimon et al., 2019).
However, information extraction in non-standard domains like the clinical domain is a challenging problem due to the large number of complex terms and unusual document structures (Lee et al., 2020). In addition, pre-trained language models (PLM) such as BERT (Devlin et al., 2019) that demonstrated superior performance for many NLP tasks are typically trained on standard domains, such as web texts, news articles or Wikipedia. Despite showing some robustness across languages and domains (Conneau et al., 2020) these models still achieve their best performance when applied to targets similar to their pre-training corpora which can limit their applicability in many situations (Gururangan et al., 2020). One way to overcome this domain-gap is the adaptation of existing language models to the new target domain or training a new domain-specific model from scratch (Beltagy et al., 2019; Lee et al., 2020). Several recent works have shown that this kind of adaptation boosts performance for downstream tasks in non-standard domains by, e.g., pre-training with masked language modeling (MLM) objectives on documents from the target domain (Weber et al., 2019; Naseem et al., 2021).
While all the previous methods help to build high-performing model architectures, often there is also a lack of annotated data in the clinical domain which is usually needed for all deep-learning-based models. On the one hand, this domain has high requirements regarding the removal or masking of protected health information (PHI) of individuals (Uzuner et al., 2007; Stubbs et al., 2015) which is particularly worthy of protection and can prevent data publication. On the other hand, information extraction tasks are often specific to their target domain and clinical concepts are only found very infrequently outside EHRs which limits reusability of existing resources. Possible solutions for the low-resource problem can be multi-task learning (Khan et al., 2020; Mulyar et al., 2021) or transfer Learning (Lee et al., 2018; Peng et al., 2019) across similar corpora from the clinical domain. However, transferring knowledge is particularly challenging in the clinical domain as biomedical NLP models have problems generalizing to new entities (Kim and Kang, 2021). Therefore, one has to carefully select the transfer sources (Lange et al., 2021b).
Over the last years, we have participated in a series of shared tasks on information extraction in the Spanish clinical domain (Marimon et al., 2019; Miranda-Escalada et al., 2020; Lima-López et al., 2021). With our systems, we were able to outperform the other participants and won the competitions twice. The winning systems were task-agnostic and utilized domain-adapted language models and word embeddings (Lange et al., 2019), as well as improved training routines for transformer models (Lange et al., 2021a). Based on our findings and lessons learned during the competitions, we propose in this paper a robust model architecture and training procedure for concept extraction in the clinical domain that is task- and language-agnostic. We introduce a new Spanish clinical language model CLIN-XES (Clinical XLM-R) that outperforms existing transformer models on Spanish corpora and exemplifies the benefits of cross-language domain adaptation for English tasks as well. For this, we perform a broad evaluation of ten clinical information extraction tasks from two languages (English and Spanish), including low-resource settings. Finally, we perform cross-task transfer experiments and show that this can boost performance by more than 47 points for few-shot training. Our results demonstrate great and consistent improvements compared to standard transformer models across all tasks in both languages. We release both, CLIN-XES as well as its English counterpart CLIN-XEN.
Approach
In this paper, we introduce new pre-trained language models and propose a robust model architecture to perform concept extraction in the clinical domain for English and Spanish. The overall model architecture is shown in Figure 1 and our proposed model components are highlighted. First, the input is computed on subword-level instead of the usual word-level, which eliminates the need for external tokenization. In addition, the input is enriched with its cross-sentence context to capture a wider document context. Second, the input is processed by a transformer model that is adapted to our target domain. Third, the model output is computed using a conditional random field (CRF) output layer to address long annotations. Then, an ensemble over models trained on different training splits is computed that reduces variance and captures the complementary knowledge from all models. Finally, we experiment with cross-task model transfer to further improve the model in few-shot settings.
In summary, the contributions of this paper are as follows:
We study the impact of domain-adaptive pre-training for clinical concept extraction for different embedding types and publish new language models that are adapted to the clinical domain. We show that this PLM outperforms other publicly available embeddings and models in our settings and we also show that cross-language domain adaptations works for English tasks as well.
We perform a broad evaluation of ten clinical sequence labeling tasks across two languages, including low-resource and transfer settings. By this, we demonstrate how our methods can further boost already high-performing transformer models by using advanced training methods and effective changes in the architecture.
Our models outperform the state-of-the-art methods for clinical and biomedical concept extraction, as well as various other transformer models for all ten tasks.
We make our new domain-adapted CLIN-X language models and the source code for fine-tuning the concept extraction models using our methods publicly available.
Materials and Methods
In this section, we start with a brief description of the input representations. Then, we discuss our proposed architectural choices as well as the advanced training methods.
State-of-the-art methods for concept extraction typically rely on word embeddings or language models as input representations. The standard approach is the pre-training of these models on large-scale unannotated datasets once and their reuse as powerful representations for many downstream applications (Collobert et al., 2011). Phan et al. (2019) have shown that contextual information helps in particular in the medical domain, e.g., due to the high number of synonyms. Thus, we focus on the usage of contextualized embeddings in this work, which are most often retrieved from transformer language models nowadays. This is either done with auto-regressive language modeling (Peters et al., 2018) or masked language modeling (Devlin et al., 2019), which we use in this paper.
A popular way to approach the challenges of NLP in non-standard domains is the inclusion of domain knowledge via domain-specific embeddings (Friedrich et al., 2020). For this, word embeddings or language models are pre-trained or further specialized on documents of the target domain. These embeddings can be used in downstream applications. This kind of domain adaptation has shown great benefits in practice (Gururangan et al., 2020), thus, we explore domain- and language-adaptive pre-training of transformer models in this paper.
The CLIN-X pre-trained language model.
At the time of writing, there is no Spanish clinical transformer publicly available. Thus, we train and publish the CLIN-XES language model. The model is based on the multilingual XLM-R transformer, which was trained on 100 languages and showed superior performance in many different tasks across languages and can even outperform monolingual models in certain settings (Conneau et al., 2020). Even though XLM-R was pre-trained on 53GB of Spanish documents, this was only 2% of the overall training data. To steer this model towards the Spanish clinical domain, we sample documents from the Scielo archive and the MeSpEn resources (Villegas et al., 2018). The resulting corpus has a size of 790MB and is highly specific for our target setting. We initialize CLIN-X using the pre-trained XLM-R weights and train masked language modeling (MLM) on the clinical corpus for 3 epochs which roughly corresponds to 32k steps. Nonetheless, this model is still multilingual and we demonstrate the positive impact of cross-language domain adaptation by applying this model to English tasks. In addition to the Spanish CLIN-XES model, we release an English version CLIN-XEN trained on clinical Pubmed abstracts (850MB) filtered following Haynes et al. (2005) for a direct comparison of our methods in a monolingual setting. This allows researchers and practitioners to address the English clinical domain with an out-of-the-box tailored model so that our transfer methods do not have to be applied. Pubmed is used with the courtesy of the U.S. National Library of Medicine.
2 Concept Extraction Model
In the following, we describe the architectural choices we made compared to the standard transformer model for sequence labeling as proposed by Devlin et al. (2019).
Information extraction tasks are typically performed on the token level, while most transformers work on finer subwords instead. Thus, the input representations from transformers for tokens are either retrieved from the first subword or the average (Devlin et al., 2019). In contrasts, we perform concept extraction directly on the subword level. By doing this, there is no need for external tokenization besides the subword segmentation of the transformer. Note that the usage of domain-specific subwords is still often beneficial compared to the general domain segmentation (Beltagy et al., 2019; Lee et al., 2020).
Cross-sentence context.
Transformers are suited to incorporate information from a larger context. Luoma and Pyysalo (2020) showed that context information from neighboring sentences has positive effects for named entity recognition on the general domain. Finkel et al. (2004) also showed the positive impact of context for clinical concept extraction. We follow these approaches and add context information to the input similar to Schweter and Akbik (2020). We incorporate the context of 100 subwords to the left and right and use the document boundaries to set the context limits as all corpora are clearly separated in documents.
Conditional Random Field Output.
As Kim and Kang (2021) have shown, entity recognition models in the biomedical domain tend to memorize training instances and their labels. This can result in incorrect label encodings as the model fails to generalize. A conditional random field (CRF, Lafferty et al., 2001) can constrain these incorrect sequences as the Viterbi algorithm is used for decoding. In addition, the CRF has advantages when it comes to long entities covering multiple tokens (Lima-López et al., 2021) that appear frequently in the clinical domain.
3 Training on Data Splits.
Having a robust model architecture is a good starting point for NLP in the clinical domain. However, even more important might be the actual training procedure of the model. Thus, we discuss standard and random splits, as well as ensemble models over these splits in the following.
Typically, each dataset is divided into training, development and test splits. The training split is used each epoch to train the model parameters and the best training epoch is selected based on the evaluation score on the development set. Finally, the held-out test set is used by the selected model to compute the final score. These data splits are helpful to compare performances of different models on standardized data, however, using the standard training split without modifications may not result in optimal performance (Gorman and Bedrick, 2019).
Further random splits.
The training and development parts can be further randomly divided into separate parts. Then, parts can be used for training and one part as the validation set for early stopping similar to cross-fold validation. An ensemble based on models trained on the different data splits should be more powerful than the single models as each of them encodes complementary knowledge which helps to reduce variance and biases (Clark et al., 2019). In our experiments, we use so that we get 5 different settings with unique training sets and we train one model for each setting. Note that we do not change or use the test set at all to ensure comparability to previous results.
Training on all available instances.
Recent works sometimes finds that there is no need for a held-out development set and that these labeled instances might be better used during the training. For example, Luoma and Pyysalo (2020) have shown that training on the combined training and development sets boosts performance for named entity recognition remarkably. By this, the model has access to the most data during training and model selection is based on the training loss. However, the training loss is not as meaningful as a stopping criterion and its hard to pick the best model checkpoint. We will compare to this method as an alternative to our split-based experiments.
4 Transfer Learning
Many NLP tasks suffer from a lack of labeled data. This includes non-standard domains like the clinical domain in particular. One solution to improve performance in these domains is the usage of resources from a related task in a transfer process. For example, Hofer et al. (2018) have shown that few-shot NER in the biomedical domain can be improved by transferring trained weights from a similar task. We perform a similar kind of model transfer by transferring the transformer to the new target.
However, not all transfer sources are actually useful as many can lead to negative transfer (Lange et al., 2021b). Thus, we first have to predict a suitable transfer source. We follow Lange et al. (2021b) and compute similarities between our datasets using their proposed model similarity measure. This has been shown to work well across different tasks and domains. The similarity between two models is computed based on the neural feature representations for the target datasets between two task-specific trained models. In our experiments, we study the effect of transfer from different sources in comparison to standard single-task training. Further, we will investigate this kind of transfer in low-resource settings, when the target task has only limited training resources.
In addition to the other methods, ensembling can be used to combine multiple model predictions into one. This ensemble is usually better than a single model – in particular if the models or their training data differ to some degree. We either create ensembles by majority voting (Clark et al., 2019) of training runs that vary by their random seed (standard splits) or their training data (random splits).
Results
This section describes the experimental setup starting with tasks, datasets and implementation details, and discusses the results for our experiments.
Many datasets for natural language processing in specialized domains are published in the context of shared tasks – competitions to evaluate different systems and approaches. Besides English, the clinical domain is well addressed for Spanish, and there exists an active community of researchers for natural language processing of Spanish clinical texts. Thus, in the context of the IberLEF workshop series (Iberian Language Evaluation Forum), several shared tasks have been proposed by the Barcelona Supercomputing Center concerning concept extraction in the clinical domain (Marimon et al., 2019; Gonzalez-Agirre et al., 2019; Miranda-Escalada et al., 2020; Lima-López et al., 2021). In addition to datasets of these shared tasks for Spanish, we consider four English datasets published during a series of shared tasks of the i2b2 project (Uzuner et al., 2007, 2011; Sun et al., 2013; Stubbs et al., 2015). Information on the dataset sizes are given in Table 1 and 2 for Spanish and English, respectively. Note that the Meddoprof and i2b2 2012 corpora consist of two different extraction tasks each. Thus, we consider both tracks as separated tasks in this work resulting in a total of ten tasks. Following the evaluations in the shared tasks, we use the strict micro for all datasets as evaluation metric.
2 Experimental Setup and Implementation Details
We use eight NVIDIA V100 (32GB) GPUs for pre-training the CLIN-X models. The training takes less than 1 day with a batch size of 4 per device and a sequence length of up to 512 subwords. The models were trained with the huggingface trainer for MLM.
Sequence Labeling.
The sequence labeling models were trained on single NVIDIA V100 GPUs up to 20 hours depending on the dataset size. The models were trained using the flair framework with the AdamW optimizer with an initial learning rate of and a batch size of 16 for 20 epochs. The model selection was performed on the development score if trained on standard or random splits or the training loss otherwise.
Transfer and Low-Resource Experiments.
The median model according to the development score on the source dataset was taken for transfer and used for the initialization of the target model. Except for the initialization, the training was identical to the single task training. The low-resource settings were created by limiting the data splits to the first sentences without shuffling. The test set is not changed and remains identical.
3 Evaluation of Embeddings
The choice of input embeddings has a large impact on downstream performance and may be the most important factor. Table 3 shows the average performance of several different embeddings and transformer models for the two languages. As expected, the monolingual transformers (BERT, BETO) excel at their target language, but cannot compete with multilingual models (mBERT, XLM-R) when applied to an unseen language. The lower part of Table 3 lists domain-specific variants of the embeddings which are generally more powerful in our domain-specific setting. We see that our CLIN-X models perform best for their respective languages. Furthermore, the CLIN-XES performs almost as well as the CLIN-XEN model on the English datasets, for which it was not explicitly trained. This shows, that the domain adaptation of multilingual models can also help for texts from other languages of the same domain. Due to CLIN-XES stable performance across all tasks and languages, we will use this model for the following ablations and transfer experiments.
4 Evaluation of Training Methods
The foundation for all following concept extraction models is the CLIN-XES transformer, as it has shown robust results across all tasks. For comparison to fixed standard splits, we train the models on different random splits. We see in Table 4 that in particular ensembles over random splits are a lot better than the standard splits and also all training instances. While the median performance is roughly similar for all methods, the random splits offer a lot more variety in training instances and allow for better maximum performance models. Thus, the ensemble based on random splits achieves also much higher numbers.
5 Evaluation of Concept Extraction Models
The lower part of Table 4 lists an ablation study of our individual model components. For example, adding cross-sentence context to the transformers boosts performance across all tasks by 0.5 F1 on average. Performing concept extraction on the subword level helps even further. This is particularly helpful considering that no external tokenization is needed, which can be challenging in the clinical domain (Lange et al., 2020). The CRF helps for both languages, though the differences are larger for Spanish, as the two MEDDOPROF tasks have particularly long annotations (2.53 tokens per annotation on average). The same holds for the BIOSE labels, that have the smallest impact of all components, but consistently improve upon the standard BIO labels. As each of our proposed methods improves the transformer even further, we use the combination of all methods in the following as our model architecture.
6 Evaluation of Transfer Learning
In addition to the training based on random splits, we explore the effects of transfer learning. For this, we simulate low-resource settings where we limit the annotated data of the target dataset between 250 labeled sentences up to 7500 sentences, roughly the size of the smallest corpus. The results are given in Table 5 and Table 6 for English and Spanish, respectively.
Large positive transfer happens in most settings, particularly for the low-resource settings with up to (+47.3 points) for Meddoprof when only 250 labeled sentences are available. The improvements in the full data scenario are below 1 F1. However, there is also negative transfer, in particular using i2b2 2012-T and Cantemist datasets as transfer sources often result in negative transfer. The source selection is also crucial in low-resource scenarios, as not every source is equally beneficial. Using the model similarity measure from Lange et al. (2021b) we are able to predict good transfer sources in all settings; often the best source is selected.
7 Comparison to State-of-the-Art Models
As our results demonstrate, we have proposed a robust model for the clinical domain that works well across the different tasks in both languages. Finally, we compare CLIN-X to various transformer models as introduced earlier. We also compare to HunFlair (Weber et al., 2021), the current state-of-the-art for concept extraction in the biomedical domain. We use their model architecture based on clinical flair and fasttext embeddings and train models accordingly on our datasets. In addition, we compare to our NLNDE submissions for the Spanish shared tasks and the ClinicalBERT by Alsentzer et al. (2019) for the English datasets.
The results for each task are shown in Table 7. The CLIN-X language models in combination with our model architecture outperform the other transformers and HunFlair by a large margin. CLIN-X is able to utilize the domain knowledge obtained from the additional pre-training with further improvements from the ensembling over random splits. Even though CLIN-X works best in combination with our model architecture, CLIN-X based on the standard transformer architecture with a single classification layer already outperforms the existing models on 8 out of 10 tasks.
We tested statistical significance between CLIN-XES with and without transfer learning – highlighted with asterisks in Table 7. We find that all differences for English are significant, while only one difference for Spanish is significant. This might indicate the complementary relationship of domain adaptation and model transfer learning. As CLIN-X was explicitly adapted to Spanish, additional transfer is not necessary in high-resource settings. In contrast, the cross-language domain adaptation for English can still be improved with transfer from related sources, where CLIN-XES +Transfer has also notably higher performances in 3 out of 5 settings compared to CLIN-XEN which is adapted to English.
Conclusion
In this paper, we described the newly pre-trained CLIN-X language models for the clinical domain. We have shown that CLIN-X sets the new state the of the art results for ten clinical concept extraction tasks in two languages. We demonstrated the positive impact of other model components, such as ensembles over random splits and cross-sentence context and we have studied the effects of cross-task transfer learning from different clinical corpora. Using a model similarity measure, we found good transfer sources for almost all datasets in general and for low-resource scenarios in particular. We are convinced that the new CLIN-X language models will help boosting performance for various Spanish and English clinical information extraction tasks with our or other model architectures.