Clinical Concept Extraction with Contextual Word Embedding
Henghui Zhu, Ioannis Ch. Paschalidis, Amir Tahmasebi
Introduction
Automatic clinical concept extraction is a crucial step in transforming the abandon of unstructured patient clinical data into a set of actionable information. A major challenge towards developing clinical Named Entity Recognition (NER) tools is access to a corpus of labeled data. The 2010i2b2/VA challenge released such corpus of annotated clinical notes to facilitate the development of clinical concept extraction systems. Since the release of the 2010 i2b2/VA dataset, there have been numerous efforts on developing NER tools reported in the literature. In earlier works, Conditional Random Fields (CRF) with token-level features were proposed for concept extraction . Nevertheless, such approaches require manual feature engineering, which makes its application limited. Recently, with the advancement of deep learning models and its usage for natural language processing, Recurrent Neural Network (RNN) models, such as Long Short-Term Memory (LSTM), have been widely used for deriving contextual features for CRF model training and have demonstrated promising performance for clinical concept extraction tasks . Although RNN models provide a good set of features for NER, they depend heavily on the word embedding models , which are not good at dealing with the complex characteristics of word use under different linguistic contexts.
Recently, Peters et al. proposed a contextual word-embedding model (ELMo), which claims to improve the performance of various NLP tasks such as sentiment analysis, question answering and sequence labeling . By including the features of a language model, the word representations of the ELMo contain richer information compared to a standard word embedding such as skip-gram or Glove . Although the ELMo model is shown to have a good performance in some NER tasks such as the CoNLL 2003 NER task, it is trained on a corpus in a general domain and as a result, does not demonstrate a desired performance for a clinical concept extraction task. This could be due to the fact that clinical text is structured differently compared to text in a general domain and therefore, the language model trained on a general domain corpus fails to generalize well on it.
In this work, we train an ELMo model in a corpus with a mixture of clinical reports and relevant Wikipedia pages in the clinical domain. Next, a bidirectional LSTM-CRF model is applied to identify the clinical concepts. This model is trained and tested on the 2010 i2b2/VA challenge . Our proposed model achieves the best performance among others the state-of-the-art models by 3.4% in terms of F1-score.
Dataset
In this work, we used the data provided by the 2010 i2b2/VA challenge for training a clinical concept extraction system. Due to the restrictions introduced by Institutional Review Board (IRB), only a portion of data from the original dataset is available. The released dataset consists of clinical summaries from three different medical sites: Partners Healthcare, Beth Israel Deaconess Medical Center, and the University of Pittsburgh Medical Center. There are three clinical concepts annotated in this corpus: problems, tests, and treatments. There are 170 summaries for training and 256 for test. The dataset statistics are shown in Table 1.
Methods
The proposed model consists of an ELMo model trained using a corpus in a clinical domain and a bidirectional LSTM-CRF model for the clinical context extraction.
ELMo is a recently proposed model that extends a word embedding model with features produced by a bidirectional language model. It has been shown that the utilization of ELMo for different NLP tasks results in improved performance compared to other types of word embedding models such as skip-gram and Glove. Although released several pretrained ELMo models, all of them are trained on a general text corpus . We believe the clinical reports are structured differently compared to such general language corpora. Therefore, it is necessary to train an ELMo model specifically for the clinical domain in order to achieve the desired performance for clinical NLP tasks including concept extraction. To do so, the following corpora were considered for training the ELMo model.
Wikipedia pages with titles that are items (medical concepts) in a standard clinical ontology, known as SNOMED CT . The following sections were excluded: ‘See also’, ‘References’, ‘Further reading’ and ‘External links’. Furthermore, if a term has more than one Wikipedia page, we exclude all these pages in the corpus for avoiding introducing ambiguity.
The discharge summaries and radiology reports from MIMIC-III dataset . Phrase normalization is performed on the deidentified entities in the dataset, e.g., converting all names to ‘John Does’, removing the prefixes and suffixes for deidentified data and numbers. The reason for the normalization is to transform the deidentified reports to be more similar to actual reports.
This paper uses our in-house built sentence segmentation toolhttps://github.com/noc-lab/simple_sentence_segment to detect sentence boundary. NLTK is used for tokenization. Statistics of the training corpus is shown in Table 2. For training the ELMo model, we use the default hyperparameter settings shown in . In detail, a character-based Convolutional Neural Network (char-CNN) embedding layer is used with character embeddings dimension , filter widths and number of filters $$. Next, a two-layer bidirectional LSTM (bi-LSTM) with 4,096 hidden units in each layer is considered. After each char-CNN embedding and LSTM layer, the output is projected to 512 dimensions and a high-way connection is applied. In addition, we define the vocabulary in the language model as the tokens that appear not less than 5 times in the corpus.
We randomly split the whole corpus into a training corpus (90%) and a testing corpus (10%). We train an ELMo model using the training corpus for epochs. The average perplexity in the testing corpus is . On the other hand, the ELMo modelThe ‘Original’ model in section ‘Pre-trained ELMo Models’ at https://allennlp.org/elmo trained in a corpus of a general domain only achieves a perplexity of . Despite the fact that comparing these two perplexities is not fair due to the different vocabularies in the two language models, such large gap between them suggests that training a specific ELMo model for the clinical domain is necessary.
2 Bidirectional LSTM CRF model for NER
We use a bidirectional LSTM-CRF model for the NER task. The architecture of the proposed model is shown in Figure 1. The input is a list of tokens. Contextual word embeddings are generated as a learnable aggregation of char-CNN word-embedding layer and two bi-LSTM layers for the language model. Suppose is the context-independent token representation for the th token in the sentence produced by the character CNN layer and denote as the token layer. Also denote and the two bidirectional LSTM layers in ELMo with forward and backward language models. Following , we use
as the features for the NER model, where is a scale factor and ’s are softmax-normalized weights. During training the NER model, the parameters of the ELMo model is fixed while and ’s are learnable parameters. Next, a two layer bidirectional LSTM is applied with ’s as input. Finally, a linear-chain CRF layer is applied for predicting the label of each token. In this paper, the BIO-tagging format is used, as shown in Figure 1.
Results
A two-layer bidirectional LSTM is used for NER task, each of which consists of 256 hidden states. For regularization, dropout is applied to the LSTM layer with a rate of . We train the model with the Adam optimizer using a learning rate of , a batch size of 32, and 200 epochs.
Three different scenarios of the proposed ELMo-based model were considered and compared in this study: Two BiLSTM-CRF models were trained using 1) an ELMo model trained on a general domain corpus , referred to as “ELMo(General) + BiLSTM-CRF (Single)"; and 2) an ELMo model trained on a clinical corpus as described before, referred to as “ELMo(Clinical) + BiLSTM-CRF (Single)". The training was performed 10 times starting with 10 different random seeds and the mean and standard deviation of the performance metrics were reported. Besides, we also trained an ensemble model based on the most voted label by the 10 models for each token. The performance of our models and several previously published baseline models for the 2010 i2b2/VA challenge is reported in Table 3 in terms of precision, recall and F1-score for exact class spans using the definition given in . As expected, it can be observed from the table that an ELMo model trained using domain-specific data results in significant improvement in performance. The BiLSTM-CRF model with ELMo trained on the clinical corpus outperforms other alternatives. Furthermore, it was observed that our best model yielded similar performance among three types of named entities (problem, treatment and test).
Conclusions
Contextual word embedding approach such as ELMo has exhibited a promising performance in many natural language processing tasks. In this paper, we trained a domain-specific ELMo model using a clinical domain-specific corpus and furthermore, utilized it for building a clinical concept extraction tool. We trained and tested the proposed model using the dataset provided by the 2010 i2b2/VA challenge. To the best of our knowledge, our model yields the best performance compared to the reported prior work by a significant margin of 3.4% in terms of F1-score.
From the results reported in this work, one can conclude that training a domain-specific language model is essential to achieve high performance for NER tasks. Nevertheless, access to a large size domain-specific corpus is a known challenge for training a language model. In this work, we demonstrated an effective yet simple approach to create such corpus for a clinical domain by filtering a general domain corpus such as wiki-pages using a domain-specific ontology such as SNOMED CT.
Research partially supported by the ONR under MURI N00014-16-1-2832, by the NSF under grants DMS-1664644, CNS-1645681, CCF-1527292, and IIS-1237022, by the NVIDIA Corporation with the donation of a Titan Xp GPU, and by the Center for Information and Systems Engineering.