Cross-type Biomedical Named Entity Recognition with Deep Multi-Task Learning

Xuan Wang, Yu Zhang, Xiang Ren, Yuhao Zhang, Marinka Zitnik, Jingbo Shang, Curtis Langlotz, Jiawei Han

Introduction

Biomedical named entity recognition (BioNER) is one of the most fundamental task in biomedical text mining that aims to automatically recognize and classify biomedical entities (e.g., genes, proteins, chemicals and diseases) from text. BioNER can be used to identify new gene names from text (Smith et al. 2008). It also serves as a primitive step of many downstream applications, such as relation extraction (Cokol et al. 2005) and knowledge base completion (Szklarczyk et al. 2017; Wei et al. 2013; Xie et al. 2013; Szklarczyk et al. 2015).

BioNER is typically formulated as a sequence labeling problem whose goal is to assign a label to each word in a sentence. State-of-the-art BioNER systems often require handcrafted features (e.g., capitalization, prefix and suffix) to be specifically designed for each entity type (Ando 2007; Leaman and Lu 2016; Zhou and Su 2004; Lu et al. 2015). This feature generation process takes the majority of time and cost in developing a BioNER system (Leser and Hakenberg 2005), and leads to highly specialized systems that cannot be directly used to recognize new types of entities. The accuracy of the resulting BioNER tools remains a limiting factor in the performance of biomedical text mining pipelines (Huang and Lu 2015).

Recent NER studies consider neural network models to automatically generate quality features (Chiu and Nichols 2016; Ma and Hovy 2016; Lample et al. 2016; Liu et al. 2018). Crichton et al. 2017 took each word token and its surrounding context words as input into a convolutional neural network (CNN). Habibi et al. 2017 adopted the model from Lample et al. 2016 and used word embeddings as input into a bidirectional long short-term memory-conditional random field (BiLSTM-CRF) model. These neural network models free experts from manual feature engineering. However, these models have millions of parameters and require very large datasets to reliably estimate the parameters. This poses a major challenge for biomedicine, where datasets of this scale are expensive and slow to create and thus neural network models cannot realize their potential performance to the fullest (Camacho et al. 2018). Although neural network models can outperform traditional sequence labeling models (e.g., CRF models (Lafferty et al. 2001)), they are still outperfomed by handcrafted feature-based systems in multiple domains (Crichton et al. 2017).

One direction to address the above challenge is to use the labeled data of different entity types to augment the training signals for each of them, as information like word semantics and grammatical structure may be shared across different datasets. However, simply combining all datasets and training one single model over multiple entity types can introduce many false negatives because each dataset is typically specifically annotated for one or only a few entity types. For example, combining dataset A for gene recognition and dataset B for chemical recognition will result in missing chemical entity labels in dataset A and missing gene entity labels in dataset B. Multi-task learning (MTL) (Collobert and Weston 2008; Søgaard and Goldberg 2016) offers a solution to this issue by collectively training a model on several related tasks, so that each task benefits model learning in other tasks without introducing additional errors. MTL has been successfully applied in natural language processing (Collobert and Weston 2008), speech recognition (Deng et al. 2013), computer vision (Girshick 2015) and drug discovery (Ramsundar et al. 2015). But MTL is less commonly used and has seen limited success in BioNER so far. Crichton et al. 2017 explored MTL with a CNN model for BioNER. However, Crichton et al. 2017 only considers word-level features as input, ignoring character-level lexical information which are often crucial for modeling biomedical entities (e.g. -ase could be an important subword feature for gene/protein entity recognition). As a result, their best performing multi-task CNN model does not outperform state-of-the-art systems that use on handcrafted features (Crichton et al. 2017).

In this paper, we propose a new multi-task learning framework using char-level neural models for BioNER. The proposed framework, despite being simple and not requiring any feature engineering, achieves excellent benchmark performance. Our multi-task model is built upon a single-task neural network model (Liu et al. 2018). In particular, we consider a BiLSTM-CRF model with an additional context-dependent BiLSTM layer for modeling character sequences. A prominent advantage of our multi-task model is that inputs from different datasets can efficiently share both character- and word-level representations, by reusing parameters in the corresponding BiLSTM units. We compare the proposed multi-task model with state-of-the-art BioNER systems and baseline neural network models on 15 benchmark BioNER datasets and observe substantially better performance. We further show through detailed experimental analysis on 5 datasets that the proposed approach adds marginal computational overhead and outperforms strong baseline neural models that do not consider multi-task learning, suggesting that multi-task learning plays an important role in its success. Altogether, this work introduces a new text-mining approach that can help scientists exploit knowledge buried in biomedical literature in a systematic and unbiased way.

Background

Let Φ\Phi denote the set of labels indicating whether a word is part of a specific entity type or not. Given a sequence of words w={w1,w2,...,wn}\boldsymbol{w}=\{w_{1},w_{2},...,w_{n}\}, the output is a sequence of labels y={y1,y2,...,yn}\boldsymbol{y}=\{y_{1},y_{2},...,y_{n}\}, yi∈Φy_{i}\in\Phi. For example, given a sentence ‘‘…. including the RING1 …", the output should be ‘‘… O O S-GENE …" in which ‘‘O" indicates a non-entity type and ‘‘S-GENE" indicates a single-token GENE type.

2 Long Short-Term Memory (LSTM)

Long short-term memory neural network is a specific type of recurrent neural network that models dependencies between elements in a sequence through recurrent connections (Fig. 1). The input to an LSTM network is a sequence of vectors X={x1,x2,...,xT}\boldsymbol{X}=\{\boldsymbol{x}_{1},\boldsymbol{x}_{2},...,\boldsymbol{x}_{T}\}, where vector xi\boldsymbol{x}_{i} is a representation vector of a word in the input sentence. The output is a sequence of vectors H={h1,h2,...,hT}\boldsymbol{H}=\{\boldsymbol{h}_{1},\boldsymbol{h}_{2},...,\boldsymbol{h}_{T}\}, where hi\boldsymbol{h}_{i} is a hidden state vector. At step tt of the recurrent calculation, the network takes xt,ct−1,ht−1\boldsymbol{x}_{t},\boldsymbol{c}_{t-1},\boldsymbol{h}_{t-1} as inputs and produces ct,ht\boldsymbol{c}_{t},\boldsymbol{h}_{t} via the following intermediate calculations:

where σ(⋅)\sigma(\cdot) and tanh(⋅)tanh(\cdot) denote element-wise sigmoid and hyperbolic tangent functions, respectively, and ⊙\odot denotes element-wise multiplication. The it\boldsymbol{i}_{t}, ft\boldsymbol{f}_{t} and ot\boldsymbol{o}_{t} are referred to as input, forget, and output gates, respectively. The gt\boldsymbol{g}_{t} and ct\boldsymbol{c}_{t} are intermediate calculation steps. At t=1t=1, h0\boldsymbol{h}_{0} and c0\boldsymbol{c}_{0} are initialized to zero vectors. The trainable parameters are Wj\boldsymbol{W}^{j},Uj\boldsymbol{U}^{j} and bj\boldsymbol{b}^{j} for j∈{i,f,o,g}j\in\{i,f,o,g\}.

The LSTM architecture described above can only process the input in one direction. The bi-directional long short-term memory (BiLSTM) model improves the LSTM by feeding the input to the LSTM network twice, once in the original direction and once in the reversed direction. Outputs from both directions are concatenated to represent the final output. This design allows for detection of dependencies from both previous and subsequent words in a sequence.

3 Bi-directional Long Short-Term Memory-Conditional Random Field (BiLSTM-CRF)

A naive way of applying the BiLSTM network to sequence labeling is to use the output hidden state vectors to make independent tagging decisions. However, in many sequence labeling tasks such as BioNER, it is useful to also model the dependencies across output tags. The BiLSTM-CRF network adds a conditional random field (CRF) layer on top of a BiLSTM network. This BiLSTM-CRF network takes the input sequence X={x1,x2,...,xn}\boldsymbol{X}=\{\boldsymbol{x}_{1},\boldsymbol{x}_{2},...,\boldsymbol{x}_{n}\} to predict an output label sequence y={y1,y2,...,yn}\boldsymbol{y}=\{y_{1},y_{2},...,y_{n}\}. A score is defined as:

where P\boldsymbol{P} is an n×kn\times k matrix of the output from the BiLSTM layer, nn is the sequence length, kk is the number of distinct labels, A\boldsymbol{A} is a (k+2)×(k+2)(k+2)\times(k+2) transition matrix and Ai,j\boldsymbol{A}_{i,j} represents the transition probability from the ii-th label to the jj-th label. Note that two additional labels and are used to represent the start and end of a sentence, respectively. We further define YX\boldsymbol{Y}_{\boldsymbol{X}} as all possible sequence labels given the input sequence X\boldsymbol{X}. The training process maximizes the log-probability of the label sequence y\boldsymbol{y} given the input sequence X\boldsymbol{X}:

A three-layer BiLSTM-CRF architecture is employed by Lample et al. 2016 and Habibi et al. 2017 to jointly model the word and the character sequences in the input sentence. In this architecture, the first BiLSTM layer takes character embedding sequence of each word as input, and produces a character-level representation vector for this word as output. This character-level vector is then concatenated with a word embedding vector, and fed into a second BiLSTM layer. Lastly, a CRF layer takes the output vectors from the second BiLSTM layer, and outputs the best tag sequence by maximizing the log-probability in Equation 1.

In practice, the character embedding vectors are randomly initialized and co-trained during the model training process. The word embedding vectors are retrieved directly from a pre-trained word embedding lookup table. The classical Viterbi algorithm is used to infer the final labels for the CRF model. The three-layer BiLSTM-CRF model is a differentiable neural network architecture that can be trained by backpropagation.

Deep multi-task learning for BioNER

The vanilla BiLSTM-CRF model can learn high-quality representations for words that appeared in the training dataset. However, it often fails to generalize to out-of-vocabulary (OOV) words (i.e., words that did not appear in the training dataset) because they don’t have a pre-trained word embedding. These OOV words are common in biomedical text (67.21% OOV words of the datasets in Table 1). Therefore, for the baseline single-task BioNER model, we use a neural network architecture that better handles OOV words. As shown in Fig. 2, our single-task model consists of three layers. In the first layer, a BiLSTM network is used to model the character sequence of the input sentence. We use character embedding vectors as input to the network. Hidden state vectors at the word boundaries of this character-level BiLSTM are then selected and concatenated with word embedding vectors to form word representations. Next, these word representation vectors are fed into a word-level BiLSTM layer (i.e., the upper BiLSTM layer in Fig. 2). Lastly, output of this word-level BiLSTM is fed into the a CRF layer for label prediction. Compared to the vanilla BiLSTM-CRF model, a major advantage of this model is that it can infer the meaning of an out-of-vocabulary word from its character sequence and other characters around it. For example, the model is now able to infer that ‘‘RING2’’ likely represents a gene symbol, even though then network may have only seen the word ‘‘RING1’’ during training.

2 Multi-task models (MTMs)

An important characteristic of the BioNER task is the limited availability of supervised training data. We propose a multi-task learning approach to address this problem by training different BioNER models on datasets with different entity types while sharing parameters across these models. We hypothesize that the proposed approach can make more efficient use of the data and encourage the models to learn representations of words and characters (which are shared between multiple corpora) in a more effective and generalized way.

We give a formal definition of the multi-task setting as the following. Given mm datasets, for i∈{1,...,m}i\in\{1,...,m\}, each dataset DiD_{i} consists of nin_{i} training samples, i.e., Di={wji,yji}j=1niD_{i}=\{\boldsymbol{w}_{j}^{i},y_{j}^{i}\}_{j=1}^{n_{i}}. We denote the training matrix for each dataset as Xi={x1i,...,xnii}\boldsymbol{X}^{i}=\{\boldsymbol{x}_{1}^{i},...,\boldsymbol{x}_{n_{i}}^{i}\} (Xi\boldsymbol{X}^{i} is the feature representation of the input word sequence wji\boldsymbol{w}_{j}^{i}) and the labels for each dataset as yi={y1i,...,ynii}\boldsymbol{y}^{i}=\{y_{1}^{i},...,y_{n_{i}}^{i}\}. The model parameters include the word-level BiLSTM parameters (θiw\theta_{i}^{w}), the character-level BiLSTM parameters (θic\theta_{i}^{c}) and the output CRF parameters (θio\theta_{i}^{o}). A multi-task model therefore consists of mm different models, each trained on a separate dataset, while sharing part of the model parameters across datasets. The loss function LL of the multi-task model is:

The log-likelihood term is shown in Equation 1 and λi\lambda_{i} is a positive hyper-parameter that controls the contribution of each dataset. We observed that our multi-task model is able to achieve very competitive performance with λi=1\lambda_{i}=1 on all datasets that we evaluated on and therefore use this value in our experiments. However, we believe that the performance can be improved with further tuned λi\lambda_{i} values.

We propose three different multi-task models, as illustrated in Fig. 3. These three models differ in which part of the model parameters (θiw,θic,θio\theta_{i}^{w},\theta_{i}^{c},\theta_{i}^{o}) are shared across multiple datasets:

In this model, θic=θc\theta^{c}_{i}=\theta^{c} are shared among different tasks. All datasets are iteratively used to train the model. When a dataset is used, the parameters updated during the training are θc\theta^{c} and θiw\theta^{w}_{i}. The detailed architecture of this multi-task model is shown in Fig. 3(a).

MTM-W

In this model, θiw=θw\theta^{w}_{i}=\theta^{w} are shared among different tasks. When a dataset is used, the parameters updated during the training are θw\theta^{w} and θic\theta^{c}_{i}. The detailed architecture of this multi-task model is shown in Fig. 3(b).

MTM-CW

In this model, θic=θc\theta^{c}_{i}=\theta^{c} and θiw=θw\theta^{w}_{i}=\theta^{w} are shared among different tasks. Each dataset has its specific θio\theta^{o}_{i} for label prediction. MTM-CW shared the most information across tasks compared with the other two multi-task models. It enables sharing both character- and word-level information between different biomedical entities, while the other two models only enable sharing part of the information. The detailed architecture of this multi-task model is shown in Fig. 3(c).

Experimental setup

We test our method on the same 15 datasets used by Crichton et al. 2017, and find our model achieves substantially better performance on 14 of them compared with baseline neural network models. Due to space limit, here we report detailed results of the multi-task model on 5 main datasets (Table 1), which altogether cover major biomedical entity types (e.g., genes, proteins, chemicals, diseases). We also include full results on all the 15 datasets in Supplementary Material: Performance comparison on 15 datasets. The performance of the multi-task model is slightly different when trained on 5 datasets compared with trained on 15 datasets (shown in Supplementary Material: Performance comparison on 15 datasets), as the MTL model has access to more data. In our experiments, we follow the experiment setup of Crichton et al. 2017 and divide each dataset into training, development and test sets. We use training and development sets to train the final model. All datasets are publicly available. All datasets can be downloaded from: \hrefhttps://github.com/cambridgeltl/MTL-Bioinformatics-2016https://github.com/cambridgeltl/MTL-Bioinformatics-2016. As part of preprocessing, word labels are encoded using an IOBES scheme. In this scheme, for example, a word describing a gene entity is tagged with ‘‘B-Gene’’ if it is at the beginning of the entity, ‘‘I-Gene’’ if it is in the middle of the entity, and ‘‘E-Gene’’ if it is at the end of the entity. Single-word gene entities are tagged with ‘‘S-Gene’’. All other words not describing entities of interest are tagged as ‘O’. Next, we briefly describe the 5 main datasets and their corresponding state-of-the-art BioNER systems.

The state-of-the-art system reported for the BioCreative II gene mention recognition task adopts semi-supervised learning method with alternating structure optimization (Ando 2007).

BC4CHEMD

The state-of-the-art system reported for the BioCreative IV chemical entity mention recognition task is the CHEMDNER system (Lu et al. 2015), which is based on mixed conditional random fields with Brown clustering of words.

BC5CDR

The state-of-the-art system reported for the most recent BioCreative V chemical and disease mention recognition task is the TaggerOne system (Leaman and Lu 2016), which uses a semi-Markov model for joint entity recognition and normalization.

NCBI-Disease

The NCBI disease dataset was initially introduced for disease name recognition and normalization. It has been widely used for a lot of applications. The state-of-the-art system on this dataset is also the TaggerOne system (Leaman and Lu 2016).

JNLPBA

The state-of-the-art system (Zhou and Su 2004) for the 2004 JNLPBA shared task on biomedical entity (gene/protein, DNA, RNA, cell line, cell type) recognition uses a hidden markov model (HMM). Although this task and the model is a bit old compared with the others, it still remains a competitive benchmark method for comparison.

2 Evaluation metrics

We report the performance of all the compared methods on the test set. We deem each predicted entity as correct only if both the entity boundary and entity types are the same as the ground-truth annotation (i.e., exact match). Then we calculate the precision, recall and F1 scores on all datasets and macro-averaged F1 scores on all entity types. For error analysis, we compare the ratios of false positive (FP) and false negative (FN) labels in the single-task and the multi-task models and include the results in Supplementary Material: Error analysis.

The test set of the BC2GM dataset is constructed slightly differently compared to the test sets of other datasets. BC2GM additionally provides a list of alternative answers for each entity in the test set. A predicted entity is deemed correct as long as it matches the ground truth or one of the alternative answers. We refer to this measurement as alternative match and report scores under both exact match and alternative match for the BC2GM dataset.

3 Pre-trained word embeddings

We initialize the word embedding matrix with pre-trained word vectors from Pyysalo et al. 2013 in all experiments. The pre-trained word vectors can be download from: \hrefhttp://bio.nlplab.org/http://bio.nlplab.org/. These word embeddings are trained using a skip-gram model, as described in Mikolov et al. 2013. These word vectors are trained on three different datasets: (1) abstracts from the PubMed database, (2) abstracts from the PubMed database together with full-text articles from the PubMed Central (PMC), and (3) the entire Pubmed database of abstracts and full-text articles together with the Wikipedia corpus. We found the third set of word vectors lead to best results on development set and therefore used it for the model development. We provide a full comparison of different word embeddings in Supplementary Material: Performance of Word Embeddings. In all experiments, we replace rare words (i.e., words with a frequency of less than 5) with a special token, whose embedding is randomly initialized and fine-tuned during model training.

4 Training details

All the neural network models are trained on one GeForce GTX 1080 GPU. To train our neural models, we use a learning rate of 0.01 with a decay rate of 0.05 applied to every epoch of training. The dimensions of word and character embedding vectors are set to be 200 and 30, respectively (Liu et al. 2018). We adopted 200 (best performance among 100, 200 and 300) for both character- and word-level BiLSTM layers. Note that Liu et al. 2018 considers advanced strategies, such as highway structures, to further improve performance. We did not observe any significant performance boost with these advanced strategies, thus do not adopt these strategies in this work. The performance of the model variations with these advanced strategies can be found in Supplementary Material: Performance of Model Variations. To train the baseline neural network models, we use the default parameter settings as used in their paper (Lample et al. 2016; Habibi et al. 2017; Ma and Hovy 2016) because we found the default parameters also lead to almost optimal performance on the development set.

Results

We compare the proposed single-task (Section 3.1) and multi-task models (Section 3.2) with state-of-the-art BioNER systems (reported for each dataset) and three neural network models from Crichton et al. 2017, Lample et al. 2016; Habibi et al. 2017, and Ma and Hovy 2016. The evaluation metrics include precision, recall and F1 score (Tsai et al. 2006) (Table 2). We denote results of the best system priorly reported for each dataset as ‘‘Dataset Benchmark". For method proposed by Crichton et al. 2017, we quote their experiment results directly. For other neural network models, we repeart each experiment three times with the mean and standard deviation reported (Table 2). To directly compare with the results in Crichton et al. 2017, we measure statistical significance with the same t-test as used in their paper.

We observe that the MTM-CW model achieves significantly higher F1 scores than state-of-the-art benchmark systems (column Dataset Benchmark in Table 2) on all of the five datasets. Following established practice in the literature, we use exact match to compare benchmark performance on all the datasets except for the BC2GM, where we report benchmark performance based on alternative match. Furthermore, MTM-CW generally achieves significantly higher F1 scores than other neural network models. These results show that the proposed multi-task learning neural network significantly outperforms state-of-the-art systems and other neural networks. In particular, the MTM-CW model consistently achieves a better performance than the single task model, demonstrating that multi-task learning is able to successfully leverage information across different datasets and mutually enhance performance on each single task. We further investigate the performance of three multi-task models (MTM-C, MTM-W, and MTM-CW, Table 3). Results show that the best performing multi-task model is MTM-CW, indicating the importance of morphological information captured by character-level BiLSTM as well as lexical and contextual information captured by word-level BiLSTM.

2 Performance on major biomedical entity types

We also conduct more fine-grained comparison of all models on four major biomedical entity types: genes/proteins, chemicals, diseases and cell lines since they are the most often annotated entity types (Fig. 4). Each entity type comes from multiple datasets: genes/proteins from BC2GM and JNLPBA, chemicals from BC4CHEMD and BC5CDR, diseases from BC5CDR and NCBI-Disease, and cell lines from JNLPBA.

The MTM-CW model performs consistently better than the neural network model (Habibi et al. 2017) on all four entity types. It also outperforms the state-of-the-art systems (Benchmark in Fig. 4) on three entity types except for cell lines. These results further confirm that the multi-task neural network model achieves a significantly better performance compared with state-of-art systems and other neural network models for BioNER.

3 Integration of biomedical entity dictionaries

A biomedical entity dictionary is a manually-curated list of entity names that belong to a specific entity type. Traditional BioNER systems make heavy use of these dictionaries in addition to other data. To study whether our approach can benefit from the use of entity dictionaries, we retrieve biomedical entity dictionaries for three entity types (i.e., genes/proteins, chemicals and diseases) from the Comparative Toxicogenomics Database (CTD) (Davis et al. 2017). We use these entity dictionaries in a neural network model in two different ways: (1) dictionary post-processing to match the ‘O’-labeled entities with the dictionary to reduce the false negative rate, or (2) dictionary feature to provide additional information about words into the word-level BiLSTM. This dictionary feature indicates whether a word sequence consisting of a word and its neighbors is present in a dictionary. We consider word sequences of up to six words, which adds 21 additional dimensions for each entity type. We compare the performance of MTM-CW with and without adding dictionaries (Table 4).

We observe no significant performance improvement when biomedical entity dictionaries are included into the MTM-CW model at the pre-processing stage. Moreover, including dictionaries at the post-processing stage even hurts the performance. This is presumably due to a higher false positive rate introduced by the dictionaries, when some words share the surface name with dictionary entities but do not share the same meaning or entity types. These results indicate that our multi-task model, by sharing information at both the character and word levels, is able to learn effective data representations and generalize to new data without the use of external lexicon resources.

4 Comparison on training time

All of the neural network models are trained on one GeForce GTX 1080 GPU. We compare the average training time (seconds per sentence) of our method on the 5 main datasets with the baseline neural models in Table 2. Since our multi-task model requires training on the 5 datasets together, we calculate and compare the average training time on all datasets instead of on each individual one. We find that our single-task neural model STM is the most efficient among the neural models and almost halves the training time (0.71 s/sent.) when compared to Lample et al. 2016; Habibi et al. 2017 (1.59 s/sent.). Compared to the single-task model STM, our multi-task model MTM-CW achieves 8.0% overall F1 improvements with only 5.1% additional training time. The reason that MTM-CW is slightly slower compared with STM is that it takes a few more epochs for MTM-CW to reach convergence when trained on 5 datasets together.

5 Case study

To investigate the major advantages of the multi-task models compared with the single task models, we examine some sentences with predicted labels (Table 5). The true labels and the predicted labels of each model are underlined in a sentence.

One major challenge of BioNER is to recognize a long entity with integrity. In Case 1, the true gene entity is ‘‘endo-beta-1,4-glucanase-encoding genes’’. The single-task model tends to break this whole entity into two parts separated by a comma, while the multi-task model can detect this gene entity as a whole. This result could due to the co-training of multiple datasets containing long entity training examples. Another challenge is to detect the correct boundaries of biomedical entities. In Case 2, the correct protein entity is ‘‘SMase’’ in the phrase ‘‘SMase - sphingomyelin complex structure’’. The single-task models recognize the whole phrase as a protein entity. Our multi-task model is able to detect the correct right boundary of the protein entity, probably also due to seeing more examples from other datasets which may contain ‘‘sphingomyelin’’ as a non-chemical entity. In Case 3, the adjective words ‘‘human’’ and ‘‘complement factor’’ in front of ‘‘H deficiency’’ should be included as part of the true entity. The single-task models missed the adjective words while the multi-task model is able to detect the correct right boundary of the disease entity. In summary, the multi-task model works better at dealing with two critical challenges for BioNER: (1) recognizing long entities with integrity and (2) detecting the correct left and right boundaries of biomedical entities. Both improvements come from collectively training multiple datasets with different entity types and sharing useful information between datasets.

Conclusion

We proposed an neural multi-task learning approach for biomedical named entity recognition. The proposed approach, despite being simple and not requiring manual feature engineering, outperformed state-of-the-art systems and several strong neural network models on benchmark BioNER datasets. We also showed through detailed analysis that the strong performance is achieved by the multi-task model with only marginally added training time, and confirmed that the large performance gains of our approach mainly come from sharing character- and word-level information between biomedical entity types.

Lastly, we highlight several future directions to improve the multi-task BioNER model. First, combining single-task and multi-task models might be a fruitful direction. Second, by further resolving the entity boundary and type conflict problem, we could build a unified system for recognizing multiple types of biomedical entities with high performance and efficiency.

References