Attending to Characters in Neural Sequence Labeling Models

Marek Rei, Gamal K. O. Crichton, Sampo Pyysalo

Introduction

Many NLP tasks, including named entity recognition (NER), part-of-speech (POS) tagging and shallow parsing can be framed as types of sequence labeling. The development of accurate and efficient sequence labeling models is thereby useful for a wide range of downstream applications. Work in this area has traditionally involved task-specific feature engineering – for example, integrating gazetteers for named entity recognition, or using features from a morphological analyser in POS-tagging. Recent developments in neural architectures and representation learning have opened the door to models that can discover useful features automatically from the data. Such sequence labeling systems are applicable to many tasks, using only the surface text as input, yet are able to achieve competitive results [Collobert et al. (2011, Irsoy and Cardie (2014].

Current neural models generally make use of word embeddings, which allow them to learn similar representations for semantically or functionally similar words. While this is an important improvement over count-based models, they still have weaknesses that should be addressed. The most obvious problem arises when dealing with out-of-vocabulary (OOV) words – if a token has never been seen before, then it does not have an embedding and the model needs to back-off to a generic OOV representation. Words that have been seen very infrequently have embeddings, but they will likely have low quality due to lack of training data. The approach can also be sub-optimal in terms of parameter usage – for example, certain suffixes indicate more likely POS tags for these words, but this information gets encoded into each individual embedding as opposed to being shared between the whole vocabulary.

In this paper, we construct a task-independent neural network architecture for sequence labeling, and then extend it with two different approaches for integrating character-level information. By operating on individual characters, the model is able to infer representations for previously unseen words and share information about morpheme-level regularities. We propose a novel architecture for combining character-level representations with word embeddings using a gating mechanism, also referred to as attention, which allows the model to dynamically decide which source of information to use for each word. In addition, we describe a new objective for model training where the character-level representations are optimised to mimic the current state of word embeddings.

We evaluate the neural models on 8 datasets from the fields of NER, POS-tagging, chunking and error detection in learner texts. Our experiments show that including a character-based component in the sequence labeling model provides substantial performance improvements on all the benchmarks. In addition, the attention-based architecture achieves the best results on all evaluations, while requiring a smaller number of parameters.

Bidirectional LSTM for sequence labeling

We first describe a basic word-level neural network for sequence labeling, following the models described by ?) and ?), and then propose two alternative methods for incorporating character-level information.

Figure 1 shows the general architecture of the sequence labeling network. The model receives a sequence of tokens (w1,...,wT)(w_{1},...,w_{T}) as input, and predicts a label corresponding to each of the input tokens. The tokens are first mapped to a distributed vector space, resulting in a sequence of word embeddings (x1,...,xT)(x_{1},...,x_{T}). Next, the embeddings are given as input to two LSTM [Hochreiter and Schmidhuber (1997] components moving in opposite directions through the text, creating context-specific representations. The respective forward- and backward-conditioned representations are concatenated for each word position, resulting in representations that are conditioned on the whole sequence:

We include an extra narrow hidden layer on top of the LSTM, which proved to be a useful modification based on development experiments. An additional hidden layer allows the model to detect higher-level feature combinations, while constraining it to be small forces it to focus on more generalisable patterns:

where WdW_{d} is a weight matrix between the layers, and the size of dtd_{t} is intentionally kept small.

Finally, to produce label predictions, we use either a softmax layer or a conditional random field (CRF, ?)). The softmax calculates a normalised probability distribution over all the possible labels for each word:

where P(yt=k∣dt)P(y_{t}=k|d_{t}) is the probability of the label of the tt-th word (yty_{t}) being kk, KK is the set of all possible labels, and Wo,kW_{o,k} is the kk-th row of output weight matrix WoW_{o}. To optimise this model, we minimise categorical crossentropy, which is equivalent to minimising the negative log-probability of the correct labels:

Following ?), we can also use a CRF as the output layer, which conditions each prediction on the previously predicted label. In this architecture, the last hidden layer is used to predict confidence scores for the word having each of the possible labels. A separate weight matrix is used to learn transition probabilities between different labels, and the Viterbi algorithm is used to find an optimal sequence of weights. Given that yy is a sequence of labels [y1,...,yT][y_{1},...,y_{T}], then the CRF score for this sequence can be calculated as:

where At,ytA_{t,y_{t}} shows how confident the network is that the label on the tt-th word is yty_{t}. Byt,yt+1B_{y_{t},y_{t+1}} shows the likelihood of transitioning from label yty_{t} to label yt+1y_{t+1}, and these values are optimised during training. The output from the model is the sequence of labels with the largest score s(y)s(y), which can be found efficiently using the Viterbi algorithm. In order to optimise the CRF model, the loss function maximises the score for the correct label sequence, while minimising the scores for all other sequences:

where Y~\widetilde{Y} is the set of all possible label sequences.

Character-level sequence labeling

Distributed embeddings map words into a space where semantically similar words have similar vector representations, allowing the models to generalise better. However, they still treat words as atomic units and ignore any surface- or morphological similarities between different words. By constructing models that operate over individual characters in each word, we can take advantage of these regularities. This can be particularly useful for handling unseen words – for example, if we have never seen the word cabinets before, a character-level model could still infer a representation for this word if it has previously seen the word cabinet and other words with the suffix -s. In contrast, a word-level model can only represent this word with a generic out-of-vocabulary representation, which is shared between all other unseen words.

Research into character-level models is still in fairly early stages, and models that operate exclusively on characters are not yet competitive to word-level models on most tasks. However, instead of fully replacing word embeddings, we are interested in combining the two approaches, thereby allowing the model to take advantage of information at both granularity levels. The general outline of our approach is shown in Figure 2. Each word is broken down into individual characters, these are then mapped to a sequence of character embeddings (c1,...,cR)(c_{1},...,c_{R}), which are passed through a bidirectional LSTM:

We then use the last hidden vectors from each of the LSTM components, concatenate them together, and pass the result through a separate non-linear layer.

where WmW_{m} is a weight matrix mapping the concatenated hidden vectors from both LSTMs into a joint word representation mm, built from individual characters.

We now have two alternative feature representations for each word – xtx_{t} from Section 2 is an embedding learned on the word level, and m(t)m^{(t)} is a representation dynamically built from individual characters in the tt-th word of the input text. Following ?), one possible approach is to concatenate the two vectors and use this as the new word-level representation for the sequence labeling model:

This approach, also illustrated in Figure 2, assumes that the word-level and character-level components learn somewhat disjoint information, and it is beneficial to give them separately as input to the sequence labeler.

Attention over character features

Alternatively, we can have the word embedding and the character-level component learn the same semantic features for each word. Instead of concatenating them as alternative feature sets, we specifically construct the network so that they would learn the same representations, and then allow the model to decide how to combine the information for each specific word.

We first construct the word representation from characters using the same architecture – a bidirectional LSTM operates over characters, and the last hidden states are used to create vector mm for the input word. Instead of concatenating this with the word embedding, the two vectors are added together using a weighted sum, where the weights are predicted by a two-layer network:

where Wz(1)W^{(1)}_{z}, Wz(2)W^{(2)}_{z} and Wz(3)W^{(3)}_{z} are weight matrices for calculating zz, and σ()\sigma() is the logistic function with values in the range $.Thevector. The vectorzhasthesamedimensionsashas the same dimensions asxororm$, acting as the weight between the two vectors. It allows the model to dynamically decide how much information to use from the character-level component or from the word embedding. This decision is done for each feature separately, which adds extra flexiblity – for example, words with regular suffixes can share some character-level features, whereas irregular words can store exceptions into word embeddings. Furthermore, previously unknown words are able to use character-level regularities whenever possible, and are still able to revert to using the generic OOV token when necessary.

The main benefits of character-level modeling are expected to come from improved handling of rare and unseen words, whereas frequent words are likely able to learn high-quality word-level embeddings directly. We would like to take advantage of this, and train the character component to predict these word embeddings. Our attention-based architecture requires the learned features in both word representations to align, and we can add in an extra constraint to encourage this. During training, we add a term to the loss function that optimises the vector mm to be similar to the word embedding xx:

Equation 12 maximises the cosine similarity between m(t)m^{(t)} and xtx_{t}. Importantly, this is done only for words that are not out-of-vocabulary – we want the character-level component to learn from the word embeddings, but this should exclude the OOV embedding, as it is shared between many words. We use gtg_{t} to set this cost component to for any OOV tokens.

While the character component learns general regularities that are shared between all the words, individual word embeddings provide a way for the model to store word-specific information and any exceptions. Therefore, while we want the character-based model to shift towards predicting high-quality word embeddings, it is not desireable to optimise the word embeddings towards the character-level representations. This can be achieved by making sure that the optimisation is performed only in one direction; in Theano [Bergstra et al. (2010], the disconnected_grad function gives the desired effect.

Datasets

We evaluate the sequence labeling models and character architectures on 8 different datasets. Table 1 contains information about the number of labels and dataset sizes for each of them.

CoNLL00: The CoNLL-2000 dataset [Tjong Kim Sang and Buchholz (2000] is a frequently used benchmark for the task of chunking. Wall Street Journal Sections 15-18 from the Penn Treebank are used for training, and Section 20 as the test data. As there is no official development set, we separated some of the training set for this purpose.

CoNLL03: The CoNLL-2003 corpus [Tjong Kim Sang and De Meulder (2003] was created for the shared task on language-independent NER. We use the English section of the dataset, containing news stories from the Reuters Corpushttp://about.reuters.com/researchandstandards/corpus/.

PTB-POS: The Penn Treebank POS-tag corpus [Marcus et al. (1993] contains texts from the Wall Street Journal, annotated for part-of-speech tags. The PTB label set includes 36 main tags and an additional 12 tags covering items such as punctuation.

FCEPUBLIC: The publicly released subset of the First Certificate in English (FCE) dataset contains short essays written by language learners and manual corrections by examiners [Yannakoudakis et al. (2011]. We use a version of this corpus converted into a binary error detection task, where each token is labeled as being correct or incorrect in the given context.

BC2GM: The BioCreative II Gene Mention corpus [Smith et al. (2008] consists of 20,000 sentences from biomedical publication abstracts and is annotated for mentions of the names of genes, proteins and related entities using a single NE class.

CHEMDNER: The BioCreative IV Chemical and Drug [Krallinger et al. (2015] NER corpus consists of 10,000 abstracts annotated for mentions of chemical and drug names using a single class. We make use of the official splits provided by the shared task organizers.

JNLPBA: The JNLPBA corpus [Kim et al. (2004] consists of 2,404 biomedical abstracts and is annotated for mentions of five entity types: cell line, cell type, dna, rna, and protein. The corpus was derived from GENIA corpus entity annotations for use in the shared task organized in conjuction with the BioNLP 2004 workshop.

GENIA-POS: The GENIA corpus [Ohta et al. (2002] is one of the most widely used resources for biomedical NLP and has a rich set of annotations including parts of speech, phrase structure syntax, entity mentions, and events. Here, we make use of the GENIA POS annotations, which cover 2,000 PubMed abstracts (approx. 20,000 sentences). We use the same 210-document test set as ?), and additionally split off a sample of 210 from the remaining documents as a development set.

Experiment settings

For data prepocessing, all digits were replaced with the character ’0’. Any words that occurred only once in the training data were replaced by the generic OOV token for word embeddings, but were still used in the character-level components. The word embeddings were initialised with publicly available pretrained vectors, created using word2vec [Mikolov et al. (2013], and then fine-tuned during model training. For the general-domain datasets we used 300-dimensional vectors trained on Google Newshttps://code.google.com/archive/p/word2vec/; for the biomedical datasets we used 200-dimensional vectors trained on PubMed and PMChttp://bio.nlplab.org/. The embeddings for characters were set to length 5050 and initialised randomly.

The LSTM layer size was set to 200200 in each direction for both word- and character-level components. The hidden layer dd has size 5050, and the combined representation mm has the same length as the word embeddings. CRF was used as the output layer for all the experiments – we found that this gave most benefits to tasks with larger numbers of possible labels. Parameters were optimised using AdaDelta [Zeiler (2012] with default learning rate 1.01.0 and sentences were grouped into batches of size 6464. Performance on the development set was measured at every epoch and training was stopped if performance had not improved for 7 epochs; the best-performing model on the development set was then used for evaluation on the test set. In order to avoid any outlier results due to randomness in the model initialisation, we trained each configuration with 10 different random seeds and present here the averaged results.

When evaluating on each dataset, we report the measures established in previous work. Token-level accuracy is used for PTB-POS and GENIA-POS; F0.5F_{0.5} score over the erroneous words for FCEPUBLIC; the official evaluation script for BC2GM which allows for alternative correct entity spans; and microaveraged mention-level F1F_{1} score for the remaining datasets.

Results

While optimising the hyperparameters for each dataset separately would likely improve individual performance, we conduct more controlled experiments on a task-independent model. Therefore, we use the same hyperparameters from Section 6 on all datasets, and the development set is only used for the stopping condition. With these experiments, we wish to determine 1) on which sequence labeling tasks do character-based models offer an advantange, and 2) which character-based architecture performs better.

Results for the different model architectures on all 8 datasets are shown in Table 2. As can be seen, including a character-based component in the sequence labeling architecture improves performance on every benchmark. The NER datasets have the largest absolute improvement – the model is able to learn character-level patterns for names, and also improve the handling of any previously unseen tokens.

Compared to concatenating the word- and character-level representations, the attention-based character model outperforms the former on all evaluations. The mechanism for dynamically deciding how much character-level information to use allows the model to better handle individual word representations, giving it an advantage in the experiments. Visualisation of the attention values in Figure 3 shows that the model is actively using character-based features, and the attention areas vary between different words.

The results of this general tagging architecture are competitive, even when compared to previous work using hand-crafted features. The network achieves 97.27% on PTB-POS compared to 97.55% by ?), and 72.70% on JNLPBA compared to 72.55% by ?). In some cases, we are also able to beat the previous best results – 87.99% on BC2GM compared to 87.48% by ?), and 41.88% on FCEPUBLIC compared to 41.1% by ?). ?) report a considerably higher result of 90.94% on CoNLL03, indicating that the chosen hyperparameters for the baseline system are suboptimal for this specific task. Compared to the experiments presented here, their model used the IOBES tagging scheme instead of the original IOB, and embeddings pretrained with a more specialised method that accounts for word order.

It is important to also compare the parameter counts of alternative neural architectures, as this shows their learning capacity and indicates their time requirements in practice. Table 3 contains the parameter counts on three representative datasets. While keeping the model hyperparameters constant, the character-level models require additional parameters for the character composition and character embeddings. However, the attention-based model uses fewer parameters compared to the concatenation approach. When the two representations are concatenated, the overall word representation size is increased, which in turn increases the number of parameters required for the word-level bidirectional LSTM. Therefore, the attention-based character architecture achieves improved results even with a smaller parameter footprint.

Related work

There is a wide range of previous work on constructing and optimising neural architectures applicable to sequence labeling. ?) described one of the first task-independent neural tagging models using convolutional neural networks. They were able to achieve good results on POS tagging, chunking, NER and semantic role labeling, without relying on hand-engineered features. ?) experimented with multi-layer bidirectional Elman-style recurrent networks, and found that the deep models outperformed conditional random fields on the task of opinion mining. ?) described a bidirectional LSTM model with a CRF layer, which included hand-crafted features specialised for the task of named entity recognition. ?) evaluated a range of neural architectures, including convolutional and recurrent networks, on the task of error detection in learner writing. The word-level sequence labeling model described in this paper follows the previous work, combining useful design choices from each of them. In addition, we extended the model with two alternative character-level architectures, and evaluated its performance on 8 different datasets.

Character-level models have the potential of capturing morpheme patterns, thereby improving generalisation on both frequent and unseen words. In recent years, there has been an increase in research into these models, resulting in several interesting applications. ?) described a character-level neural model for machine translation, performing both encoding and decoding on individual characters. ?) implemented a language model where encoding is performed by a convolutional network and LSTM over characters, whereas predictions are given on the word-level. ?) proposed a method for learning both word embeddings and morphological segmentation with a bidirectional recurrent network over characters. There is also research on performing parsing [Ballesteros et al. (2015] and text classification [Zhang et al. (2015] with character-level neural models. ?) proposed a neural architecture that replaces word embeddings with dynamically-constructed character-based representations. We applied a similar method for operating over characters, but combined them with word embeddings instead of replacing them, as this allows the model to benefit from both approaches. ?) described a model where the character-level representation is combined with word embeddings through concatenation. In this work, we proposed an alternative architecture, where the representations are combined using an attention mechanism, and evaluated both approaches on a range of tasks and datasets. Recently, ?) have also described a related method for the task of language modelling, combining characters and word embeddings using gating.

Conclusion

Developments in neural network research allow for model architectures that work well on a wide range of sequence labeling datasets without requiring hand-crafted data. While word-level representation learning is a powerful tool for automatically discovering useful features, these models still come with certain weaknesses – rare words have low-quality representations, previously unseen words cannot be modeled at all, and morpheme-level information is not shared with the whole vocabulary.

In this paper, we investigated character-level model components for a sequence labeling architecture, which allow the system to learn useful patterns from sub-word units. In addition to a bidirectional LSTM operating over words, a separate bidirectional LSTM is used to construct word representations from individual characters. We proposed a novel architecture for combining the character-based representation with the word embedding by using an attention mechanism, allowing the model to dynamically choose which information to use from each information source. In addition, the character-level composition function is augmented with a novel training objective, optimising it to predict representations that are similar to the word embeddings in the model.

The evaluation was performed on 8 different sequence labeling datasets, covering a range of tasks and domains. We found that incorporating character-level information into the model improved performance on every benchmark, indicating that capturing features regarding characters and morphmes is indeed useful in a general-purpose tagging system. In addition, the attention-based model for combining character representations outperformed the concatenation method used in previous work in all evaluations. Even though the proposed method requires fewer parameters, the added ability of controlling how much character-level information is used for each word has led to improved performance on a range of different tasks.

References