Hierarchical Contextualized Representation for Named Entity Recognition
Ying Luo, Fengshun Xiao, Hai Zhao
Introduction
Named Entity Recognition (NER) is one of the fundamental tasks in natural language processing (NLP) that intends to identify words or phrases as the proper names of PER (Person), ORG (Organization), LOC (Location), etc. Currently, most state-of-the-art NER systems (?; ?; ?; ?) employ BiRNNs, specially BiLSTM (?) as the encoder to extract the sequential information.
BiLSTM architectures exist limitations in making full use of global information. First, at each time step, BiLSTM takes current word embedding and past summary states as inputs, making it difficult to capture sentence-level information. (?) simultaneously model the sub-states for individual words and an overall sentence-level state. (?) use a global contextual encoder and mean pooling strategy to capture sentence-level features, though they ignore the different importance of words in the same sentence. Second, though BiLSTM updates the parameters with the iteration of all training instances, it only consumes one instance during both training and predicting. This nature prevents the model from effectively capturing document (dataset)-level information, e.g. for a unique token, its representations in training instances are indicative for recognizing the concerned token. (?) use a pooling operation on different contextualized embeddings to generate global word representations. While they only consider the changes of embeddings for each unique word.
In this paper, we propose a hierarchical contextualized representation architecture to enhance NER modeling. For sentence-level representation, inspired by (?), we embed labels in the same space with word embeddings, and label embeddings are learned from attention mechanism computed with word embeddings to ensure that, each word embedding is much closer to their corresponding label embedding, and farther to other label embeddings. Then, the similarity between a word embedding and its nearest label embedding is regarded as a confidence score for this word. Hence, words with higher confidence scores contribute more to sentence-level representation. The sentence-level representations are then assigned to each token and fed to the encoder as shown in Figure 1. For document-level representation, we adopt a key-value memory component (?) which memorizes all the word embeddings of training instances and their corresponding representations. The attention mechanism is adopted to compute the output of the memory component. The retrieved document-level representation is fused with the original hidden state and fed to the decoder as shown in Figure 1. In this case, the training instances are not only used to train the model parameters, but also involved in inference.
To verify the effectiveness our model, we conduct extensive experiments on three benchmark NER datasets. Experimental results on these benchmarks suggest that our model can achieve state-of-the-art performance on CoNLL-2003 (91.96 without external knowledge and 93.37 with BERT), OntoNotes 5.0 (87.98 without external knowledge and 90.30 with BERT), 87.08 on CoNLL-2002, meaning that our model truly learns and benefits from useful contextualized representations.
Our contributions in this paper are summarized as follows.
We are the first to introduce hierarchical contextualized representations, namely sentence-level and document-level representation, for NER to take full advantage of non-local information.
We introduce the label embedding attention mechanism for sentence-level representation and propose an effective approach to distill document-level information using key-value memory network.
The evaluation results on three benchmark NER datasets show that our model outperforms all previously reported results without external knowledge. Furthermore, with pre-trained language model BERT, we establish new state-of-the-art results on CoNLL-2003 and ontonotes 5.0 datasets.
Related work
Neural Named Entity Recognition Recently, with the development of deep neural network in a wide range of NLP tasks (?; ?; ?; ?; ?; ?), neural network based models build reliable NER systems without hand-crafted features or task-specific knowledge. (?) firstly proposed the BiLSTM-CRF architecture, which is used by most state-of-the-art models. Later, character-level embeddings are concatenated to enhance the representation of rare and out-of-vocabulary words, these embeddings are generated with LSTM (?), CNN (?), and recently IntNet (?). (?) stack BiLSTMs with residual connections between different layers of BiLSTM to add more representational power. More recently, pre-trained language models from huge corpus are adopted to enhance the representation of words (?; ?; ?).
Sentence-level Representation has been adopted to eliminate the limitations of RNNs due to their sequential nature. (?) leverage RNN models to learn sentence-level patterns for NER reranking. (?) model the sub-states for individual words and an overall sentence-level state simultaneously to capture local and non-local contexts. (?) use contextual layer and relation layer to model the relations between words in sentences, and then use gates to fuse local context features into global ones. (?) simplify sentence-level state to average of the hidden states of each individual word from an independent global contextual encoder. Inspired by (?) which use label information to construct text-sequence representations, we adopt label embedding attention to enhance the sentence-level representation learned from an independent BiLSTM. The use of sentence-level information in (?) can be seen as a special case of our model where the attention weight vector is a uniform distribution that assigns equal probabilities to all the words in the sentence.
Document-level Representation (?) considers the dependency structure of word sequence as the global information. (?) dynamically aggregates contextualized embeddings for each unique string and then use a pooling operation to generate a global word representation from these contextualized instances for NER. Different from their work which only uses the contextual word embeddings, our memory component uses the key-value memory networks to memorize the word representations (key) and the hidden states (value) from the sequence labeling encoder. The attention mechanism is then called to calculate the document-level representation.
Model
This section presents our NER model in detail. The overall model architecture is shown in Figure 2, which consists of four components: a decoder (top part), a sequence labeling encoder (upper right part), a sentence-level encoder (bottom right part), and a document-level encoder (left part).
We adopt the IntNet-LSTM-CRF model proposed by (?) as our baseline, which consists of three parts: representation module, sequence labeling encoder and decoder module.
Token Representation Given a sequence of tokens , for each word , we concatenate the word-level and character-level embedding as the joint word representation . is the pre-trained word embedding. The character-level embedding is learned from IntNet, which is a funnel-shaped wide convolutional neural architecture for learning representations of the internal structure of words. The network comprises of convolutions, which implies convolutional blocks. In each convolutional block, the first layer is the convolution which transforms the input, then the concatenation of convolutions with different kernel sizes in the second layer is fed to the next convolutional block. Direct connections from every other layer to all subsequent layers are used like dense connections.
Sequence Labeling Encoder The concatenation of word-level and character-level embeddings is then fed into the sequence labeling BiLSTM, which represents the sequential information at each step.
where and are trainable parameters, respectively.
Decoder Conditional random field (CRF) (?) has been widely used in state-of-the-art NER models (?; ?) to help make decisions when considering strong connections between output tags. During decoding, the Viterbi algorithm is applied to search the label sequence with the highest probability. For being a predicted sequence of labels with same length as . We define its score as:
where represents the transmission score from the to , is the score of the tag of the word from the sequence labeling encoder.
The CRF model defines a family of conditional probability over all possible tag sequences :
during training, we consider the maximum log probability of the correct sequence of tags. While decoding, we search the label sequence with maximum score:
Sentence-level Representation
Document-level Representation
In terms of memory network, we introduce document-aware representations of the unique word in training instances as an extra knowledge source to help the prediction. Memory network was originally proposed by (?) in the domain of question answering (QA) for prediction, where the long-term memory acts as a dynamic knowledge base. (?) further introduces a key-value memory networks, which utilize different encodings in the addressing and output stages. The keys are designed to help match the question, while the values are to generate the response.
We adopt the key-value memory component M to memorize document-level contextualized representation. Memory slots are defined as pairs of vectors . In each single slot, the key represents the word embedding , and the value is the corresponding hidden states from sequence labeling encoder for each token in training instances. So the same word may occupy in many different slots because of changing embeddings and representations under different contexts. Table 1 shows an example of using training instances to help indicate the NE type of queried token.
Memory Update The word embeddings are fine-tuned during training and used to update the key part of the memory. The sequence labeling encoder generates the hidden states to update the value part. Supposing the states of the -th token is changed after computation, the -th slot in the memory will be rewritten. Each memory slot will be updated once in one epoch.
Memory Query For the -th word in the sentence, we distill all the contextualized representations for this word in the memory through an inverted index that finds a subset of size , where the inverted index records the positions of the unique word in the memory as shown in Table 1. represents the number of occurrences of this word among the training instances.
The attention operation is called to compute the weight of document-level representation. For the unique word, the memory key is used as the attention key, the memory value is used as the attention value. Then the embedding of the queried word serves as the attention query . Here, we consider three compatibility functions :
where represents the dimension of word embeddings.
Memory Response The document-level representation is computed as:
where is a hyperparameter, indicating how much document-aware information is adopted, 0 for document-level representation only and 1 for discarding all document-level information at all.
Experiment
Our proposed representations are evaluated on three benchmark NER datasets: CoNLL-2003 (?) and OntoNotes 5.0 (?) English NER datasets, CoNLL-2002 Spanish NER (?) dataset.
CoNLL-2003 English NER consists of 22,137 sentences totally and is split into 14,987, 3,466 and 3,684 sentences for the training, development set and test sets, respectively. It is tagged with four linguistic entity types (PER, LOC, ORG, MISC).
CoNLL-2002 Spanish NER consists of 11,752 sentences totally and is split into 8,322, 1,914 and 1,516 sentences for the training, development and test sets, respectively. It is also tagged with four linguistic entity types (PER, LOC, ORG, MISC).
OntoNotes 5.0 consists of 76,714 sentences from a wide variety of sources (magazine, telephone conversation, newswire, etc.). Following (?; ?), we use the portion of the dataset with gold-standard named entity annotations, and thus exclude the New Testaments portion. It is tagged with eighteen entity types (PERSON, CARDINAL, LOC, PRODUCT, etc.).
Metric We use the BIOES sequence labeling scheme instead of BIO for these three datasets during training. As for test, we convert the prediction results back to the BIO scheme and use the standard conlleval script to compute the score.
Setup
Pre-trained Word Embeddings. For the CoNLL-2003 and OntoNotes 5.0 English datasets, we use the publicly available pre-trained 100 GloVe (?) embeddings. For CoNLL-2002 Spanish dataset, we train 64 GloVe embeddings with the minimum frequency of occurrence as 3, and the window size of 5. The word embeddings are fine-tuned during training. Character Embeddings. We train the IntNet character embeddings (?). The dimension of character embeddings is 32, which is randomly initialized, the filter size of the initial convolution is 32 and that of other convolutions is 16. Different from (?), we set filters as size [3; 5] for all the kernels, and the number of convolutional layers is 7. Parameters. We follow the work of (?), and conduct optimization with the stochastic gradient descent Code will be available at https://github.com/cslydia/Hire-NER.. The batch size is set as 10, the initial learning rate is set to 0.015 and will shrunk by 5% after each epoch. The hidden size of sequence labeling encoder and the sentence-level encoder are set as 256 and 128, respectively. We apply dropout to embeddings and hidden states with a rate of 0.5. The used to fuse original hidden state and document-level representation is set as 0.3 empirically. For each type of NEs, we randomly select hundreds of NEs, and calculate the average of the word embeddings as its label embedding.
Results and Comparisons
Tables 2, 3, 4 compare our model to existing state-of-the-art approaches on the three benchmark datasets. Our model surpasses previous state-of-the-art approaches on all the three datasets. On CoNLL-2003 dataset, we compare our model with the state-of-the-art models, including the models that use global information to enhance the representation (?; ?; ?; ?). We also incorporate pre-trained language model BERT (?) for fair comparisons with the models which also use pre-trained language models or other external knowledge. Some of the results (?; ?) are not comparable to our results directly, because their final models are trained on both training and development datasets. On CoNLL-2002 Spanish dataset, our model achieves 87.08 score without external knowledge, which surpasses previous best score by 0.4. Considering that the above two datasets are relatively small, we further conduct experiment on a much more large OntoNotes 5.0 dataset, which also has more entity types. We compare our model with the previous model that also reported results on it (?; ?; ?). As shown in Table 4, our model shows a significant advantage on this dataset, which outperforms previous state-of-the-art results substantially at 87.08 (+0.31) without BERT, and 90.30 (+0.59) with BERT. More notably, our model without external knowledge surpasses the previous model (?), which use extra lexicon information of 120 entity types from Wikipedia. Overall, the comparisons on these three benchmark datasets well demonstrate that our model truly learns and benefits from useful sentence-level and document-level representation without the support from external knowledge.
Ablation Study
In this experiment, we individually adopt two hierarchical contextualized representations to enhance the representation of tokens: sentence-level representation for assigning the sentence state to each token and document-level representation for inference. Table 5 shows the score raise and relative error reduction brought by each of the two hierarchical representation on the three benchmark datasets. We discover that both sentence-level and document-level representations enhance the baseline. By combing these two representations together, we get a larger gain of 0.36 / 0.43 / 0.40, respectively.
We further analyze the two hierarchical representations by adopting different strategies. (?) perform mean pooling over all the tokens to generate sentence-level representation. We further conduct experiments to investigate the three compatibility functions used to employ memorized information. As shown in Table 6, compared with the mean pooling strategy, our label-embedding attention mechanism raises the score by 0.25. Among the three compatibility functions to compute the weight of query word and memorized slots, cosine similarity performs best, while dot-product performs worst. (?) use scaled dot-product to counteract the dot products growth in magnitude, showing better than dot product. Cosine similarity calculates the inner product of word vectors with unit length, and can further solve the inconsistency between the embeddings and the similarity measurement. Thus, we eventually adopt cosine similarity as the compatibility function.
Memory Size and Time Consuming
Figure 3 illustrates our model performance and time proportion compared to the baseline with respect to the max queried subset size for each unique word in the memory query step. For words occurring more than times in the corpus (these words are more likely to be stop words when is large), we only randomly select slots in the subset to compute the document-level representation. For fair comparisons, we keep the IntNet layer, sequence labeling layer and CRF layer the same for all the experiments. The consumed time of our model is only 19% more than the baseline on CoNLL-2003 dataset even with the max memory size as 500. Therefore, our model brings slight increase on time consumption. When is less than 500, the larger may incorporate more useful contextualized representation for practice words and improve the results accordingly, when is larger than 500, which may involve more stop words, our model drops slightly.
Improvement Discussion
Table 7 presents the score of in-both-vocabulary words (IV), out-of-training-vocabulary words (OOTV), out-of-embedding-vocabulary words (OOEV), and out-of-both-vocabulary words (OOBV) on CoNLL-2003 datset. According to our statistic, 63.40% / 52.43% / 84.68% of the NEs in the test set of CoNLL-2003, CoNLL-2002, and OntoNotes datasets are located in the IV part, respectively. Therefore, it is of great importance to focus on this part. We adopt memory network to memorize and retrieve the global representations and use the memorized training instances directly to participate in inference, which greatly improves both the precision and recall of the NEs in IV part, in which our model outperforms baseline by 0.39 in terms of score. For OOV NEs, sentence-level representation can help these concerned tokens aware of the entire sentence, thus enhance the performance. The improvement is 0.44 / 0.43 score for OOTV NEs and OOBV NEs, respectively.
Conclusions
In this paper, we adopt hierarchical contextualized representations to enhance the performance of named entity recognition (NER). Our model makes full use of the training instances and the spatial information of the embedding space by incorporating sentence-level representation and document-level representation. We consider the importance of words in the sentences and weight their contributions with the label embedding attention for the sentence-level representation. For words shown in training instances, we memorize the representations of these instances, and involve these representations for inference during test. Empirical results on three benchmark datasets (CoNLL-2003 and Ontonotes 5.0 English datasets, CoNLL-2002 Spanish dataset) show that our model outperforms previous state-of-the-art systems with or without pre-trained language models respectively.