TENER: Adapting Transformer Encoder for Named Entity Recognition

Hang Yan, Bocao Deng, Xiaonan Li, Xipeng Qiu

Introduction

The named entity recognition (NER) is the task of finding the start and end of an entity in a sentence and assigning a class for this entity. NER has been widely studied in the field of natural language processing (NLP) because of its potential assistance in question generation Zhou et al. (2017), relation extraction Miwa and Bansal (2016), and coreference resolution Fragkou (2017). Since Collobert et al. (2011), various neural models have been introduced to avoid hand-crafted features Huang et al. (2015); Ma and Hovy (2016); Lample et al. (2016).

NER is usually viewed as a sequence labeling task, the neural models usually contain three components: word embedding layer, context encoder layer, and decoder layer Huang et al. (2015); Ma and Hovy (2016); Lample et al. (2016); Chiu and Nichols (2016); Chen et al. (2019); Zhang et al. (2018); Gui et al. (2019b). The difference between various NER models mainly lies in the variance in these components.

Recurrent Neural Networks (RNNs) are widely employed in NLP tasks due to its sequential characteristic, which is aligned well with language. Specifically, bidirectional long short-term memory networks (BiLSTM) Hochreiter and Schmidhuber (1997) is one of the most widely used RNN structures. (Huang et al., 2015) was the first one to apply the BiLSTM and Conditional Random Fields (CRF) Lafferty et al. (2001) to sequence labeling tasks. Owing to BiLSTM’s high power to learn the contextual representation of words, it has been adopted by the majority of NER models as the encoder Ma and Hovy (2016); Lample et al. (2016); Zhang et al. (2018); Gui et al. (2019b).

Recently, Transformer Vaswani et al. (2017) began to prevail in various NLP tasks, like machine translation Vaswani et al. (2017), language modeling Radford et al. (2018), and pretraining models Devlin et al. (2018). The Transformer encoder adopts a fully-connected self-attention structure to model the long-range context, which is the weakness of RNNs. Moreover, Transformer has better parallelism ability than RNNs. However, in the NER task, Transformer encoder has been reported to perform poorly Guo et al. (2019), our experiments also confirm this result. Therefore, it is intriguing to explore the reason why Transformer does not work well in NER task.

In this paper, we analyze the properties of Transformer and propose two specific improvements for NER.

The first is that the sinusoidal position embedding used in the vanilla Transformer is aware of distance but unaware of the directionality. In addition, this property will lose when used in the vanilla Transformer. However, both the direction and distance information are important in the NER task. For example in Fig 1, words after “in” are more likely to be a location or time than words before it, and words before “Inc.” are mostly likely to be of the entity type “ORG”. Besides, an entity is a continuous span of words. Therefore, the awareness of distance might help the word better recognizes its neighbor. To endow the Transformer with the ability of direction- and distance-awareness, we adopt the relative positional encoding Shaw et al. (2018); Huang et al. (2019); Dai et al. (2019). instead of the absolute position encoding. We propose a revised relative positional encoding that uses fewer parameters and performs better.

The second is an empirical finding. The attention distribution of the vanilla Transformer is scaled and smooth. But for NER, a sparse attention is suitable since not all words are necessary to be attended. Given a current word, a few contextual words are enough to judge its label. The smooth attention could include some noisy information. Therefore, we abandon the scale factor of dot-production attention and use an un-scaled and sharp attention.

With the above improvements, we can greatly boost the performance of Transformer encoder for NER.

Other than only using Transformer to model the word-level context, we also tried to apply it as a character encoder to model word representation with character-level information. The previous work has proved that character encoder is necessary to capture the character-level features and alleviate the out-of-vocabulary (OOV) problem Lample et al. (2016); Ma and Hovy (2016); Chiu and Nichols (2016); Xin et al. (2018). In NER, CNN is commonly used as the character encoder. However, we argue that CNN is also not perfect for representing character-level information, because the receptive field of CNN is limited, and the kernel size of the CNN character encoder is usually 3, which means it cannot correctly recognize 2-gram or 4-gram patterns. Although we can deliberately design different kernels, CNN still cannot solve patterns with discontinuous characters, such as “un..ily” in “unhappily” and “unnecessarily”. Instead, the Transformer-based character encoder shall not only fully make use of the concurrence power of GPUs, but also have the potentiality to recognize different n-grams and even discontinuous patterns. Therefore, in this paper, we also try to use Transformer as the character encoder, and we compare four kinds of character encoders.

In summary, to improve the performance of the Transformer-based model in the NER task, we explicitly utilize the directional relative positional encoding, reduce the number of parameters and sharp the attention distribution. After the adaptation, the performance raises a lot, making our model even performs better than BiLSTM based models. Furthermore, in the six NER datasets, we achieve state-of-the-art performance among models without considering the pre-trained language models or designed features.

Related Work

Collobert et al. (2011) utilized the Multi-Layer Perceptron (MLP) and CNN to avoid using task-specific features to tackle different sequence labeling tasks, such as Chunking, Part-of-Speech (POS) and NER. In Huang et al. (2015), BiLSTM-CRF was introduced to solve sequence labeling questions. Since then, the BiLSTM has been extensively used in the field of NER Chiu and Nichols (2016); Dong et al. (2016); Yang et al. (2018); Ma and Hovy (2016).

Despite BiLSTM’s great success in the NER task, it has to compute token representations one by one, which massively hinders full exploitation of GPU’s parallelism. Therefore, CNN has been proposed by Strubell et al. (2017); Gui et al. (2019a) to encode words concurrently. In order to enlarge the receptive field of CNNs, (Strubell et al., 2017) used iterative dilated CNNs (ID-CNN).

Since the word shape information, such as the capitalization and n-gram, is important in recognizing named entities, CNN and BiLSTM have been used to extract character-level information Chiu and Nichols (2016); Lample et al. (2016); Ma and Hovy (2016); Strubell et al. (2017); Chen et al. (2019).

Almost all neural-based NER models used pre-trained word embeddings, like Word2vec and Glove Pennington et al. (2014); Mikolov et al. (2013). And when contextual word embeddings are combined, the performance of NER models will boost a lot Peters et al. (2017, 2018); Akbik et al. (2018). ELMo introduced by Peters et al. (2018) used the CNN character encoder and BiLSTM language models to get contextualized word representations. Except for the BiLSTM based pre-trained models, BERT was based on Transformer Devlin et al. (2018).

2 Transformer

Transformer was introduced by (Vaswani et al., 2017), which was mainly based on self-attention. It achieved great success in various NLP tasks. Since the self-attention mechanism used in the Transformer is unaware of positions, to avoid this shortage, position embeddings were used Vaswani et al. (2017); Devlin et al. (2018). Instead of using the sinusoidal position embedding Vaswani et al. (2017) and learned absolute position embedding, Shaw et al. (2018) argued that the distance between two tokens should be considered when calculating their attention score. Huang et al. (2019) reduced the computation complexity of relative positional encoding from O(l2d)O(l^{2}d) to O(ld)O(ld), where ll is the length of sequences and dd is the hidden size. Dai et al. (2019) derived a new form of relative positional encodings, so that the relative relation could be better considered.

where QtQ_{t} is the query vector of the ttth token, jj is the token the ttth token attends. KjK_{j} is the key vector representation of the jjth token. The softmax is along the last dimension. Instead of using one group of WqW_{q}, WkW_{k}, WvW_{v}, using several groups will enhance the ability of self-attention. When several groups are used, it is called multi-head self-attention, the calculation can be formulated as follows,

The output of the multi-head attention will be further processed by the position-wise feed-forward networks, which can be represented as follows,

2.2 Position Embedding

The self-attention is not aware of the positions of different tokens, making it unable to capture the sequential characteristic of languages. In order to solve this problem, (Vaswani et al., 2017) suggested to use position embeddings generated by sinusoids of varying frequency. The ttth token’s position embedding can be represented by the following equations

where ii is in the range of [0,d2][0,\frac{d}{2}], dd is the input dimension. This sinusoid based position embedding makes Transformer have an ability to model the position of a token and the distance of each two tokens. For any fixed offset kk, PEt+kPE_{t+k} can be represented by a linear transformation of PEtPE_{t} Vaswani et al. (2017).

Proposed Model

In this paper, we utilize the Transformer encoder to model the long-range and complicated interactions of sentence for NER. The structure of proposed model is shown in Fig 2. We detail each parts in the following sections.

To alleviate the problems of data sparsity and out-of-vocabulary (OOV), most NER models adopted the CNN character encoder Ma and Hovy (2016); Ye and Ling (2018); Chen et al. (2019) to represent words. Compared to BiLSTM based character encoder (Lample et al., 2016; Ghaddar and Langlais, 2018), CNN is more efficient. Since Transformer can also fully exploit the GPU’s parallelism, it is interesting to use Transformer as the character encoder. A potential benefit of Transformer-based character encoder is to extract different n-grams and even uncontinuous character patterns, like “un..ily” in “unhappily” and “uneasily”. For the model’s uniformity, we use the “adapted Transformer” to represent the Transformer introduced in next subsection.

The final word embedding is the concatenation of the character features extracted by the character encoder and the pre-trained word embeddings.

2 Encoding Layer with Adapted Transformer

Although Transformer encoder has potential advantage in modeling long-range context, it is not working well for NER task. In this paper, we propose an adapted Transformer for NER task with two improvements.

Inspired by the success of BiLSTM in NER tasks, we consider what properties the Transformer lacks compared to BiLSTM-based models. One observation is that BiLSTM can discriminatively collect the context information of a token from its left and right sides. But it is not easy for the Transformer to distinguish which side the context information comes from.

Although the dot product between two sinusoidal position embeddings is able to reflect their distance, it lacks directionality and this property will be broken by the vanilla Transformer attention. To illustrate this, we first prove two properties of the sinusoidal position embeddings.

For an offset kk and a position tt, PEt+kTPEtPE_{t+k}^{T}PE_{t} only depends on kk, which means the dot product of two sinusoidal position embeddings can reflect the distance between two tokens.

Based on the definitions of Eq.(10) and Eq.(11), the position embedding of tt-th token is {seequation} PE_t = [ sin(c_0t)cos(c_0t)⋮sin(c_d2-1t)cos(c_d2-1t) ], where dd is the dimension of the position embedding, cic_{i} is a constant decided by ii, and its value is 1/100002i/d1/10000^{2i/d}.

where Eq.(12) to Eq.(13) is based on the equation cos⁡(x−y)=sin⁡(x)sin⁡(y)+cos⁡(x)cos⁡(y)\cos(x-y)=\sin(x)\sin(y)+\cos(x)\cos(y). ∎

For an offset kk and a position tt, PEtTPEt−k=PEtTPEt+kPE_{t}^{T}PE_{t-k}=PE_{t}^{T}PE_{t+k}, which means the sinusoidal position embeddings is unware of directionality.

Let j=t−kj=t-k, according to property 1, we have

The relation between dd, kk and PEtTPEt+kPE_{t}^{T}PE_{t+k} is displayed in Fig 3. The sinusoidal position embeddings are distance-aware but lacks directionality.

However, the property of distance-awareness also disappears when PEtPE_{t} is projected into the query and key space of self-attention. Since in vanilla Transformer the calculation between PEtPE_{t} and PEt+kPE_{t+k} is actually PEtTWqTWkPEt+kPE_{t}^{T}W_{q}^{T}W_{k}PE_{t+k}, where Wq,WkW_{q},W_{k} are parameters in Eq.(3). Mathematically, it can be viewed as PEtTWPEt+kPE_{t}^{T}WPE_{t+k} with only one parameter WW. The relation between PEtTPEt+kPE_{t}^{T}PE_{t+k} and PEtTWPEt+kPE_{t}^{T}WPE_{t+k} is depicted in Fig 4.

Therefore, to improve the Transformer with direction- and distance-aware characteristic, we calculate the attention scores using the equations below:

because sin⁡(−x)=−sin⁡(x),cos⁡(x)=cos⁡(−x)\sin(-x)=-\sin(x),\cos(x)=\cos(-x). This means for an offset tt, the forward and backward relative positional encoding are the same with respect to the cos⁡(cit)\cos(c_{i}t) terms, but is the opposite with respect to the sin⁡(cit)\sin(c_{i}t) terms. Therefore, by using Rt−jR_{t-j}, the attention score can distinguish different directions and distances.

The above improvement is based on the work (Shaw et al., 2018; Dai et al., 2019). Since the size of NER datasets is usually small, we avoid direct multiplication of two learnable parameters, because they can be represented by one learnable parameter. Therefore we do not use WkW_{k} in Eq.(17). The multi-head version is the same as Eq.(8), but we discard WoW_{o} since it is directly multiplied by W1W_{1} in Eq.(9).

2.2 Un-scaled Dot-Product Attention

The vanilla Transformer use the scaled dot-product attention to smooth the output of softmax function. In Eq.(5), the dot product of key and value matrices is divided by the scaling factor dk\sqrt{d_{k}}.

We empirically found that models perform better without the scaling factor dk\sqrt{d_{k}}. We presume this is because without the scaling factor the attention will be sharper. And the sharper attention might be beneficial in the NER task since only few words in the sentence are named entities.

3 CRF Layer

In order to take advantage of dependency between different tags, the Conditional Random Field (CRF) was used in all of our models. Given a sequence s=[s1,s2,...,sT]\mathbf{s}=[s_{1},s_{2},...,s_{T}], the corresponding golden label sequence is y=[y1,y2,...,yT]\mathbf{y}=[y_{1},y_{2},...,y_{T}], and Y(s)\mathbf{Y}(\mathbf{s}) represents all valid label sequences. The probability of y\mathbf{y} is calculated by the following equation

where f(yt−1,yt,s)f(\mathbf{y}_{t-1},\mathbf{y}_{t},\mathbf{s}) computes the transition score from yt−1\mathbf{y}_{t-1} to yt\mathbf{y}_{t} and the score for yt\mathbf{y}_{t}. The optimization target is to maximize P(y∣s)P(\mathbf{y}|\mathbf{s}). When decoding, the Viterbi Algorithm is used to find the path achieves the maximum probability.

Experiment

We evaluate our model in two English NER datasets and four Chinese NER datasets.

(1) CoNLL2003 is one of the most evaluated English NER datasets, which contains four different named entities: PERSON, LOCATION, ORGANIZATION, and MISC Sang and Meulder (2003).

(2) OntoNotes 5.0 is an English NER dataset whose corpus comes from different domains, such as telephone conversation, newswire. We exclude the New Testaments portion since there is no named entity in it Chen et al. (2019); Chiu and Nichols (2016). This dataset has eleven entity names and seven value types, like CARDINAL, MONEY, LOC.

(3) Weischedel (2011) released OntoNotes 4.0. In this paper, we use the Chinese part. We adopted the same pre-process as Che et al. (2013).

(4) The corpus of the Chinese NER dataset MSRA came from news domain Levow (2006).

(5) Weibo NER was built based on text in Chinese social media Sina Weibo Peng and Dredze (2015), and it contained 4 kinds of entities.

(6) Resume NER was annotated by Zhang and Yang (2018).

Their statistics are listed in Table 1. For all datasets, we replace all digits with “0”, and use the BIOES tag schema. For English, we use the Glove 100d pre-trained embedding Pennington et al. (2014). For the character encoder, we use 30d randomly initialized character embeddings. More details on models’ hyper-parameters can be found in the supplementary material. For Chinese, we used the character embedding and bigram embedding released by Zhang and Yang (2018). All pre-trained embeddings are finetuned during training. In order to reduce the impact of randomness, we ran all of our experiments at least three times, and its average F1 score and standard deviation are reported.

We used random-search to find the optimal hyper-parameters, hyper-parameters and their ranges are displayed in the supplemental material. We use SGD and 0.9 momentum to optimize the model. We run 100 epochs and each batch has 16 samples. During the optimization, we use the triangle learning rate Smith (2017) where the learning rate rises to the pre-set learning rate at the first 1% steps and decreases to 0 in the left 99% steps. The model achieves the highest development performance was used to evaluate the test set. The hyper-parameter search range and other settings can be found in the supplementary material. Codes are available at https://github.com/fastnlp/TENER.

2 Results on Chinese NER Datasets

We first present our results in the four Chinese NER datasets. Since Chinese NER is directly based on the characters, it is more straightforward to show the abilities of different models without considering the influence of word representation.

As shown in Table 2, the vanilla Transformer does not perform well and is worse than the BiLSTM and CNN based models. However, when relative positional encoding combined, the performance was enhanced greatly, resulting in better results than the BiLSTM and CNN in all datasets. The number of training examples of the Weibo dataset is tiny, therefore the performance of the Transformer is abysmal, which is as expected since the Transformer is data-hungry. Nevertheless, when enhanced with the relative positional encoding and unscaled attention, it can achieve even better performance than the BiLSTM-based model. The superior performance of the adapted Transformer in four datasets ranging from small datasets to big datasets depicts that the adapted Transformer is more robust to the number of training examples than the vanilla Transformer. As the last line of Table 2 depicts, the scaled attention will deteriorate the performance.

3 Results on English NER datasets

The comparison between different NER models on English NER datasets is shown in Table 3. The poor performance of the Transformer in the NER datasets was also reported by Guo et al. (2019). Although performance of the Transformer is higher than Guo et al. (2019), it still lags behind the BiLSTM-based models Ma and Hovy (2016). Nonetheless, the performance is massively enhanced by incorporating the relative positional encoding and unscaled attention into the Transformer. The adaptation not only makes the Transformer achieve superior performance than BiLSTM based models, but also unveil the new state-of-the-art performance in two NER datasets when only the Glove 100d embedding and CNN character embedding are used. The same deterioration of performance was observed when using the scaled attention. Besides, if ELMo was used Peters et al. (2018), the performance of TENER can be further boosted as depicted in Table 4.

4 Analysis of Different Character Encoders

The character-level encoder has been widely used in the English NER task to alleviate the data sparsity and OOV problem in word representation. In this section, we cross different character-level encoders (BiLSTM, CNN, Transformer encoder and our adapted Transformer encoder (AdaTrans for short) ) and different word-level encoders (BiLSTM, ID-CNN and AdaTrans) to implement the NER task. Results on CoNLL2003 and OntoNotes 5.0 are presented in Table 5a and Table 5b, respectively.

The ID-CNN encoder is from Strubell et al. (2017), and we re-implement their model in PyTorch. For different combinations, we use random search to find its best hyper-parameters. Hyper-parameters for character encoders were fixed. The details can be found in the supplementary material.

For the results on CoNLL2003 dataset which is depicted in Table 5a, the AdaTrans performs as good as the BiLSTM in different character encoder scenario averagely. In addition, from Table 5b, we can find the pattern that the AdaTrans character encoder outpaces the BiLSTM and CNN character encoders when different word-level encoders being used. Moreover, no matter what character encoder being used or none being used, the AdaTrans word-level encoder gets the best performance. This implies that when the number of training examples increases, the AdaTrans character-level and word-level encoder can better realize their ability.

5 Convergent Speed Comparison

We compare the convergent speed of BiLSTM, ID-CNN, Transformer, and TENER in the development set of the OntoNotes 5.0. The curves are shown in Fig 5. TENER converges as fast as the BiLSTM model and outperforms the vanilla Transformer.

Conclusion

In this paper, we propose TENER, a model adopting Transformer Encoder with specific customizations for the NER task. Transformer Encoder has a powerful ability to capture the long-range context. In order to make the Transformer more suitable to the NER task, we introduce the direction-aware, distance-aware and un-scaled attention. Experiments in two English NER tasks and four Chinese NER tasks show that the performance can be massively increased. Under the same pre-trained embeddings and external knowledge, our proposed modification outperforms previous models in the six datasets. Meanwhile, we also found the adapted Transformer is suitable for being used as the English character encoder, because it has the potentiality to extract intricate patterns from characters. Experiments in two English NER datasets show that the adapted Transformer character encoder performs better than BiLSTM and CNN character encoders.

References

Supplemental Material

We exploit four kinds of character encoders. For all character encoders, the randomly initialized character embeddings are 30d. The hidden size of BiLSTM used in the character encoder is 50d in each direction. The kernel size of CNN used in the character encoder is 3, and we used 30 kernels with stride 1. For Transformer and adapted Transformer, the number of heads is 3, and every head is 10d, the dropout rate is 0.15, the feed-forward dimension is 60. The Transformer used the sinusoid position embedding. The number of parameters for the character encoder (excluding character embedding) when using BiLSTM, CNN, Transformer and adapted Transformer are 35830, 3660, 8460 and 6600 respectively. For all experiments, the hyper-parameters of character encoders stay unchanged.

2 Hyper-parameters

The hyper-parameters and search ranges for different encoders are presented in Table 6, Table 7 and Table 8.