Part-of-Speech Tagging with Bidirectional Long Short-Term Memory Recurrent Neural Network

Peilu Wang, Yao Qian, Frank K. Soong, Lei He, Hai Zhao

Introduction

Bidirectional long short-term memory [Hochreiter and Schmidhuber, 1997, Schuster and Paliwal, 1997] (BLSTM) is a type of recurrent neural network (RNN) that can incorporate contextual information from long period of fore-and-aft inputs. It has been proven a powerful model for sequential labeling tasks. For applications in natural language processing (NLP), it has helped achieve superior performance in language modeling [Sundermeyer et al., 2012, Sundermeyer et al., 2015], language understanding [Yao et al., 2013], and machine translation [Sundermeyer et al., 2014]. Since part-of-speech (POS) tagging is a typical sequential labeling task, it seems natural to expect BLSTM RNN can also be effective for this task.

As a neural network model, it is awkward for BLSTM RNN to make use of conventional NLP features, such as morphological features. Since these features are discrete and has to be represented as one-hot vector to be used, using rich this type of features leads to too large input layer to maintain and update. Therefore, we avoid using such features except word form and simple capital features, instead we involve word embedding. Word embedding is a low dimensional real-valued vector used to represent word. It is considered containing part of syntactic and semantic information and has shown a very attractive feature for various of language processing tasks [Collobert and Weston, 2008, Turian et al., 2010a, Collobert et al., 2011]. Word embedding can be obtained by training a neural network model, especially, a neural network language model [Bengio et al., 2006, Mikolov et al., 2010] or a neural network designed for a specific task [Collobert et al., 2011, Mikolov et al., 2013a, Pennington et al., 2014a]. Currently many word embeddings trained on quite large corpora are available on line. However, these embeddings are trained by neural networks that are very different from BLSTM RNN. This inconsistency is supposed as an shortcoming to make the most of these trained word embeddings. To conquer this shortcoming, we also propose a novel method to train word embedding on unlabeled data with BLSTM RNN.

The main contributions of this work include: First, it shows an effective way to use BLSTM RNN for POS tagging task and achieves a state-of-the-art tagging accuracy. Second, a novel method for training word embedding is proposed. Finally, we demonstrate that competitive tagging accuracy can be obtained without using morphological features, which makes this approach more practical to tag a language that lacks of necessary morphological knowledge.

Methods

Given a sentence w1,w2,...,wnw_{1},w_{2},...,w_{n} with tags y1,y2,...,yny_{1},y_{2},...,y_{n}, BLSTM RNN is used to predict the tag probability distribution of each word. The usage is illustrated in Figure 1.

Here wi‾\underline{w_{i}} is the one hot representation of the current word. It is a binary vector of dimension ∣V∣|V| where VV is the vocabulary. To reduce ∣V∣|V|, each letter in input word is transferred into lower case. To still keep the upper case information, a function f(wi)f(w_{i}) is introduced to indicate the original case information of word wiw_{i}. More specifically, f(wi)f(w_{i}) returns a three-dimensional binary vector to tell if wiw_{i} is full lowercase, full uppercase or leading with a capital letter. The input vector IiI_{i} of the neural network is computed as:

where W1W_{1} and W2W_{2} are weight matrixes connecting two layers. W1wi‾W_{1}\underline{w_{i}} is the word embedding of wiw_{i} which has a much smaller dimension than wi‾\underline{w_{i}}. In practice, W1W_{1} is implemented as a lookup table, W1wi‾W_{1}\underline{w_{i}} is returned by referring to the word embedding of wiw_{i} stored in this table. To use word embeddings trained by other task or method, we just need to initialize this lookup table with those external embeddings. For words without corresponding external embeddings, their word embeddings are initialized with uniformly distributed random values, ranging from -0.1 to 0.1. The implementation of BLSTM layer is detailed descripted in [Graves, 2012] and therefore is skipped in this paper. This layer incorporates information from the past and future histories when making prediction for current word and is updated as a function of the entire input sentence. The output layer is a softmax layer whose dimension is the number of tag types. It outputs the tag probability distribution of input word wiw_{i}. All weights are trained using backpropagation and gradient descent algorithm to maximize the likelihood on training data:

The obtained probability distribution of each step is supposed independent with each other. The utilization of contextual information strictly comes from the BLSTM layer. Thus, in inference phase, the likeliest tag yi′y^{\prime}_{i} of input word wiw_{i} can just be chose as:

2 Word Embedding

In this section, we propose a novel method to train word embedding on unlabeled data with BLSTM RNN. In this approach, BLSTM RNN is also used to do a tagging task, but only has two types of tags to predict: incorrect/correct. The input is a sequence of words which is a normal sentence with some words replaced by randomly chosen words. For those replaced words, their tags are 0 (incorrect) and for those that are not replaced, their tags are 1 (correct). Although it is possible that some replaced words are also reasonable in the sentence, they are still considered “incorrect”. Then BLSTM RNN is trained to minimize the binary classification error on the training corpus. The neural network structure is the same as that in Figure 1. When the neural network is trained, W1W_{1} contains all trained word embeddings.

Experiments

BLSTM RNN systems in our experiments are implemented with CURRENT [Weninger et al., 2014], a machine learning library for RNN which adopts GPU acceleration. The activation function of input layer is identity function, hidden layer is logistic function, while the output layer uses softmax function for multiclass classification. Neural network is trained using statistical gradient descent algorithm with constant learning rate.

The part-of-speech tagged data used in our experiments is the Wall Street Journal data from Penn Treebank III [Marcus et al., 1993]. Training, development and test sets are split following setup in [Collins, 2002]. Table 1 lists the detailed information of the three data sets.

To train word embedding, we uses North American news [Graff, 2008] as the unlabeled data. This corpus contains about 536 million words. It is tokenized using the Penn Treebank tokenizer script https://www.cis.upenn.edu/~treebank/tokenization.html. All consecutive digits occurring within a word are replaced with the symbol “#”. For example, both words “Tel192” and “Tel6” are transferred to the same word “Tel#”.

2 Hidden Layer Size

We evaluate different sizes of hidden layer in BLSTM RNN to pick up the best structure for later experiments. The input layer size is set to 100 and output layer size is fixed as 45 in all experiments. The accuracies on WSJ test set are shown in Figure 2.

It shows that hidden layer size has a limited impact on performance when it becomes large enough. To keep a good trade-off of accuracy, model size and running time, we choose 100 which is the smallest layer size to get “reasonable” performance as the hidden layer size in all the following experiments.

3 POS Tagging Accuracies

Table 2 compares the performance of our systems with other baseline systems.

Baseline systems. Four typical systems are chosen as baseline systems. [Toutanova et al., 2003] is one of the most commonly used approaches which is also known as Stanford tagger. [Huang et al., 2012] is the system reports best accuracy on WSJ test set (97.35%). In fact, [Spoustová et al., 2009] reports a higher accuracy (97.44%), but this work relies on multiple trained taggers and combines their tagging results. Here we focus on single model tagging algorithm and therefore do not include this work as baseline. Besides, [Moore, 2014] (97.34%) and [Shen et al., 2007] (97.33%) also reach accuracy above 97.3%. These two systems plus [Huang et al., 2012] are considered as current state-of-the-art systems. All these systems rely on rich morphological features. In contrast, [Collobert et al., 2011] NN only uses word form and capital features. [Collobert et al., 2011] NN+WE also incorporates word embeddings trained on unlabeled data like our approach. The main difference is that [Collobert et al., 2011] uses feedforward neural network instead of BLSTM RNN.

BLSTM-RNN is the system described in Section 2.1 which only uses word form and capital features. The vocabulary we used in this experiment is all words appearing in WSJ Penn Treebank training set, merging with the most common 100,000 words in North American news corpus, plus one single “UNK” symbol for replacing all out of vocabulary words.

Without the help of morphological features, it is not surprising that BLSTM-RNN falls behind the state-of-the-art system. However, BLSTM-RNN surpasses [Collobert et al., 2011] NN which is also neural network based method and uses the same input features. It is consistent with [Fernandez et al., 2014, Fan et al., 2014], in which BLSTM RNN outperforms feedforward neural network.

BLSTM-RNN+WE. To construct corpus for training word embeddings, about 20% words in normal sentences of North American news corpus are replaced with randomly selected words. Then BLSTM RNN is trained to judge which word has been replaced as described in Section 2.2. The vocabulary for this task contains the 100,000 most common words in North American news corpus and one special “UNK” symbol. When training is finished, word embedding lookup table (W1W_{1}) in BLSTM RNN for POS tagging is initialized with the trained word embeddings. The following training and testing are the same as previous experiment.

Table 2 shows the results of using word embeddings trained on the first 10 million words (WE(10m)), first 100 million words (WE(100m)) and all 530 million words (WE(all)) of North American news corpus. While WE(10m) does not show much help for the improvement, WE(100m) and WE(all) significantly boosts the performance. It shows that BLSTM RNN can benefit from word embeddings trained on large unlabeled corpus and larger training corpus leads to a better performance. This suggests that the result may be further improved by using even bigger unlabeled data set. With the help of GPU, WE(all) can be trained in about one day (23 hrs). The training time increases linearly with the training corpus size.

WE(all) reduces over 20% error rate of BLSTM-RNN and lets the result comparable with [Toutanova et al., 2003]. Note that this result is obtained without using any morphological features. Current state-of-the-art systems [Moore, 2014, Shen et al., 2007, Huang et al., 2012] all utilize morphological features proposed in [Ratnaparkhi, 1996] which involves nn-gram prefix and suffix (nn = 1 to 4). Moreover, [Shen et al., 2007] also involves prefix and suffix of length from 5 to 9. [Moore, 2014] adds extra elaborately designed features, including flags indicating if word ends with −ed-ed or −ing-ing, etc. In practice, many languages with rich morphological forms lack of necessary or effective morphological processing tools. In these cases, a POS tagger that does not rely on morphological features is more realistic for use.

BLSTM-RNN+WE(all)+suffix2. In this experiment, we add bigram suffix of each word as extra feature. These last 2 characters are represented as one-hot vector and appended to the original extra feature vector (f(wi)f(w_{i})). The other configuration follows BLSTM-RNN+WE(all). The additional feature furthermore pushes up the accuracy and lets the approach get the state-of-the-art performance (97.40%). However, adding more morphological features such as trigram suffix does not further improve the performance. One possible reason is that adding such feature brings a much longer extra feature vector which needs retuning parameters such as learning rate and hidden layer size to get the optimum performance.

4 Different Word Embeddings

In this experiment, six types of published well-trained word embeddings are evaluated. The basic information of involved word embeddings and results are listed in Table 3 where RCV1 represents the Reuters Corpus Volume 1 news set. The OOV (out of vocabulary) column indicates the rate of words in vocabulary of BLSTM RNN for POS tagging that are not covered by external word embedding vocabulary. The usage of word embeddings is the same as in BLSTM-RNN+WE experiment except that input layer size here is equal to the dimension of external word embedding.

All word embeddings bring about higher accuracy. However, none of them can enhance BLSTM RNN tagging to get a competitive accuracy, despite of larger corpora that they are trained on and lower OOV rate. [Pennington et al., 2014b]1 (97.12%) has the highest accuracy among them but it is still lower than [Toutanova et al., 2003] (97.24%). Although more experiments are needed to judge which word embeddings are better, this experiment at least shows word embeddings trained by BLSTM RNN are essential in our POS tagging approach to achieve a superior performance.

Conclusions

In this paper, BLSTM RNN is proposed for POS tagging and training word embedding. Combined with word embedding trained on big unlabeled data, this approach gets state-of-the-art accuracy on WSJ test set without using rich morphological features. BLSTM RNN with word embedding is expected as an effective solution for tagging tasks and worth further exploration.

References