An Attentive Neural Architecture for Fine-grained Entity Type Classification
Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel
Introduction
Entity type classification is the task of assigning semantic types to mentions of entities in sentences. Identifying the types of entities is useful for various natural language processing tasks, such as relation extraction [Ling and Weld, 2012], question answering [Lee et al., 2006], and knowledge base population [Carlson et al., 2010]. Unfortunately, most entity type classification systems use a relatively small number of types (e.g. person, organization, location, time, and miscellaneous [Grishman and Sundheim, 1996]) which may be too coarse-grained for some NLP applications [Sekine, 2008]. To address this shortcoming, a series of recent work has investigated entity type classification with a large set of fine-grained types [Lee et al., 2006, Ling and Weld, 2012, Yosef et al., 2012, Yogatama et al., 2015, Del Corro et al., 2015].
Existing fine-grained entity type classification systems have used approaches ranging from sparse binary features to dense vector representations of entities to model the entity mention and its context. However, no previously proposed system has attempted to learn to recursively compose representations of entity context. For example, one can see that a phrase “got a Ph.D. from” is indicative of the next words being an educational institution, something which would be helpful for fine-grained entity type classification.
In this work our main contributions are two-fold:
A first model for fine-grained entity type classification that learns to recursively compose representations for the context of each mention and attains state-of-the-art performance on a well-established dataset.
The observation that by incorporating an attention mechanism into our model, we not only achieve better performance, but also are able to observe that the model learns contextual linguistic expressions that indicate fine-grained category memberships of an entity.
Related Work
To the best of our knowledge, ?) were the first to address the task of fine-grained entity type classification. They defined 147 fine-grained entity types and evaluated a conditional random fields-based model on a manually annotated Korean dataset. ?) advocated the necessity of a large set of types for entity type classification and defined types which served as a basis for future work on fine-grained entity type classification.
?) defined a set of types based on Freebase and created a training dataset from Wikipedia using a distant supervision method inspired by ?). For evaluation, they created a small manually annotated dataset of newspaper articles and also demonstrated that their system, FIGER, could improve the performance of a relation extraction system by providing fine-grained entity type predictions as features. ?) organised types in a hierarchical taxonomy, with several hundreds of types at different levels. Based on this taxonomy they developed a multi-label hierarchical classification system. In ?) the authors proposed to use label embeddings to allow information sharing between related labels. This approach lead to improvements on the FIGER dataset, and they also demonstrated that fine-grained labels can be used as features to improve coarse-grained entity type classification performance. ?) introduced the most fine-grained entity type classification system to-date, it operates on the the entire WordNet hierarchy with more than types.
While all previous models relied on hand-crafted features, ?) defined types and created a two-part neural classifier. They used a recurrent neural networks to recursively obtain a vector representation of each entity mention and used a fixed-size window to capture the context of each mention. The key difference between our work and theirs lies in that we use recursive neural networks to compose context representations and that we employ an attention mechanism to allow our model to focus on relevant expressions.
Models
At inference, the type is predicted if is greater than or is the maximum value . The motivation of the former is that it acts as a cut-off, while the latter enforces the constraint that each mention is assigned at least one type.
2 General Model
Note that we did not include a bias term in the above formulation since the type distribution in the training and test corpus could potentially be significantly different due to domain differences. That is, in logistic regression, a bias fits to the empirical distribution of types in the training set, which would lead to bad performance on a test set that has a different type distribution.
The loss for a prediction when the true labels are encoded in a binary vector is the following cross entropy loss function:
3 Mention Representation
During our experiments we were surprised by the fact that unlike the observations made by ?), complex neural models did not work well for learning mention representations compared to the simpler model described above. One possible explanation for this would be labeling discrepancies between the training and test set. For example, the label time is assigned to days of the week (e.g. “Friday”, “Monday”, and “Sunday”) in the test set, but not in the training set, whereas explicit dates (e.g. “Feb. 24” and “June 4th”) are assigned the time label in both the training and test set. This may be harmful for complex models due to their tendency to overfit on the training data.
4 Context Representation
We compare three methods for computing context representations.
Applying the same averaging approach as for the mention representation for both the left and right context. Thus, the concatenation of those two vectors becomes the representation of the context:
4.2 LSTM Encoder
For the left context, the model reads sequences from left to right to produce the outputs . For the right context, the model reads sequences from right to left to produce the outputs . Then the representation is obtained by concatenating and :
A more detailed formulation of the LSTM used in this work can be found in ?).
4.3 Attentive Encoder
While an LSTM can encode sequential data, it still finds it difficult to learn long-term dependencies. Inspired by recent work using attention mechanisms for natural language processing [Hermann et al., 2015, Rocktäschel et al., 2015], we circumvent this problem by introducing a novel attention mechanism. We also hypothesize that by incorporating an attention mechanism the model can recognize informative expressions for the classification and make the model behavior more interpretable.
The computation of the attention mechanism is as follows. Firstly, for both the right and left context, we encode the sequences using bi-directional LSTMs [Graves, 2012]. We denote the outputs as and .
Experiment
To train and evaluate our model we use the publicly available FIGER dataset with 112 fine-grained types from ?). The sizes of our datasets are for training, for development, and for testing. Note that the train and development sets were created from Wikipedia, whereas the test set is a manually annotated dataset of newspaper articles.
2 Pre-trained Word Embeddings
The only features used by our model are pre-trained word embeddings that were not updated during training to help the model generalize for words not appearing in the training set. Specifically, we used the freely available dimensional cased word embeddings trained on 840 billion tokens from the Common Crawl supplied by ?). As embeddings for out-of-vocabulary words, we used the embedding of the “unk” token from the pre-trained embeddings.
3 Evaluation Criteria
Following ?), we evaluate the model performances by strict, loose macro, and loose micro measures. For the -th instance, let the set of the predicted types be , and the set of the true types be . Then the precisions and recall for each measure are computed as follows.
Where is the total number of instances.
4 Hyperparameter Settings
As hyperparameters, all three models used the same dimensional word embeddings, the hidden-size of the LSTM was set to , and the hidden-layer size of the attention module was set to . We used Adam [Kingma and Ba, 2014] as our optimization method with a learning rate of with a mini-batch size of . As a regularizer we used dropout with probability applied to the mention representation.
The context window size was set to and mention window size was set to . It should be noted that our approach is not restricted to using fixed window sizes, rather this is an implementation detail arising from current limitations of the machine learning library used when handling dynamic-width recurrent neural networks. For each epoch we iterated over the training data set ten times and then evaluated the model performance on the development set. After training we picked up the best model on the development set as our final model and report the performance on the test set. Our model implementation was done in Python using the TensorFlow [Abadi et al., 2015] machine learning library.
5 Results
The performance of the various models are summarized Tables 1 and 2. We see that the Averaging base line performs well in spite of its relative simplicity, the LSTM model shows some improvements, and the attention model performs better than any previously proposed method. In Figure 2, we visualize the attentions for several instances that were manually selected from the development set. It is clear that our proposed model is attending over expressions relevant for the entity types such as immediately adjacent to the mention such as “starring” and “Republican Governor”, as well as more distant expressions such as “filmmakers”.
Conclusion
In this paper, we proposed a novel state-of-the-art neural network architecture with an attention mechanism for the task of fine-grained entity type classification. We also demonstrated that the model can successfully learn to attend over expressions that are important for the classification of fine-grained types.
Acknowledgments
This work was supported by CREST-JST, JSPS KAKENHI Grant Number 15H01702, a Marie Curie Career Integration Award, and an Allen Distinguished Investigator Award. We would like to thank the anonymous reviewers and Koji Matsuda for their helpful comments and feedback.