Finer Grained Entity Typing with TypeNet
Shikhar Murty, Patrick Verga, Luke Vilnis, Andrew McCallum
Introduction
Recognizing entities and their types is a core problem in natural language processing, underlying complex natural language understanding problems in relation extraction (Yaghoobzadeh et al., 2017), knowledge base construction, question answering (Lee et al., 2006), and query comprehension (Dalton et al., 2014). Early attempts at entity recognition focused only on very coarse grained types (Tjong Kim Sang and De Meulder, 2003; Hovy et al., 2006). More recently, there has been growing interest in models explicitly focused on entity typing with finer grained typesets e.g. FIGER (Ling and Weld, 2012).
The increasingly sophisticated natural language understanding tasks undertaken by the machine learning community often require commensurately more sophisticated world knowledge. This world knowledge, often organized hierarchically in ontologies, motivates our creation of a new fine-grained, deep, and high-quality dataset of hierarchical types.
Despite the increasing focus on fine-grained typing, existing typesets still contain only on the order of 100 different types. Further, these typesets are either endowed with only a shallow hierarchy, typically on the order of two levels deep or don’t have links to existing KBs (see Table 1). In this work, we advocate for larger, deeper typesets and models that exploit the inherently hierarchical nature of these types. To this end, we present TypeNet, an expert-annotated type hierarchy containing 1941 individual types, with an average depth of 7.8.
We also evaluate several models for fine-grained entity typing, and establish a strong baseline of 74.8 MAP on the CoNLL-YAGO dataset (Hoffart et al., 2011) for our best model. With each entity having on the order of 30 types, there are clearly exciting opportunities for improvement from future research. Additionally, we investigate multi-task models that explicitly incorporate the hierarchical relations between types into the learning objective.
Dataset Creation
We now discuss TypeNethttps://github.com/iesl/TypeNet, a new dataset of entity types for extremely fine grained entity typing. TypeNet was created by manually aligning Freebase types to noun synsets from the WordNet hierarchy (Fellbaum, 1998), naturally producing a hierarchical type set.
This was done by first filtering out all Freebase types that were linked to 20 entities, and then filtering Freebase API types. The Freebase API types we filtered were those in the domain "/freebase","/dataworld","/schema", "/atom", "/scheme" and "/topics".
For each Freebase type in our filtered set, we generate a list of candidate WordNet synsets through a substring match. The annotator then attempted to map the Freebase type to one or more synset in the candidate list with a parent-of, child-of or equivalence link by examining definitions of each synset and example entities of the Freebase type. If no match was found, the annotator queried the online WordNet API until an appropriate synset was found.
This procedure was carried out by two separate annotators independently after which conflicts were discussed and resolved. The annotators were conservative with assigning equivalence links resulting in a greater number of child-of links. The final dataset contained 13 parent-of, 727 child-of, and 380 equivalence links. Note that some Freebase types have multiple child-of links to WordNet. Finally, all the ancestors of the Freeebase types (following the manually created child-of and WordNet hypernym links) were added to construct the dataset.
We also carried out a procedure to add an additional set of 614 fb fb links. This was done by computing conditional probabilities of freebase types given other freebase types from a collection of 5 million randomly chosen freebase entities. We then threshold these probabilities to 0.7, and manually filter the resulting links.
Model
We now describe the various neural models we use for our experiments.
Input Layer: We represent a mention as a sequence of word vectors where each vector is of a fixed dimension . To obtain a mention vector representation, we use a Convolutional Neural Network (CNN) based architecture. The CNN learns mention representations from sliding w-gram features of the mention. For a mention with tokens represented as vectors, the CNN outputs vectors, which are then max-pooled along every dimension to obtain :
where is a CNN filter of size , and is a bias vector of size . We then concatenate with another vector obtained by averaging the surface form of the entity to which the mention links. We do this to provide our model explicit signal about the entity present in the mention. The concatenation is then passed through a series of affine, ReLU and affine transforms to obtain the final mention representation (See Fig-2):
Loss Function: Like prior work (Shimaoka et al., 2017), we model entity typing as a multi-label problem, and for a given mention, produce a vector of scores corresponding to each type. We optimize a mention-typing loss over each minibatch of pairs, where is the set of gold types annotated for the mention :
where is some function indicating the compatibility score for mention being in type , using the interpretation of a type as a set defined by a unary predicate (“has type t”).
In this work, we also introduce a structure loss among the types to incorporate the hierarchy. For this, we have a separate minibatch of pairs, where is the set of ancestor types for the type :
The exact scoring functions used for different models are summarized in Table 2, but are either variations of binary cross entropy or order embedding loss. We experiment with models whose loss is either alone, or a weighted combination of and .
Hyperparameters: We use pretrained 300 dimensional case sensitive GloVe vectors by Pennington et al. (2014) and a CNN with a filter width of 5. The type vectors are all 300 dimensional initialized using Glorot initialization (Glorot et al., 2011). We use dropout (Srivastava et al., 2014) as described in Fig-2 and optimize using Adam (Kingma and Ba, 2014). We tune our hyperparameters via grid search and early stopping on the development set.
Results
Dataset and Evaluation metrics: To perform our experiments, we use the CoNLL-YAGO (Hoffart et al., 2011) dev/test split and a subsampled version of Wikipedia (2016/09/20 dump) for training. We obtain labels for mentions via distant supervision, by assuming as positive types all the TypeNet types of the entity linked to a mention. We do not perform any heuristic pruning/denoising of these types, even though not all of them are relevant for a mention. This is because most pruning methods in literature are either harsh for extremely fine types (Gillick et al., 2014), or did not give an increase in performance (Shimaoka et al., 2017). However, we plan on improving this in future work and release a gold test data set.
To obtain the TypeNet types of an entity, we filter out all its Freebase types present in TypeNet, and finally for every type filtered out, add all its ancestors from TypeNet, giving us an average of 30.73 types per mention, much greater than earlier datasets such as FIGER (GOLD) (Ling and Weld, 2012) for which the average was 1.73 types per mention.
Since we have on average 30x more types per entities, we use Mean Average Precision (MAP) to measure performance unlike prior work on fine grained entity typing (Shimaoka et al., 2017). Our results are summarized in Table 3.
Discussion: We observe that the CNN encoder model works best, and multitasking mention typing with structure decreases performance if the structure is modeled using a dot product. This is expected since hypernymy is an asymmetric relation. Furthermore, modeling structure with a bilinear objective improves performance over a dot product objective but fails to perform better than the regular model. We believe this is due to the nature of our test data which is made up mostly of leaf type predictions from a non-diverse set of types (people, locations, and organizations). A more diverse dataset requiring predictions at different depths of the hierarchy could use this structure more effectively, since we posit that without using structure, the model will incorrectly predict leaf types when only parent types are true.
Interestingly, we observe the order embeddings model from Vendrov et al. (2016) to have a poor performance for our task. We attribute this to the fact that the loss function is poorly suited to the problem since it uses unrelated concepts as negative examples, which in the traditional order embedding model actually implies a reversal of the parent-child relation, rather than simply forcing the types to be unrelated. For example, consider a hypernym link person organism, and a negative example, person stadium. The loss function from Vendrov et al. (2016) attempts to increase the order violation between person and stadium, making stadium a hyponym of person. We also observe particularly poor performance combining the order embeddings with the CNN encoder.
Related work
Table-1 summarizes existing hierarchical type systems including popular data sets such as FIGER (Ling and Weld, 2012) and Gillick et al. (2014).
Del Corro et al. (2015) considered the task of extremely fine grained entity typing. They use manually crafted rules and patterns (Hearst patters, appositives, etc). Hearst (1992) to extract candidate entity types that match Wordnet synsets. They apply an optional KB type filtering step for entity by matching a candidate type to any of ’s KB types if is a string match of any , , or . We instead manually annotated the exact mapping from 1081 Freebase types to the specific WordNet synset sense allowing us to leverage distant supervision to trained supervised classifiers.
The knowledge base Yago (Suchanek et al., 2008) includes integration with WordNet and type hierarchies have been derived from its type system (Yosef et al., 2012). However the links between entity types and WordNet types are performed heuristically whereas TypeNet contains gold links between Freebase and WordNet.
There has been a growing interest in learning representations of hierarchically organized objects. Vilnis and McCallum (2016) proposed Gaussian embeddings which learn containment properties of words by approximating them with Gaussian distributions. Vendrov et al. (2016) introduced order embeddings by minimizing an order violation loss. Recently Nickel and Kiela (2017) proposed Poincaré embeddings.
Conclusion and Future Work
We introduced TypeNet, a human labeled alignment between Freebase entity types and WordNet synsets. We used this typeset to distantly label the CoNLL-YAGO entity linking dataset and reported initial results with several models comparable to the state-of-the-art models previously used on pre-existing datasets e.g. Shimaoka et al. (2017).
We additionally present results from models incorporating a structure loss over the type hierarchy, which does not appear to be required by the CoNLL-YAGO dataset, but should be helpful on more diverse datasets from different domains e.g. ClueWeb, which we will explore in future work and encourage the community to do the same.
We are exploring more sophisticated methods of incorporating the type hierarchy into the typing loss, as well as joint models for related tasks such as simultaneous typing and entity linking.
We are excited to see what the community will do with TypeNet, the largest and deepest entity type hierarchy with manual alignment to Freebase. We hope this will spur improvements in fine-grained entity and mention typing, linking and associated downstream tasks.