Non-distributional Word Vector Representations
Manaal Faruqui, Chris Dyer
Introduction
Distributed representations of words have been shown to benefit a diverse set of NLP tasks including syntactic parsing [Lazaridou et al., 2013, Bansal et al., 2014], named entity recognition [Guo et al., 2014] and sentiment analysis [Socher et al., 2013]. Additionally, because they can be induced directly from unannotated corpora, they are likewise available in domains and languages where traditional linguistic resources do not exhaust. Intrinsic evaluations on various tasks are helping refine vector learning methods to discover representations that captures many facts about lexical semantics [Turney, 2001, Turney and Pantel, 2010].
Yet induced word vectors do not look anything like the representations described in most lexical semantic theories, which focus on identifying classes of words [Levin, 1993, Baker et al., 1998, Schuler, 2005, Miller, 1995]. Though expensive to construct, conceptualizing word meanings symbolically is important for theoretical understanding and interpretability is desired in computational models.
Our contribution to this discussion is a new technique that constructs task-independent word vector representations using linguistic knowledge derived from pre-constructed linguistic resources like WordNet [Miller, 1995], FrameNet [Baker et al., 1998], Penn Treebank [Marcus et al., 1993] etc. In such word vectors every dimension is a linguistic feature and 1/0 indicates the presence or absence of that feature in a word, thus the vector representations are binary while being highly sparse (). Since these vectors do not encode any word cooccurrence information, they are non-distributional. An additional benefit of constructing such vectors is that they are fully interpretable i.e, every dimension of these vectors maps to a linguistic feature unlike distributional word vectors where the vector dimensions have no interpretability.
Of course, engineering feature vectors from linguistic resources is established practice in many applications of discriminative learning; e.g., parsing [McDonald and Pereira, 2006, Nivre, 2008] or part of speech tagging [Ratnaparkhi, 1996, Collins, 2002]. However, despite a certain common inventories of features that re-appear across many tasks, feature engineering tends to be seen as a task-specific problem, and engineered feature vectors are not typically evaluated independently of the tasks they are designed for. We evaluate the quality of our linguistic vectors on a number of tasks that have been proposed for evaluating distributional word vectors. We show that linguistic word vectors are comparable to current state-of-the-art distributional word vectors trained on billions of words as evaluated on a battery of semantic and syntactic evaluation benchmarks.Our vectors can be downloaded at: https://github.com/mfaruqui/non-distributional
Linguistic Word Vectors
We construct linguistic word vectors by extracting word level information from linguistic resources. Table 1 shows the size of vocabulary and number of features induced from every lexicon. We now describe various linguistic resources that we use for constructing linguistic word vectors.
WordNet [Miller, 1995] is an English lexical database that groups words into sets of synonyms called synsets and records a number of relations among these synsets or their members. For a word we look up its synset for all possible part of speech (POS) tags that it can assume. For example, film will have Synset.Film.V.01 and Synset.Film.N.01 as features as it can be both a verb and a noun. In addition to synsets, we include the hyponym (for ex. Hypo.CollageFilm.N.01), hypernym (for ex. Hyper:Sheet.N.06) and holonym synset of the word as features. We also collect antonyms and pertainyms of all the words in a synset and include those as features in the linguistic vector.
Supsersenses.
WordNet partitions nouns and verbs into semantic field categories known as supsersenses [Ciaramita and Altun, 2006, Nastase, 2008]. For example, lioness evokes the supersense SS.Noun.Animal. These supersenses were further extended to adjectives [Tsvetkov et al., 2014].http://www.cs.cmu.edu/~ytsvetko/adj-supersenses.tar.gz We use these supsersense tags for nouns, verbs and adjectives as features in the linguistic word vectors.
FrameNet.
FrameNet [Baker et al., 1998, Fillmore et al., 2003] is a rich linguistic resource that contains information about lexical and predicate-argument semantics in English. Frames can be realized on the surface by many different word types, which suggests that the word types evoking the same frame should be semantically related. For every word, we use the frame it evokes along with the roles of the evoked frame as its features. Since, information in FrameNet is part of speech (POS) disambiguated, we couple these feature with the corresponding POS tag of the word. For example, since appreciate is a verb, it will have the following features: Verb.Frame.Regard, Verb.Frame.Role.Evaluee etc.
Emotion & Sentiment.
?) constructed two different lexicons that associate words to sentiment polarity and to emotions resp. using crowdsourcing. The polarity is either positive or negative but there are eight different kinds of emotions like anger, anticipation, joy etc. Every word in the lexicon is associated with these properties. For example, cannibal evokes Pol.Neg, Emo.Disgust and Emo.Fear. We use these properties as features in linguistic vectors.
Connotation.
?) construct a lexicon that contains information about connotation of words that are seemingly objective but often allude nuanced sentiment. They assign positive, negative and neutral connotations to these words. This lexicon differs from ?) in that it has a more subtle shade of sentiment and it extends to many more words. For example, delay has a negative connotation Con.Noun.Neg, floral has a positive connotation Con.Adj.Pos and outline has a neutral connotation Con.Verb.Neut.
Color.
Most languages have expressions involving color, for example green with envy and grey with uncertainly are phrases used in English. The word-color associtation lexicon produced by ?) using crowdsourcing lists the colors that a word evokes in English. We use every color in this lexicon as a feature in the vector. For example, Color.Red is a feature evoked by the word blood.
Part of Speech Tags.
The Penn Treebank [Marcus et al., 1993] annotates naturally occurring text for linguistic structure. It contains syntactic parse trees and POS tags for every word in the corpus. We collect all the possible POS tags that a word is annotated with and use it as features in the linguistic vectors. For example, love has PTB.Noun, PTB.Verb as features.
Synonymy & Antonymy.
We use Roget’s thesaurus [Roget, 1852] to collect sets of synonymous words.http://www.gutenberg.org/ebooks/10681 For every word, its synonymous word is used as a feature in the linguistic vector. For example, adoration and affair have a feature Syno.Love, admissible has a feature Syno.Acceptable. The synonym lexicon contains 25,338 words after removal of multiword phrases. In a similar manner, we also use antonymy relations between words as features in the word vector. The antonymous words for a given word were collected from ?).https://archive.org/details/synonymsantonyms00ordwiala An example would be of impartiality, which has features Anto.Favoritism and Anto.Injustice. The antonym lexicon has 10,355 words. These features are different from those induced from WordNet as the former encode word-word relations whereas the latter encode word-synset relations.
After collecting features from the various linguistic resources described above we obtain linguistic word vectors of length 172,418 dimensions. These vectors are 99.9% sparse i.e, each vector on an average contains only 34 non-zero features out of 172,418 total features. On average a linguistic feature (vector dimension) is active for 15 word types. The linguistic word vectors contain 119,257 unique word types. Table 2 shows linguistic vectors for some of the words.
Experiments
We first briefly describe the evaluation tasks and then present results.
We evaluate our word representations on three different benchmarks to measure word similarity. The first one is the widely used WS-353 dataset [Finkelstein et al., 2001], which contains 353 pairs of English words that have been assigned similarity ratings by humans. The second is the RG-65 dataset [Rubenstein and Goodenough, 1965] of 65 words pairs. The third dataset is SimLex [Hill et al., 2014] which has been constructed to overcome the shortcomings of WS-353 and contains 999 pairs of adjectives, nouns and verbs. Word similarity is computed using cosine similarity between two words and Spearman’s rank correlation is reported between the rankings produced by vector model against the human rankings.
Sentiment Analysis.
NP-Bracketing.
2 Linguistic Vs. Distributional Vectors
We compare both sparse and dense linguistic vectors to three widely used distributional word vector models. The first two are the pre-trained Skip-Gram [Mikolov et al., 2013]https://code.google.com/p/word2vec and Glove [Pennington et al., 2014]http://www-nlp.stanford.edu/projects/glove/ word vectors each of length 300, trained on 300 billion and 6 billion words respectively. We used latent semantic analysis (LSA) to obtain word vectors from the SVD decomposition of a word-word cooccurrence matrix [Turney and Pantel, 2010]. These were trained on 1 billion words of Wikipedia with vector length 300 and context window of 5 words.
3 Results
Table 3 shows the performance of different word vector types on the evaluation tasks. It can be seen that although Skip-Gram, Glove & LSA perform better than linguistic vectors on WS-353, the linguistic vectors outperform them by a huge margin on SimLex. Linguistic vectors also perform better at RG-65. On sentiment analysis, linguistic vectors are competitive with Skip-Gram vectors and on the NP-bracketing task they outperform all distributional vectors with a statistically significant margin (p 0.05, McNemar’s test ?)). We append the sparse linguistic vectors to Skip-Gram vectors and evaluate the resultant vectors as shown in the bottom row of Table 3. The combined vector outperforms Skip-Gram on all tasks, showing that linguistic vectors contain useful information orthogonal to distributional information.
It is evident from the results that linguistic vectors are either competitive or better to state-of-the-art distributional vector models. Sparse linguistic word vectors are high dimensional but they are also sparse, which makes them computationally easy to work with.
Discussion
Linguistic resources like WordNet have found extensive applications in lexical semantics, for example, for word sense disambiguation, word similarity etc. [Resnik, 1995, Agirre et al., 2009]. Recently there has been interest in using linguistic resources to enrich word vector representations. In these approaches, relational information among words obtained from WordNet, Freebase etc. is used as a constraint to encourage words with similar properties in lexical ontologies to have similar word vectors [Xu et al., 2014, Yu and Dredze, 2014, Bian et al., 2014, Fried and Duh, 2014, Faruqui et al., 2015a]. Distributional representations have also been shown to improve by using experiential data in addition to distributional context [Andrews et al., 2009]. We have shown that simple vector concatenation can likewise be used to improve representations (further confirming the established finding that lexical resources and cooccurrence information provide somewhat orthogonal information), but it is certain that more careful combination strategies can be used.
Although distributional word vector dimensions cannot, in general, be identified with linguistic properties, it has been shown that some vector construction strategies yield dimensions that are relatively more interpretable [Murphy et al., 2012, Fyshe et al., 2014, Fyshe et al., 2015, Faruqui et al., 2015b]. However, such analysis is difficult to generalize across models of representation. In constrast to distributional word vectors, linguistic word vectors have interpretable dimensions as every dimension is a linguistic property.
Linguistic word vectors require no training as there are no parameters to be optimized, meaning they are computationally economical. While good quality linguistic word vectors may only be obtained for languages with rich linguistic resources, such resources do exist in many languages and should not be disregarded.
Conclusion
We have presented a novel method of constructing word vector representations solely using linguistic knowledge from pre-existing linguistic resources. These non-distributional, linguistic word vectors are competitive to the current models of distributional word vectors as evaluated on a battery of tasks. Linguistic vectors are fully interpretable as every dimension is a linguistic feature and are highly sparse, so they are computationally easy to work with.
Acknowledgement
We thank Nathan Schneider for giving comments on an earlier draft of this paper and the anonymous reviewers for their feedback.