Morph-fitting: Fine-Tuning Word Vector Spaces with Simple Language-Specific Rules
Ivan Vulić, Nikola Mrkšić, Roi Reichart, Diarmuid Ó Séaghdha, Steve Young, Anna Korhonen
Introduction
Word representation learning has become a research area of central importance in natural language processing (NLP), with its usefulness demonstrated across many application areas such as parsing Chen and Manning (2014); Johannsen et al. (2015), machine translation Zou et al. (2013), and many others Turian et al. (2010); Collobert et al. (2011). Most prominent word representation techniques are grounded in the distributional hypothesis Harris (1954), relying on word co-occurrence information in large textual corpora (Curran, 2004; Turney and Pantel, 2010; Mikolov et al., 2013; Mnih and Kavukcuoglu, 2013; Levy and Goldberg, 2014; Schwartz et al., 2015, i.a.).
Morphologically rich languages, in which “substantial grammatical information…is expressed at word level” Tsarfaty et al. (2010), pose specific challenges for NLP. This is not always considered when techniques are evaluated on languages such as English or Chinese, which do not have rich morphology. In the case of distributional vector space models, morphological complexity brings two challenges to the fore:
1. Estimating Rare Words: A single lemma can have many different surface realisations. Naively treating each realisation as a separate word leads to sparsity problems and a failure to exploit their shared semantics. On the other hand, lemmatising the entire corpus can obfuscate the differences that exist between different word forms even though they share some aspects of meaning.
2. Embedded Semantics: Morphology can encode semantic relations such as antonymy (e.g. literate and illiterate, expensive and inexpensive) or (near-)synonymy (north, northern, northerly).
We demonstrate the efficacy of morph-fitting in four languages (English, German, Italian, Russian), yielding large and consistent improvements on benchmarking word similarity evaluation sets such as SimLex-999 Hill et al. (2015), its multilingual extension Leviant and Reichart (2015), and SimVerb-3500 Gerz et al. (2016). The improvements are reported for all four languages, and with a variety of input distributional spaces, verifying the robustness of the approach.
We then show that incorporating morph-fitted vectors into a state-of-the-art neural-network DST model results in improved tracking performance, especially for morphologically rich languages. We report an improvement of 4% on Italian, and 6% on German when using morph-fitted vectors instead of the distributional ones, setting a new state-of-the-art DST performance for the two datasets.There are no readily available DST datasets for Russian.
Preliminaries In this work, we focus on four languages with varying levels of morphological complexity: English (en), German (de), Italian (it), and Russian (ru). These correspond to languages in the Multilingual SimLex-999 dataset. Vocabularies , , , are compiled by retaining all word forms from the four Wikipedias with word frequency over 10, see Tab. 3. We then extract sets of linguistic constraints from these (large) vocabularies using a set of simple language-specific if-then-else rules, see Tab. 2.A native speaker can easily come up with these sets of morphological rules (or at least with a reasonable subset of them) without any linguistic training. What is more, the rules for de, it, and ru were created by non-native, non-fluent speakers with a limited knowledge of the three languages, exemplifying the simplicity and portability of the approach. These constraints (Sect. 2.2) are used as input for the vector space post-processing Attract-Repel algorithm (outlined in Sect. 2.1).
The Attract-Repel model, proposed by Mrkšić et al. (2017b), is an extension of the Paragram procedure proposed by Wieting et al. (2015). It provides a generic framework for incorporating similarity (e.g. successful and accomplished) and antonymy constraints (e.g. nimble and clumsy) into pre-trained word vectors. Given the initial vector space and collections of Attract and Repel constraints and , the model gradually modifies the space to bring the designated word vectors closer together or further apart. The method’s cost function consists of three terms. The first term pulls the Attract examples closer together. If denotes the current mini-batch of Attract examples, this term can be expressed as:
where is the similarity margin which determines how much closer synonymous vectors should be to each other than to each of their respective negative examples. is the standard rectified linear unit Nair and Hinton (2010). The ‘negative’ example for each word in any Attract pair is the word vector closest to among the examples in the current mini-batch (distinct from its target synonym and itself). This means that this term forces synonymous words from the in-batch Attract constraints to be closer to one another than to any other word in the current mini-batch.
The second term pushes antonyms away from each other. If is the current mini-batch of Repel constraints, this term can be expressed as follows:
In this case, each word’s ‘negative’ example is the (in-batch) word vector furthest away from it (and distinct from the word’s target antonym). The intuition is that we want antonymous words from the input Repel constraints to be further away from each other than from any other word in the current mini-batch; is now the repel margin.
The final term of the cost function serves to retain the abundance of semantic information encoded in the starting distributional space. If is the initial distributional vector and is the set of all vectors present in the given mini-batch, this term (per mini-batch) is expressed as follows:
where is the L2 regularisation constant.We use hyperparameter values , , from prior work without fine-tuning. We train all models for 10 epochs with AdaGrad Duchi et al. (2011). This term effectively pulls word vectors towards their initial (distributional) values, ensuring that relations encoded in initial vectors persist as long as they do not contradict the newly injected ones.
2 Language-Specific Rules and Constraints
The fine-tuning Attract-Repel procedure is entirely driven by the input Attract and Repel sets of constraints. These can be extracted from a variety of semantic databases such as WordNet Fellbaum (1998), the Paraphrase Database Ganitkevitch et al. (2013); Pavlick et al. (2015), or BabelNet Navigli and Ponzetto (2012); Ehrmann et al. (2014) as done in prior work (Faruqui et al., 2015; Wieting et al., 2015; Mrkšić et al., 2016, i.a.). In this work, we investigate another option: extracting constraints without curated knowledge bases in a spectrum of languages by exploiting inherent language-specific properties related to linguistic morphology. This relaxation ensures a wider portability of Attract-Repel to languages and domains without readily available or adequate resources.
Extracting Attract Pairs
The core difference between inflectional and derivational morphology can be summarised in a few lines as follows: the former refers to a set of processes through which the word form expresses meaningful syntactic information, e.g., verb tense, without any change to the semantics of the word. On the other hand, the latter refers to the formation of new words with semantic shifts in meaning Schone and Jurafsky (2001); Haspelmath and Sims (2013); Lazaridou et al. (2013); Zeller et al. (2013); Cotterell and Schütze (2017).
For the Attract constraints, we focus on inflectional rather than on derivational morphology rules as the former preserve the full meaning of a word, modifying it only to reflect grammatical roles such as verb tense or case markers (e.g., (en_read, en_reads) or (de_katalanisch, de_katalanischer)). This choice is guided by our intent to fine-tune the original vector space in order to improve the embedded semantic relations.
We define two rules for English, widely recognised as morphologically simple Avramidis and Koehn (2008); Cotterell et al. (2016b). These are: (R1) if , where + ing/ed/s, then add and to the set of Attract constraints . This rule yields pairs such as (look, looks), (look, looking), (look, looked).
If is a function which strips the last character from word , the second rule is: (R2) if ends with the letter e and and , where + ing/ed, then add and to . This creates pairs such as (create, creating) and (create, created). Naturally, introducing more sophisticated rules is possible in order to cover for other special cases and morphological irregularities (e.g., sweep / swept), but in all our en experiments, is based on the two simple en rules R1 and R2.
The other three languages, with more complicated morphology, yield a larger number of rules. In Italian, we rely on the sets of rules spanning: (1) regular formation of plural (libro / libri); (2) regular verb conjugation (aspettare / aspettiamo); (3) regular formation of past participle (aspettare / aspettato); and (4) rules regarding grammatical gender (bianco / bianca). Besides these, another set of rules is used for German and Russian: (5) regular declension (e.g., asiatisch / asiatischem).
Extracting Repel Pairs
As another source of implicit semantic signals, also contains words which represent derivational antonyms: e.g., two words that denote concepts with opposite meanings, generated through a derivational process. We use a standard set of en “antonymy” prefixes: {dis, il, un, in, im, ir, mis, non, anti} Fromkin et al. (2013). If , where is generated by adding a prefix from to , then and are added to the set of Repel constraints . This rule generates pairs such as (advantage, disadvantage) and (regular, irregular). An additional rule replaces the suffix -ful with -less, extracting antonyms such as (careful, careless).
Following the same principle, we use {un, nicht, anti, ir, in, miss}, {in, ir, im, anti}, and {не, анти}. For instance, this generates an it pair (rispettoso, irrispettoso) (see Fig. 1). For de, we use another rule targeting suffix replacement: -voll is replaced by -los.
We further expand the set of Repel constraints by transitively combining antonymy pairs from the previous step with inflectional Attract pairs. This step yields additional constraints such as (rispettosa, irrispettosi) (see Fig. 1). The final and constraint counts are given in Tab. 3. The full sets of rules are available as supplemental material.
For each of the four languages we train the skip-gram with negative sampling (SGNS) model Mikolov et al. (2013) on the latest Wikipedia dump of each language. We induce 300-dimensional word vectors, with the frequency cut-off set to 10. The vocabulary sizes for each language are provided in Tab. 3.Other SGNS parameters were set to standard values Baroni et al. (2014); Vulić and Korhonen (2016b): epochs, negative samples, global learning rate: , subsampling rate: . Similar trends in results persist with . We label these collections of vectors sgns-large.
Other Starting Distributional Vectors
We also analyse the impact of morph-fitting on other collections of well-known en word vectors. These vectors have varying vocabulary coverage and are trained with different architectures. We test standard distributional models: Common-Crawl GloVe Pennington et al. (2014), SGNS vectors Mikolov et al. (2013) with various contexts (BOW = bag-of-words; DEPS = dependency contexts), and training data (PW = Polyglot Wikipedia from Al-Rfou et al. (2013); 8B = 8 billion token word2vec corpus), following Levy and Goldberg (2014) and Schwartz et al. (2015). We also test the symmetric-pattern based vectors of Schwartz et al. (2016) (SymPat-Emb), count-based PMI-weighted vectors reduced by SVD Baroni et al. (2014) (Count-SVD), a model which replaces the context modelling function from CBOW with bidirectional LSTMs Melamud et al. (2016) (Context2Vec), and two sets of en vectors trained by injecting multilingual information: BiSkip Luong et al. (2015) and MultiCCA Faruqui and Dyer (2014).
We also experiment with standard well-known distributional spaces in other languages (it and de), available from prior work Dinu et al. (2015); Luong et al. (2015); Vulić and Korhonen (2016a).
Morph-fixed Vectors
A baseline which utilises an equal amount of knowledge as morph-fitting, termed morph-fixing, fixes the vector of each word to the distributional vector of its most frequent inflectional synonym, tying the vectors of low-frequency words to their more frequent inflections. For each word , we construct a set of words consisting of the word itself and all words which co-occur with in the Attract constraints. We then choose the word from the set with the maximum frequency in the training data, and fix all other word vectors in to its word vector. The morph-fixed vectors (MFix) serve as our primary baseline, as they outperformed another straightforward baseline based on stemming across all of our intrinsic and extrinsic experiments.
Morph-fitting Variants
We analyse two variants of morph-fitting: (1) using Attract constraints only (MFit-A), and (2) using both Attract and Repel constraints (MFit-AR).
The first set of experiments intrinsically evaluates morph-fitted vector spaces on word similarity benchmarks, using Spearman’s rank correlation as the evaluation metric. First, we use the SimLex-999 dataset, as well as SimVerb-3500, a recent en verb pair similarity dataset providing similarity ratings for 3,500 verb pairs.Unlike other gold standard resources such as WordSim-353 Finkelstein et al. (2002) or MEN Bruni et al. (2014), SimLex and SimVerb provided explicit guidelines to discern between semantic similarity and association, so that related but non-similar words (e.g. cup and coffee) have a low rating. SimLex-999 was translated to de, it, and ru by Leviant and Reichart (2015), and they crowd-sourced similarity scores from native speakers. We use this dataset for our multilingual evaluation.Since Leviant and Reichart (2015) re-scored the original en SimLex, we use their en SimLex version for consistency.
Morph-fitting en Word Vectors
As the first experiment, we morph-fit a wide spectrum of en distributional vectors induced by various architectures (see Sect. 3). The results on SimLex and SimVerb are summarised in Tab. 4. The results with en sgns-large vectors are shown in Fig. 3(a). Morph-fitted vectors bring consistent improvement across all experiments, regardless of the quality of the initial distributional space. This finding confirms that the method is robust: its effectiveness does not depend on the architecture used to construct the initial space. To illustrate the improvements, note that the best score on SimVerb for a model trained on running text is achieved by Context2vec (); injecting morphological constraints into this vector space results in a gain of points.
Experiments on Other Languages
We next extend our experiments to other languages, testing both morph-fitting variants. The results are summarised in Tab. 5, while Fig. 3(a)-3(d) show results for the morph-fitted sgns-large vectors. These scores confirm the effectiveness and robustness of morph-fitting across languages, suggesting that the idea of fitting to morphological constraints is indeed language-agnostic, given the set of language-specific rule-based constraints. Fig. 3 also demonstrates that the morph-fitted vector spaces consistently outperform the morph-fixed ones.
The comparison between MFit-A and MFit-AR indicates that both sets of constraints are important for the fine-tuning process. MFit-A yields consistent gains over the initial spaces, and (consistent) further improvements are achieved by also incorporating the antonymous Repel constraints. This demonstrates that both types of constraints are useful for semantic specialisation.
Comparison to Other Specialisation Methods
We also tried using other post-processing specialisation models from the literature in lieu of Attract-Repel using the same set of “morphological” synonymy and antonymy constraints. We compare Attract-Repel to the retrofitting model of Faruqui et al. (2015) and counter-fitting Mrkšić et al. (2017a). The two baselines were trained for 20 iterations using suggested settings. The results for en, de, and it are summarised in Fig. 2. They clearly indicate that MFit-AR outperforms the two other post-processors for each language. We hypothesise that the difference in performance mainly stems from context-sensitive vector space updates performed by Attract-Repel. Conversely, the other two models perform pairwise updates which do not consider what effect each update has on the example pair’s relation to other word vectors (for a detailed comparison, see Mrkšić et al. (2017b)).
Besides their lower performance, the two other specialisation models have additional disadvantages compared to the proposed morph-fitting model. First, retrofitting is able to incorporate only synonymy/Attract pairs, while our results demonstrate the usefulness of both types of constraints, both for intrinsic evaluation (Tab. 5) and downstream tasks (see later Fig. 3). Second, counter-fitting is computationally intractable with sgns-large vectors, as its regularisation term involves the computation of all pairwise distances between words in the vocabulary.
Further Discussion
The simplicity of the used language-specific rules does come at a cost of occasionally generating incorrect linguistic constraints such as (tent, intent), (prove, improve) or (press, impress). In future work, we will study how to further refine extracted sets of constraints. We also plan to conduct experiments with gold standard morphological lexicons on languages for which such resources exist Sylak-Glassman et al. (2015); Cotterell et al. (2016b), and investigate approaches which learn morphological inflections and derivations in different languages automatically as another potential source of morphological constraints (Soricut and Och, 2015; Cotterell et al., 2016a; Faruqui et al., 2016; Kann et al., 2017; Aharoni and Goldberg, 2017, i.a.).
Goal-oriented dialogue systems provide conversational interfaces for tasks such as booking flights or finding restaurants. In slot-based systems, application domains are specified using ontologies that define the search constraints which users can express. An ontology consists of a number of slots and their assorted slot values. In a restaurant search domain, sets of slot-values could include price = [cheap, expensive] or food = [Thai, Indian, …].
The DST model is the first component of modern dialogue pipelines Young (2010). It serves to capture the intents expressed by the user at each dialogue turn and update the belief state. This probability distribution over the possible dialogue states (defined by the domain ontology) is the system’s internal estimate of the user’s goals. It is used by the downstream dialogue manager component to choose the subsequent system response Su et al. ; Su et al. (2016). The following example shows the true dialogue state in a multi-turn dialogue:
The Dialogue State Tracking Challenge (DSTC) shared task series formalised the evaluation and provided labelled DST datasets Henderson et al. (2014a, b); Williams et al. (2016). While a plethora of DST models are available based on, e.g., hand-crafted rules Wang et al. (2014) or conditional random fields Lee and Eskenazi (2013), the recent DST methodology has seen a shift towards neural-network architectures (Henderson et al., 2014c, d; Zilka and Jurcicek, 2015; Mrkšić et al., 2015; Perez and Liu, 2017; Liu and Perez, 2017; Vodolán et al., 2017; Mrkšić et al., 2017a, i.a.).
To detect intents in user utterances, most existing models rely on either (or both): 1) Spoken Language Understanding models which require large amounts of annotated training data; or 2) hand-crafted, domain-specific lexicons which try to capture lexical and morphological variation. The Neural Belief Tracker (NBT) is a novel DST model which overcomes both issues by reasoning purely over pre-trained word vectors Mrkšić et al. (2017a). The NBT learns to compose these vectors into intermediate utterance and context representations. These are then used to decide which of the ontology-defined intents (goals) have been expressed by the user. The NBT model keeps word vectors fixed during training, so that unseen, yet related words can be mapped to the right intent at test time (e.g. northern to north).
Data: Multilingual WOZ 2.0 Dataset
Our DST evaluation is based on the WOZ dataset, released by Wen et al. (2017). In this Wizard-of-Oz setup, two Amazon Mechanical Turk workers assumed the role of the user and the system asking/providing information about restaurants in Cambridge (operating over the same ontology and database used for DSTC2 Henderson et al. (2014a)). Users typed instead of speaking, removing the need to deal with noisy speech recognition. In DSTC datasets, users would quickly adapt to the system’s inability to deal with complex queries. Conversely, the WOZ setup allowed them to use sophisticated language. The WOZ 2.0 release expanded the dataset to 1,200 dialogues Mrkšić et al. (2017a). In this work, we use translations of this dataset to Italian and German, released by Mrkšić et al. (2017b).
Evaluation Setup
The principal metric we use to measure DST performance is the joint goal accuracy, which represents the proportion of test set dialogue turns where all user goals expressed up to that point of the dialogue were decoded correctly Henderson et al. (2014a). The NBT models for en, de and it are trained using four variants of the sgns-large vectors: 1) the initial distributional vectors; 2) morph-fixed vectors; 3) and 4) the two variants of morph-fitted vectors (see Sect. 3).
As shown by Mrkšić et al. (2017b), semantic specialisation of the employed word vectors benefits DST performance across all three languages. However, large gains on SimLex-999 do not always induce correspondingly large gains in downstream performance. In our experiments, we investigate the extent to which morph-fitting improves DST performance, and whether these gains exhibit stronger correlation with intrinsic performance.
Results and Discussion
The dark bars (against the right axes) in Fig. 3 show the DST performance of NBT models making use of the four vector collections. it and de benefit from both kinds of morph-fitting: it performance increases from (MFit-A) and de performance rises even more: (MFit-AR), setting a new state-of-the-art score for both datasets. The morph-fixed vectors do not enhance DST performance, probably because fixing word vectors to their highest frequency inflectional form eliminates useful semantic content encoded in the original vectors. On the other hand, morph-fitting makes use of this information, supplementing it with semantic relations between different morphological forms. These conclusions are in line with the SimLex gains, where morph-fitting outperforms both distributional and morph-fixed vectors.
English performance shows little variation across the four word vector collections investigated here. This corroborates our intuition that, as a morphologically simpler language, English stands to gain less from fine-tuning the morphological variation for downstream applications. This result again points at the discrepancy between intrinsic and extrinsic evaluation: the considerable gains in SimLex performance do not necessarily induce similar gains in downstream performance. Additional discrepancies between SimLex and downstream DST performance are detected for German and Italian. While we observe a slight drop in SimLex performance with the de MFit-AR vectors compared to the MFit-A ones, their relative performance is reversed in the DST task. On the other hand, we see the opposite trend in Italian, where the MFit-A vectors score lower than the MFit-AR vectors on SimLex, but higher on the DST task. In summary, we believe these results show that SimLex is not a perfect proxy for downstream performance in language understanding tasks. Regardless, its performance does correlate with downstream performance to a large extent, providing a useful indicator for the usefulness of specific word vector spaces for extrinsic tasks such as DST.
A standard approach to incorporating external information into vector spaces is to pull the representations of similar words closer together. Some models integrate such constraints into the training procedure, modifying the prior or the regularisation Yu and Dredze (2014); Xu et al. (2014); Bian et al. (2014); Kiela et al. (2015), or using a variant of the SGNS-style objective Liu et al. (2015); Osborne et al. (2016). Another class of models, popularly termed retrofitting, injects lexical knowledge from available semantic databases (e.g., WordNet, PPDB) into pre-trained word vectors Faruqui et al. (2015); Jauhar et al. (2015); Wieting et al. (2015); Nguyen et al. (2016); Mrkšić et al. (2016). Morph-fitting falls into the latter category. However, instead of resorting to curated knowledge bases, and experimenting solely with English, we show that the morphological richness of any language can be exploited as a source of inexpensive supervision for fine-tuning vector spaces, at the same time specialising them to better reflect true semantic similarity, and learning more accurate representations for low-frequency words.
Word Vectors and Morphology
The use of morphological resources to improve the representations of morphemes and words is an active area of research. The majority of proposed architectures encode morphological information, provided either as gold standard morphological resources Sylak-Glassman et al. (2015) such as CELEX Baayen et al. (1995) or as an external analyser such as Morfessor Creutz and Lagus (2007), along with distributional information jointly at training time in the language modelling (LM) objective (Luong et al., 2013; Botha and Blunsom, 2014; Qiu et al., 2014; Cotterell and Schütze, 2015; Bhatia et al., 2016, i.a.). The key idea is to learn a morphological composition function Lazaridou et al. (2013); Cotterell and Schütze (2017) which synthesises the representation of a word given the representations of its constituent morphemes. Contrary to our work, these models typically coalesce all lexical relations.
Another class of models, operating at the character level, shares a similar methodology: such models compose token-level representations from subcomponent embeddings (subwords, morphemes, or characters) (dos Santos and Zadrozny, 2014; Ling et al., 2015; Cao and Rei, 2016; Kim et al., 2016; Wieting et al., 2016; Verwimp et al., 2017, i.a.).
In contrast to prior work, our model decouples the use of morphological information, now provided in the form of inflectional and derivational rules transformed into constraints, from the actual training. This pipelined approach results in a simpler, more portable model. In spirit, our work is similar to Cotterell et al. (2016b), who formulate the idea of post-training specialisation in a generative Bayesian framework. Their work uses gold morphological lexicons; we show that competitive performance can be achieved using a non-exhaustive set of simple rules. Our framework facilitates the inclusion of antonyms at no extra cost and naturally extends to constraints from other sources (e.g., WordNet) in future work. Another practical difference is that we focus on similarity and evaluate morph-fitting in a well-defined downstream task where the artefacts of the distributional hypothesis are known to prompt statistical system failures.
We have presented a novel morph-fitting method which injects morphological knowledge in the form of linguistic constraints into word vector spaces. The method makes use of implicit semantic signals encoded in inflectional and derivational rules which describe the morphological processes in a language. The results in intrinsic word similarity tasks show that morph-fitting improves vector spaces induced by distributional models across four languages. Finally, we have shown that the use of morph-fitted vectors boosts the performance of downstream language understanding models which rely on word representations as features, especially for morphologically rich languages such as German.
Future work will focus on other potential sources of morphological knowledge, porting the framework to other morphologically rich languages and downstream tasks, and on further refinements of the post-processing specialisation algorithm and the constraint selection.
This work is supported by the ERC Consolidator Grant LEXICAL: Lexical Acquisition Across Languages (no 648909). RR is supported by the Intel-ICRI grant: Hybrid Models for Minimally Supervised Information Extraction from Conversations. The authors are grateful to the anonymous reviewers for their helpful suggestions.
Morphological Rules
In this supplemental material, we provide a short comprehensive overview of simple language-specific morphological rules in English (en), German (de), Italian (ut), and Russian (ru). These rules were used to build the sets of synonymous Attract and antonymous Repel constraints for our morph-fitting fine-tuning procedure. As discussed in the paper, the linguistic constraints extracted from the rules require only a comprehensive list of vocabulary words in each language. A native speaker of each language used in our experiments is able to easily come up with these sets of morphological rules (or at least with a reasonable subset of rules) without any linguistic training. What is more, the rules for German, Italian, and Russian were created by non-native and non-fluent speakers who have only a passive or limited knowledge of the three languages, exemplifying the simplicity and portability of the fine-tuning approach based on the shallow “morphological supervision”. The simplicity is also confirmed by the short time used to compile the rules, ranging from a few minutes for English to approximately two hours for Russian.
Different languages differ in their “morphological richness” (e.g., declension, verb conjugation, plural forming, gender) which consequently leads to the varying number of rules in each language. However, all four languages in our study display morphological regularities described by simple morphological rules that are exploited to build sets of Attract and Repel linguistic constraints in each language from scratch.Note that the rules for extracting Attract constraints were additionally used to generate the Morph-SimLex evaluation set, also provided as supplemental material.
Vocabularies in all four languages are labeled , , , . We add the pairs and generated by the rules to the sets of constraints iff both . After we generate all such constraints, since some constraints may have been generated by more than one rule, we remove all duplicates from the respective sets of Attract and Repel constraints.
Before we start, we will define two simple functions: (i) the function strips the last characters from the word , (ii) the function .ew(sub) tests if the word ends with a sequence of characters sub. For instance, returns , while .ew(’s’) returns False and .ew(’e’) returns True.
As discussed in the paper, we rely on only two simple inflectional morphological rules in English:
- + ’s’/’ed’/’ing’. This rule yields constraints such as (speak, speaking), (turtle, turtles), or (clean, cleaned).
- If .ew(’e’), then + ’ed’/’ing’. This rule yields constraints such as (create, creating), or (generate, generated).
Derivational Antonymy: Repel
We assume the following set of standard “antonymy” prefixes in English: {’dis’, ’il’, ’un’, ’in’, ’im’, ’ir’, ’mis’, ’non’, ’anti’}. We rely on the following derivational rules to extract Repel pairs:
- = ap + , where ap . This rule yields constraints such as (mature, immature), (allow, disallow) or (regularity, irregularity).
- If .ew(’ful’), then = + ’less’. This rule yields constraints such as (cheerful, cheerless).
As mentioned in the paper, for all four languages we further expand the set of Repel constraints by transitively combining antonymy pairs with inflectional Attract pairs. In simple words, the friend of my enemy is my enemy. This means that, given an Attract pair (allow, allows) and a Repel pair (allow, disallow), we extract another Repel pair (allows, disallow).
German Rules
Being morphologically richer than English, the German language naturally requires more rules to describe its (inflectional) morphological richness and variation. First, we capture the regular declension of nouns and adjectives by the following heuristic:
- Generate a set of words ; take the Cartesian product on and then exclude pairs with identical words. This rule generates pairs such as (schottisch, schottische), (schottischem, schottischen).
The second set of rules describes regular verb morphology, i.e., verb conjugation in the present and past tense, and the formation of regular past participles. This set of rules may be expressed as:
- If .ew(’en’), then . If .ew(’t’), then generate a set of words , else (if not .ew(’t’)), generate a set of words . We then take the Cartesian product on . Again, all pairs with identical words were discarded. This rule yields pairs such as (machen, machten), (mache, gemacht), (kaufst, kauft), or (arbeite, arbeitete) and (arbeiten, gearbeitet).
Another set of rules targets the regular formation of plural nouns:
- If .ew(’ei’) or .ew(’heit’) or .ew(’keit’) or .ew(’schaft’) or .ew(’ung’), then = + ’en’. This rule yields pairs such as (wahrheit, wahrheiten) or (gemeinschaft, gemeinschaften).
- If .ew(’in’), then = + ’nen’. This rule generates pairs such as (lehrerin, lehrerinnen) or (lektorin, lektorinnen).
- If .ew(’a’/’i’/’o’/’u’/’y’) then = + ’s’. This rule yields pairs such as (auto, autos).
- If .ew(’e’), then = + ’n’. This rule yields pairs such as (postkarte, postkarten).
- + er, where the function replaces the last occurrence of the letter ’a’,’o’ or ’u’ with ’ä’,’ö’ or ’ü’. This rule generates pairs such as (wörterbuch, wörterbücher) or (stadt, städter).
Derivational Antonymy: Repel
We assume the following set of standard “antonymy” prefixes in German: {’un’, ’nicht’, ’anti’, ’ir’, ’in’, ’miss’}. We rely on the following derivational rules to extract Repel pairs in German:
- = ap + , where ap . This rule yields constraints such as (aktiv, inaktiv), (wandelbar, unwandelbar) or (zyklone, antizyklone).
- If .ew(’voll’), then = + ’los’. This rule yields constraints such as (geschmackvoll, geschmacklos).
The set of Repel is then again transitively expanded yielding pairs such as (relevant, irrelevanter) or (aktivem, inaktiv).
Italian Rules
The first set of rules aims at capturing the regular plural forming in Italian (e.g., libro, libri) and regular differences in gender (e.g., rapido, rapida). We rely on the simple heuristic which can be expressed as follows:
- If .ew(’a’/’e’/’o’/’i’), then generate a set of words , and take the Cartesian product on discarding pairs with identical words. This rule yields pairs such as (nero, neri) or (generazione, generazioni).
- If .ew(’ga’/’ca’), then + ’he’. This rule generates pairs such as (tartaruga, tartarughe) or (bianca, bianche).
- If .ew(’go’), then + ’hi’. This rule generates pairs such as (albergo, alberghi).
The second set of rules targets regular verb conjugation in Italian and the formation of regular past participles. The following rules are used:
- If .ew(’are’), then generate a set of words ; take the Cartesian product on discarding pairs with identical words. This rule results in pairs such as (aspettare, aspettiamo).
- If .ew(’ere’), then generate a set of words ; take the Cartesian product on discarding pairs with identical words. This rule results in pairs such as (ricevere, ricevete) or (riceve, ricevuto).
- If .ew(’ire’), then generate a set of words ; take the Cartesian product on discarding pairs with identical words. This rule results in pairs such as (dormire, dormono) or (dormi, dormita).
Derivational Antonymy: Repel
We assume the following set of standard “antonymy” prefixes: {’in’, ’ir’, ’im’, ’anti’}. The following derivational rule is used to extract Repel pairs:
- = ap + , where ap . This rule yields constraints such as (attivo, inattivo) or (rispettosa, irrispettosa).
The set of Repel was then expanded as before, e.g., with additional pairs such as (rispettosa, irrispettosi) generated.
Russian Rules
The first set of rules in Russian targets the regular forming of plural in Russian. A few simple heuristics are used as follows:
- + ’и’/’ы’. This rule yields pairs such as (aльбом, aльбомы), transliterated as: (al’bom, al’bomy).
- if .ew(’a’/’я’/’ь’), then + ’и’/’ы’. This rule generates pairs such as (песня, песни): (pesnja, pesni).
- if .ew(’o’), then + ’a’. This rule generates pairs such as (письмо, письма): (pis’mo, pis’ma).
- if .ew(’e’), then + ’я’. This rule generates pairs such as (платье, платья): (plat’e, plat’ja).
The next set of rules targets regular verb conjugation of Russian verbs as well as the regular formation of past participles. We again build a simple heuristic to extract Attract pairs:
- if .ew(’ти’/’ть’), then generate a set of words and take the Cartesian product on discarding pairs with identical words. This rule yields pairs such as (варить, варите) or (заканчиваю, заканчивают), transliterated as: (varit’, varite), (zakanchivaju, zakanchivajut).
Following that, we also utilise the regularities regarding declension processes in Russian, captured by the following rules:
- if .ew(’a’), then generate a set of words and take the Cartesian product on discarding pairs with identical words. This rule yields pairs such as (работа, работой): (rabota, rabotoj).
- if .ew(’я’), then generate a set of words and take the Cartesian product on discarding pairs with identical words. This rule yields pairs such as (линия, линию): (linija, liniju).
- if .ew(’ы’), then generate a set of words and take the Cartesian product on discarding pairs with identical words. This rule yields pairs such as (работам, работами): (rabotam, rabotami).
- if .ew(’и’), then generate a set of words and take the Cartesian product on discarding pairs with identical words. This rule yields pairs such as (работам, работами): (rabotam, rabotami).
Yet another set of rules targets regular adjective comparison and gender:
- if .ew(’ый’/’ой’/’ий’), then generate a set of words . This rule yields pairs such as (быстрый, быстрее): (bystryj, bystree).
- if .ew(’ая’), then generate a set of words . This rule yields pairs such as (новая, новыe): (novaja, novye).
- if .ew(’oe’), then generate a set of words . This rule yields pairs such as (новое, новый): (novoe, novyj).
Derivational Antonymy: Repel
We assume the following set of standard “antonymy” prefixes in Russian: {не, анти’}, and simply use the following rule:
- = ap + , where ap . This rule yields constraints such as (адекватный, неадекватный) or (вирусная, антивирусная), transliterated as: (adekvatnyj, neadekvatnyj) and (virusnaja, antivirusnaja).
The further expansion of Repel constraints yields pairs such as (адекватный, неадекватная): (adekvatnyj, neadekvatnaja).
Further Discussion
We stress that the listed rules for all four languages are non-exhaustive and do not cover all possible inflectional and derivational morphological phenomena. More linguistic constraints may be extracted by resorting to more sophisticated rules covering finer-grained morphological processes (e.g., covering irregular plural forming or irregular verb conjugation and past participle forming, or non-standard declensions). Further, the listed rules, written by non-native speakers without any linguistic training in a very short time span, do not necessarily rely on established linguistic theories in each language, but are rather simple heuristics aiming to capture morphological regularities.