Bilingual Learning of Multi-sense Embeddings with Discrete Autoencoders

Simon Šuster, Ivan Titov, Gertjan van Noord

Introduction

Approaches to learning word embeddings (i.e. real-valued vectors) relying on word context have received much attention in recent years, and the induced representations have been shown to capture syntactic and semantic properties of words. They have been evaluated intrinsically [2013a, 2014, 2014] and have also been used in concrete NLP applications to deal with word sparsity and improve generalization . While most work to date has focused on developing embedding models which represent a word with a single vector, some researchers have attempted to capture polysemy explicitly and have encoded properties of each word with multiple vectors .

In parallel to this work on multi-sense word embeddings, another line of research has investigated integrating multilingual data, with largely two distinct goals in mind. The first goal has been to obtain representations for several languages in the same semantic space, which then enables the transfer of a model (e.g., a syntactic parser) trained on annotated training data in one language to another language lacking this annotation . Secondly, information from another language can also be leveraged to yield better first-language embeddings . Our paper falls in the latter, much less explored category. We adhere to the view of multilingual learning as a means of language grounding [2014b, 2013, 2012, 2010, 2009]. Intuitively, polysemy in one language can be at least partially resolved by looking at the translation of the word and its context in another language . Better sense assignment can then lead to better sense-specific word embeddings.

We propose a model that uses second-language embeddings as a supervisory signal in learning multi-sense representations in the first language. This supervision is easy to obtain for many language pairs as numerous parallel corpora exist nowadays. Our model, which can be seen as an autoencoder with a discrete hidden layer encoding word senses, leverages bilingual data in its encoding part, while the decoder predicts the surrounding words relying on the predicted senses. We strive to remain flexible as to the form of parallel data used in training and support both the use of word- and sentence-level alignments.

The second-language signal effectively improves the quality of multi-sense embeddings as seen on a variety of intrinsic tasks for English, with the results superior to that of the baseline Skip-Gram model, even though the crosslingual information is not available at test time.

This finding is robust across several settings, such as varying dimensionality, vocabulary size and amount of data.

In the extrinsic POS-tagging task, the second-language signal also offers improvements over monolingually-trained multi-sense embeddings, however, the standard Skip-Gram embeddings turn out to be the most robust in this task.

We make the implementation of all the models as well as the evaluation scripts available at http://github.com/rug-compling/bimu.

Word Embeddings with Discrete Autoencoders

Our method borrows its general structure from neural autoencoders . Autoencoders are trained to reproduce their input by first mapping their input to a (lower dimensional) hidden layer and then predicting an approximation of the input relying on this hidden layer. In our case, the hidden layer is not a real-valued vector, but is a categorical variable encoding the sense of a word. Discrete-state autoencoders have been successful in several natural language processing applications, including POS tagging and word alignment , semantic role induction and relation discovery .

More formally, our model consists of two components: an encoding part which assigns a sense to a pivot word, and a reconstruction (decoding) part recovering context words based on the pivot word and its sense. As predictions are probabilistic (‘soft’), the reconstruction step involves summation over all potential word senses. The goal is to find embedding parameters which minimize the error in recovering context words based on the pivot word and the sense assignment. Parameters of both encoding and reconstruction are jointly optimized. Intuitively, a good sense assignment should make the reconstruction step as easy as possible. The encoder uses not only words in the first-language sentence to choose the sense but also, at training time, is conditioning its decisions on the words in the second-language sentence. We hypothesize that the injection of crosslingual information will guide learning towards inducing more informative sense-specific word representations. Consequently, using this information at training time would benefit the model even though crosslingual information is not available to the encoder at test time.

We specify the encoding part as a log-linear model:

The reconstruction part predicts a context word xjx_{j} given the pivot xix_{i} and the current estimate of its ss:

where ∣V∣|\mathcal{V}| is the vocabulary size. This is effectively a Skip-Gram model [2013a] extended to rely on senses.

As sense assignments are not observed during training, the learning objective includes marginalization over word senses and thus can be written as:

in which index ii goes over all pivot words in the first language, jj over all context words to predict at each ii, and ss marginalizes over all possible senses of the word xix_{i}. In practice, we avoid the costly computation of the normalization factor in the softmax computation of Eq. (2) and use negative sampling [2013b] instead of log⁡p(xj∣xi,s,θ)\log p(x_{j}|x_{i},s,\theta):

where σ\sigma is the sigmoid non-linearity function and γx\gamma_{x} is a word embedding from the sample of negative (noisy) words NN. Optimizing the autoencoding objective is broadly similar to the learning algorithm defined for multi-sense embedding induction in some of the previous work . Note though that this previous work has considered only monolingual context.

We use a minibatch training regime and seek to optimize the objective function L(B,θ)L(\mathcal{B},\theta) for each minibatch B\mathcal{B}. We found that optimizing this objective directly often resulted in inducing very flat posterior distributions. We therefore use a form of posterior regularization where we can encode our prior expectations that the posteriors should be sharp. The regularized objective for a minibatch is defined as

2 Obtaining word representations

At test time, we construct the word representations by averaging all sense embeddings for a word xix_{i} and weighting them with the sense expectations Although our training objective has sparsity-inducing properties, the posteriors at test time are not entirely peaked, which makes weighting beneficial.:

Unlike in training, the sense prediction step here does not use the crosslingual context Ci′C^{\prime}_{i} since it is not available in the evaluation tasks. In this work, instead of marginalizing out the unobservable crosslingual context, we simply ignore it in computation.

Sometimes, even the first-language context is missing, as is the situation in many word similarity tasks. In that case, we just use the uniform average, \nicefrac1∣S∣∑s∈Sφi,s\nicefrac{{1}}{{|\mathcal{S}|}}\sum_{s\in\mathcal{S}}\varphi_{i,s}.

Word affiliation from alignments

In defining the crosslingual signal we draw on a heuristic inspired by Devlin et al. . The second-language context words are taken to be the multiset of words around and including the pivot affiliated to xix_{i}:

where xai′x^{\prime}_{a_{i}} is the word affiliated to xix_{i} and the parameter mm regulates the context window size. By choosing m=0m=0, only the affiliated word is used as l′l^{\prime} context, and by choosing m=∞m=\infty, the l′l^{\prime} context is the entire sentence (≈\approxuniform alignment). To obtain the index aia_{i}, we use the following:

If xix_{i} aligns to exactly one second-language word, aia_{i} is the index of the word it aligns to.

If xix_{i} aligns to multiple words, aia_{i} is the index of the aligned word in the middle (and rounding down when necessary).

If xix_{i} is unaligned, Ci′C^{\prime}_{i} is empty, therefore no l′l^{\prime} context is used.

We use the cdec aligner to word-align the parallel corpora.

Parameters and Set-up

We use the AdaGrad optimizer with initial learning rate set to 0.1. We set the minibatch size to 1000, the number of negative samples to 1, the sampling factor to 0.001 and the window size parameter mm to 5. All the embeddings are 50-dimensional (unless specified otherwise) and initialized by sampling from the uniform distribution between [−0.05,0.05][-0.05,0.05]. We include in the vocabulary all words occurring in the corpus at least 20 times. We set the number of senses per word to 3 (see further discussion in § 6.4 and § 7). All other parameters with their default values can be examined in the source code available online.

2 Bilingual data

In a large body of work on multilingual word representations, Europarl is the preferred source of parallel data. However, the domain of Europarl is rather constrained, whereas we would like to obtain word representations of more general language, also to carry out an effective evaluation on semantic similarity datasets where domains are usually broader. We therefore use the following parallel corpora: News Commentary (NC), Yandex-1Mhttps://translate.yandex.ru/corpus (RU-EN), CzEng 1.0 (CZ-EN) from which we exclude the EU legislation texts, and GigaFrEn (FR-EN). The sizes of the corpora are reported in Table 1. The word representations trained on the NC corpora are evaluated only intrinsically due to the small sizes.

Evaluation Tasks

We evaluate the quality of our word representations on a number of tasks, both intrinsic and extrinsic.

We are interested here in how well the semantic similarity ratings obtained from embedding comparisons correlate to human ratings. For this purpose, we use a variety of similarity benchmarks for English and report the Spearman ρ\rho correlation scores between the human ratings and the cosine ratings obtained from our word representations. The SCWS benchmark is probably the most suitable similarity dataset for evaluating multi-sense embeddings, since it allows us to perform the sense prediction step based on the sentential context provided for each word in the pair.

The other benchmarks we use provide the ratings for the word pairs without context. WS-353 contains 353 human-rated word pairs , while Agirre et al. separate this benchmark for similarity (WS-SIM) and relatedness (WS-REL). The RG-65 and the MC-30 benchmarks contain nouns only. The MTurk-287 and MTurk-771 include word pairs whose similarity was crowdsourced from AMT. Similarly, MEN is an AMT-annotated dataset of 3000 word pairs. The YP-130 and Verb-143 measure verb similarity. Rare-Word contains 2034 rare-word pairs. Finally, SimLex-999 [2014b] is intended to measure pure similarity as opposed to relatedness. For these benchmarks, we prepare the word representations by taking a uniform average of all sense embeddings per word. The evaluation is carried out using the tool described in Faruqui and Dyer [2014a]. Due to space constraints, we report the results by averaging over all benchmarks (Similarity), and include the individual results in the online repository.

2 Supersense similarity

We also evaluate on a task measuring the similarity between the embeddings—in our case uniformly averaged in the case of multi-sense embeddings—and a matrix of supersense features extracted from the English SemCor, using the Qvec tool . We choose this method because it has been shown to output scores that correlate well with extrinsic tasks, e.g. text classification and sentiment analysis. We believe that this, in combination with word similarity tasks from the previous section, can give a reliable picture of the generic quality of word embeddings studied in this work.

3 POS tagging

As our downstream evaluation task, we use the learned word representations to initialize the embedding layer of a neural network tagging model. We use the same convolutional architecture as Li and Jurafsky : an input layer taking a concatenation of neighboring embeddings as input, three hidden layers with a rectified linear unit activation function and a softmax output layer. We train for 10 epochs using one sentence as a batch. Other hyperparameters can be examined in the source code. The multi-sense word embeddings are inferred from the sentential context (weighted average), as for the evaluation on the SCWS dataset. We use the standard splits of the Wall Street Journal portion of the Penn Treebank: 0–18 for training, 19–21 for development and 22–24 for testing.

Results

We compare three embeddings models, Skip-Gram (Sg), Multi-sense (Mu) and Bilingual Multi-sense (BiMu), using our own implementation for each of them. The first two can be seen as simpler variants of the BiMu model: in Sg we omit the encoder entirely, and in Mu we omit the second-language (l′l^{\prime}) part of the encoder in Eq. (2). We train the Sg and the Mu models on the English part of the parallel corpora. Those parameters common to all methods are kept fixed during experiments. The values λ\lambda and mm for controlling the second-language signal in BiMu are set on the POS-tagging development set (cf. § 6.3).

The results on the SCWS benchmark (Table 2) show consistent improvements of the BiMu model over Sg and Mu across all parallel corpora, except on the small CZ-EN (NC) corpus. We have also measured the 95% confidence intervals of the difference between the correlation coefficients of BiMu and Sg, following the method described in Zou . According to these values, BiMu significantly outperforms Sg on RU-EN, and on French, Russian and Spanish NC corpora.I.e. counting those results in which the CI of the difference does not include 0.

Next, ignoring any language-specific factors, we would expect to observe a trend according to which the larger the corpus, the higher the correlation score. However, this is not what we find. Among the largest corpora, i.e. RU-EN, CZ-EN and FR-EN, the models trained on RU-EN perform surprisingly well, practically on par with the 23-times larger FR-EN corpus. Similarly, the quality of the embeddings trained on CZ-EN is generally lower than when trained on the 10 times smaller RU-EN corpus. One explanation for this might be different text composition of the corpora, with RU-EN matching the domain of the evaluation task better than the larger two corpora. Also, FR-EN is known to be noisy, containing web-crawled sentences that are not parallel or not natural language . Furthermore, language-dependent effects might be playing a role: for example, there are signs of Czech being the least helpful language among those studied. But while there is evidence for that in all intrinsic tasks, the situation in POS tagging does not confirm this speculation.

We relate our models to previously reported SCWS scores from the literature using 300-dimensional models in Table 3. Even though we train on a much smaller corpus than the previous works,For example, Li and Jurafsky use the concatenation of Gigaword and Wikipedia with more than 5B words. the BiMu model achieves a very competitive correlation score.

The results on similarity benchmarks and qvec largely confirm those on SCWS, despite the lack of sentential context which would allow to weight the contribution of different senses more accurately for the multi-sense models. Why, then, does simply averaging the Mu and BiMu embeddings lead to better results than when using the Sg embeddings? We hypothesize that the single-sense model tends to over-represent the dominant sense with its generic, one-vector-per-word representation, whereas the uniformly averaged embeddings yielded by the multi-sense models better encode the range of potential senses. Similar observations have been made in the context of selectional preference modeling of polysemous verbs .

In POS tagging, the relationship between Mu and BiMu models is similar as discussed above. Overall, however, neither of the multi-sense models outperforms the Sg embeddings. The neural network tagger may be able to implicitly perform disambiguation on top of single-sense Sg embeddings, similarly to what has been argued in Li and Jurafsky . The tagging accuracies obtained with Mu on CZ-EN and FR-EN are similar to the one obtained by Li and Jurafsky with their multi-sense model (93.8), while the accuracy of Sg is more competitive in our case (around 94.0 compared to 92.5), although they use a larger corpus for training the word representations.

In all tasks, the addition of the bilingual component during training increases the accuracy of the encoder for most corpora, even though the bilingual information is not available during evaluation.

Fig. 2a displays how the semantic similarity as measured on SCWS evolves as a function of increasingly larger sub-samples from FR-EN, our largest parallel corpus. The BiMu embeddings show relatively stable improvements over Mu and especially over Sg embeddings. The same performance as that of Sg at 100% is achieved by Mu and BiMu sooner, using only around 40/50% of the corpus.

2 The dimensionality and frequent words

It is argued in Li and Jurafsky that often just increasing the dimensionality of the Sg model suffices to obtain better results than that of their multi-sense model. We look at the effect of dimensionality on semantic similarity in fig. 2b, and see that simply increasing the dimensionality of the Sg model (to any of 100, 200 or 300 dimensions) is not sufficient to outperform the Mu or BiMu models. When constraining the vocabulary to 6,000 most frequent words, the representations obtain higher quality. We can see that the models, especially Sg, benefit slightly more from the increased dimensionality when looking at these most frequent words. This is according to expectations—frequent words need more representational capacity due to their complex semantic and syntactic behavior .

3 The role of bilingual signal

The degree of contribution of the second language l′l^{\prime} during learning is affected by two parameters, λ\lambda for the trade-off between the importance of first and second language in the sense prediction part (encoder) and the value of mm for the size of the window around the second-language word affiliated to the pivot. Fig. 3a suggests that the context from the second language is useful in sense prediction, and that it should be weighted relatively heavily (around 0.7 and 0.8, depending on the language).

Regarding the role of the context-window size in sense disambiguation, the WSD literature has reported both smaller (more local) and larger (more topical) monolingual contexts to be useful, see e.g. Ide and Véronis for an overview. In fig. 3b we find that considering a very narrow context in the second language—the affiliated word only or a m=1m=1 window around it—performs the best, and that there is little gain in using a broader window. This is understandable since the l′l^{\prime} representation participating in the sense selection is simply an average over all generic embeddings in the window, which means that the averaged representation probably becomes noisy for large mm, i.e. more irrelevant words are included in the window. However, the negative effect on the accuracy is still relatively small, up to around −0.1-0.1 for the models using French and Russian as the second languages, and −0.25-0.25 for Czech when setting m=∞m=\infty. The infinite window size setting, corresponding to the sentence-only alignment, performs well also on SCWS, improving on the monolingual multi-sense baseline on all corpora (Table 4).

4 The number of senses

In our work, the number of senses kk is a model parameter, which we keep fixed to 3 throughout the empirical study. We comment here briefly on other choices of k∈{2,4,5}k\in\{2,4,5\}. We have found k=2k=2 to be a good choice on the RU-EN and FR-EN corpora (but not on CZ-EN), with an around 0.20.2-point improvement over k=3k=3 on SCWS and in POS tagging. With the larger values of kk, the performance tends to degrade. For example, on RU-EN, the k=5k=5 score on SCWS is about 0.60.6 point below our default setting.

Additional Related Work

Multi-sense models. One line of research has dealt with sense induction as a separate, clustering problem that is followed by an embedding learning component . In another, the sense assignment and the embeddings are trained jointly . Neelakantan et al. propose an extension of Skip-Gram [2013a] by introducing sense-specific parameters together with the kk-means-inspired ‘centroid’ vectors that keep track of the contexts in which word senses have occurred. They explore two model variants, one in which the number of senses is the same for all words, and another in which a threshold value determines the number of senses for each word. The results comparing the two variants are inconclusive, with the advantage of the dynamic variant being virtually nonexistent. In our work, we use the static approach. Whenever there is evidence for less senses than the number of available sense vectors, this is unlikely to be a serious issue as the learning would concentrate on some of the senses, and these would then be the preferred predictions also at test time. Li and Jurafsky build upon the work of Neelakantan et al. with a more principled method for introducing new senses using the Chinese Restaurant Processes (CRP). Our experiments confirm the findings of Neelakantan et al. that multi-sense embeddings improve Skip-gram embeddings on intrinsic tasks, as well as those of Li and Jurafsky, who find that multi-sense embeddings offer little benefit to the neural network learner on extrinsic tasks. Our discrete-autoencoding method when viewed without the bilingual part in the encoder has a lot in common with their methods.

Multilingual models. The research on using multilingual information in the learning of multi-sense embedding models is scarce. Guo et al. perform a sense induction step based on clustering translations prior to learning word embeddings. Once the translations are clustered, they are mapped to a source corpus using WSD heuristics, after which a recurrent neural network is trained to obtain sense-specific representations. Unlike in our work, the sense induction and embedding learning components are entirely separated, without a possibility for one to influence another. In a similar vein, Bansal et al. use bilingual corpora to perform soft word clustering, extending the previous work on the monolingual case of Lin and Wu . Single-sense representations in the multilingual context have been studied more extensively [2015, 2014b, 2014a, 2014, 2013, 2013], with a goal of bringing the representations in the same semantic space. A related line of work concerns the crosslingual setting, where one tries to leverage training data in one language to build models for typically lower-resource languages .

The recent works of Kawakami and Dyer and Nalisnick and Ravi are also of interest. The latter work on the infinite Skip-Gram model in which the embedding dimensionality is stochastic is relevant since it demonstrates that their embeddings exploit different dimensions to encode different word meanings. Just like us, Kawakami and Dyer use bilingual supervision, but in a more complex LSTM network that is trained to predict word translations. Although they do not represent different word senses separately, their method produces representations that depend on the context. In our work, the second-language signal is introduced only in the sense prediction component and is flexible—it can be defined in various ways and can be obtained from sentence-only alignments as a special case.

Conclusion

We have presented a method for learning multi-sense embeddings that performs sense estimation and context prediction jointly. Both mono- and bilingual information is used in the sense prediction during training. We have explored the model performance on a variety of tasks, showing that the bilingual signal improves the sense predictor, even though the crosslingual information is not available at test time. In this way, we are able to obtain word representations that are of better quality than the monolingually-trained multi-sense representations, and that outperform the Skip-Gram embeddings on intrinsic tasks. We have analyzed the model performance under several conditions, namely varying dimensionality, vocabulary size, amount of data, and size of the second-language context. For the latter parameter, we find that bilingual information is useful even when using the entire sentence as context, suggesting that sentence-only alignment might be sufficient in certain situations.

Acknowledgments

We would like to thank Jiwei Li for providing his tagger implementation, and Robert Grimm, Diego Marcheggiani and the anonymous reviewers for useful comments. The computational work was carried out on Peregrine HPC cluster of the University of Groningen. The second author was supported by NWO Vidi grant 016.153.327.

References