An Empirical Analysis of NMT-Derived Interlingual Embeddings and their Use in Parallel Sentence Identification

Cristina España-Bonet, Ádám Csaba Varga, Alberto Barrón-Cedeño, Josef van Genabith

Introduction

End-to-end neural machine translation systems (NMT) emerged in 2013 as a promising alternative to statistical and rule-based systems. Nowadays, they are the state of the art for language pairs with large amounts of parallel data and have nice properties that other paradigms lack. We highlight three: being a deep learning architecture, NMT does not require manually predefined features; it allows for the simultaneous training of systems across multiple languages; and it can provide zero-shot translations, i.e. translations for language pairs not directly seen in the training data .

Multilingual neural machine translation systems (ML-NMT) have interesting features. To perform multilingual translation, the network must project all the languages into the same common embedding space. In principle this space is multilingual, but the network does more than simply locating words according to their language and meaning independently. Previous studies suggest that the network locates words according to their semantics, irrespective of their language . That is somehow reinforced by the fact that zero-shot translation is possible (though at low quality). If that is confirmed, ML-NMT systems are learning a representation akin to an interligua for a source text and such interlingual embeddings could be used to assess cross-language similarity, among other applications.

In the past, the analysis of internal embeddings in NMT systems has been limited to visualisations; e.g., showing the proximity between semantically-similar representations. In the first part of this paper, we go beyond graphical analyses and search for empirical evidence of interlinguality. We address four specific research questions. RQ1: Whether the embedding learned by the network for a source text also depends on the target language. RQ2: How distinguishable representations of semantically-similar and semantically-distant sentence pairs are. RQ3: How close representations of sentence pairs within and across languages are. RQ4: How representations evolve throughout the training. These questions are addressed by means of statistics on cosine similarities between pairs of sentences both in a monolingual and a cross-language setting. In order to do that, we perform a large number of experiments using parallel and comparable data in Arabic, English, French, German, and Spanish (arar, enen, frfr, dede, and eses onwards). The second part of the paper is devoted to an application of the findings gathered in the first part: we explore the use of the ‘‘interlingua’’ representations to extract parallel sentences from comparable corpora. In this context, comparable corpora are text data on the same topic that are not direct translations of each other but may contain fragments that are translation equivalents; e.g., Wikipedia or news articles on the same subject in different languages. We evaluate the performance of supervised classification algorithms based upon our best contextual representations when discriminating between parallel and non-parallel sentences.

The article is organised as follows. Section 2 overviews the architecture of NMT systems. Section 3 describes the related work. Section 4 details the ML-NMT engines used in our analysis, presented in Section 5. Section 6 presents a use case: using the embeddings to identify parallel sentences. The conclusions are drawn in Section 7.

Background

State-of-the-art NMT systems use an encoder–decoder architecture with recurrent neural networks (RNN) . The encoder projects source sentences into an embedding space. The decoder generates target sentences from the encoder embeddings. Let s=(x1,…,xn)s=(x_{1},\dots,x_{n}) be a source sentence of length nn. The encoder encodes ss as a set of context vectorsCalled “annotation vectors” by , who use “context vectors” to designate the vectors after the attention mechanism., one per word:

Each component of this vector is obtained by concatenating the forward (h→i\overrightarrow{\mathbf{h}}_{i}) and backward (h←i\overleftarrow{\mathbf{h}}_{i}) encoder RNN hidden states:

where ff is a recurrent unit (GRU: Gated Recurrent Units in our experiments) and ri\mathbf{r}_{i} is the embedding space representation of the source word at position ii: ri=Wx⋅xi\mathbf{r}_{i}=\mathbf{W_{x}}\cdot\mathbf{x}_{i}.

The decoder generates the output sentence t=(y1,…,ym)t=(y_{1},\dots,y_{m}) of length mm on a word-by-word basis. The recurrent hidden state of the decoder zj\mathbf{z}_{j} is computed using its previous hidden state zj−1\mathbf{z}_{j-1}, as well as the previous continuous representation of the target word tj−1\mathbf{t}_{j-1} and the weighted context vector qj\mathbf{q}_{j} at time step jj:

where gg is a non-linear function and Wy\mathbf{W_{y}} is the matrix of the target embeddings. The weighted context vector qj\mathbf{q}_{j} is calculated by the attention mechanism as described in . Its function is to assign weights to the context vectors in order to selectively focus on different source words at different time steps of the translation. To this end, a single-hidden-layer feed-forward neural network is utilised to assign relevance scores (aa, as they can be interpreted as alignment scores) to the context vectors, which are then normalised into probabilities by the softmax function:

The attention mechanism takes the decoder’s previous hidden state zj−1\mathbf{z}_{j-1} and the context vector hi\mathbf{h}_{i} as inputs and weighs them up with the trainable weight matrices Wa\mathbf{W}_{a} and Ua\mathbf{U}_{a}, respectively. Finally, the probability of a target word is given by the following softmax activation :

where Wp1,Wp2,Wp3,Wo\mathbf{W}_{p1},\mathbf{W}_{p2},\mathbf{W}_{p3},\mathbf{W}_{o} are trainable matrices.

A number of papers extend this architecture to perform multilingual translation. They use multiple encoders and/or decoders with multiple or shared attention mechanisms . A simpler approximation considers exactly the same architecture as the one-to-one NMT for many-to-many NMT using multilingual data with some additional labelling. The authors in append the tag of the target language to the source-side sentences, forcing the decoder to translate to the appropriate language. The authors in also include tags specifying the language of every source word. Both papers show how these ML-NMT architectures can improve the translation quality between under-resourced language pairs and how they can be used for zero-shot translation. Given the premise that the encoder of an NMT system projects sentences into an embedding space, we can expect the encoder of ML-NMT systems to project sentences in different languages into a common (interlingual) embedding space. One of our aims is to study the characteristics of the internal representations of the encoder module in a ML-NMT system, and validate this assumption (see Section 5).

Related Work

There is some relevant previous research on qualitative studies of the NMT embedding space. The authors in show how a monolingual NMT encoder represents sentences with similar meaning close in the embedding space. They show graphically —with two instance sentences— that clustering by meaning goes beyond a bag-of-words understanding, and that differences caused by the order of the words are reflected in the representation. The authors in go one step further and visualise the internal space in a many-to-one language NMT system. A 2D-representation of some multilingual word embeddings from the encoder after training displays translations and related words close together. Experiments in provide visual evidence of a shared space for the attention vectors in a ML-NMT setup. Sentences with the same meaning but in different languages group together, except for zero-shot translations. When a language pair has not been seen during training, the embeddings lie in a different region of the space. In the authors study the representation generated by the attention vectors; i.e. the vectors showing the activations in the layer between encoder and decoder. The activations indicate which part of the source sentence is important during decoding to produce a particular chunk of the translation. Although the attention mechanism is shared across all the languages, the relevant chunks in the source sentence can vary depending on the target language.

In contrast to previous qualitative research, we focus on the context vectors: the concatenation of the hidden states of the forward and the backward network in the encoding module —right before applying the attention mechanism. Our goal goes beyond understanding the internal representations learned by the network: we aim at finding an appropriate representation to assess multilingual similarity. With this goal in mind, we look for a representation as target-independent as possible. Similarity assessment is at the core of many natural language processing and information retrieval tasks. Paraphrase identification is essentially similarity assessment and so is the task of plagiarism detection . In multi-document summarisation finding two highly-similar pieces of information in two texts may imply it is worth adding them into a good summary. In information retrieval, particularly in question answering , a high similarity between a document and an information request is a key factor of relevance. Similarity assessment also plays an important role in MT. It is essential in MT evaluation and, in the current cross-language setting, to identify parallel corpora to feed machine translation models . Efforts have been carried out to approach cross-language versions of these tasks using interlingua or multilingual representations instead of translating the texts into one common language . Still, such representations are usually hard to design. This is where our neural context vector NMT embedding representation comes into play. A multilingual encoder offers an environment where interlingua representations are learnt in a multilingual context. To some extent, it can be thought of as a generalisation of methods that project monolingual embeddings in two different languages into a common space to obtain bilingual word embeddings .

Recently, the authors in used the context vectors (CoVe) from a deep LSTM encoder in a bilingual NMT system to complement GloVe word vectors and improve the performance on several tasks: sentiment analysis, question classification, entailment, and question answering. In their case, the purpose is to exploit the context of a word rather than the interlingual nature of its representation. Finally, in a concurrent work, describe how joint multilingual sentence representations are learned with an NMT architecture with multiple encoders and/or decoders. In their case, a sentence is represented by the last state of an LSTM or by the max pooling after a BLSTM, depending on the nature of the encoder. They go beyond a visual analysis and evaluate the equivalence among representations of the same sentence in different languages by looking at the error when recovering multilingual parallel corpora.

NMT Systems Description

We carried out experiments with two multilingual many-to-many NMT engines trained with Nematus . As in and similarly to , we trained our systems on parallel corpora for several language pairs Lii–Ljj simultaneously, adding a tag in the source sentence to account for the target language ‘‘<2Ljj>’’ (e.g., <2ar> if the target language is Arabic). Table I shows the key parameters of the engines. Since our aim is to study the capability of NMT representations to characterise similar sentences within and across languages, we selected languages for which text similarity and/or translation test sets are available.

First, we build a ML-NMT engine for arar, enen, and eses. We trained the multilingual system for the 6 language pair directions on 56 M56\,M parallel sentences; see Table I(a). We used 1024 hidden units, which correspond to 2048-dimensional context vectors. We train system S1-w after cleaning and tokenising the texts. A second system called S1-l is trained on lemmatised sentences. We used MADAMIRA for tokenisation and lemmatisation in arar. For enen and eses we used Moses for tokenisation and IXA pipeline for lemmatisation. In both cases we employ a vocabulary of 60 K60\,K tokens plus 2 K2\,K for subword units, segmented using Byte Pair Encoding (BPE) .

Second, we build a ML-NMT engine for dede, frfr, enen, and eses. We train the system with data on 4 language pairs: dede–enen, frfr–enen, eses–enen and eses–frfr. Although some corpora exist for the remaining two (eses–dede and frfr–dede), we exclude them to study these pairs as instances of zero-shot translation. We obtain ∼\sim15 M15\,M parallel sentences per language pair —for dede–enen, we oversampled to reach that amount by tripling the original sentences; see Table I(b). We use a larger vocabulary in this engine: 80 K80\,K type tokens plus 2 K2\,K for BPE, as it involves one more language than in the first system. Only tokenisation with Moses is carried out. Regarding the number of hidden units, we experiment with three configurations: S2-w-d512, S2-w-d1024, and S2-w-d2048. In all cases we used sentences no longer than 50 tokens.

For evaluation, we consider three types of test sets. The source side is always the same and is aligned to a target set that contains either: (i) literal translations of the source, (ii) highly-similar sentences (both mono- and cross-language), and (iii) unrelated sentences (both mono- and cross-language). For arar, enen, and eses we build the three kinds of pairs out of the Semantic Textual Similarity Task at SemEval 2017 (STS 2017) http://alt.qcri.org/semeval2017/task1. The task asks to assess the similarity between two texts within the range $,where, where5standsforsemanticequivalence.Weextractthesubsetofsentenceswiththehighestsimilarity,4and5,anduse140sentencesoriginallyderivedfromtheMicrosoftResearchParaphraseCorpus(MSR),and203sentencesfromWMT2008http://www.statmt.org/wmt08/shared−evaluation−task.htmltobuildourfinaltestsetwith343sentences(subSTS2017).Thesedatawereavailableforstands for semantic equivalence. We extract the subset of sentences with the highest similarity, 4 and 5, and use 140 sentences originally derived from the Microsoft Research Paraphrase Corpus (MSR), and 203 sentences from WMT2008http://www.statmt.org/wmt08/shared-evaluation-task.html to build our final test set with 343 sentences (subSTS2017). These data were available forarandandenbutnotforbut not fores,sowemanuallytranslatedtheMSRpartofthecorpusinto, so we manually translated the MSR part of the corpus intoes,andgatheredthe, and gathered theescounterpartsofWMT2008fromtheofficialset.Withthisprocess,wegeneratedthetestwithtranslations(counterparts of WMT2008 from the official set. With this process, we generated the test with translations (trad)andhighlysimilarsentencepairs() and highly similar sentence pairs (semrel).Weshuffledoneofthesidesofthetestsettogeneratetheunrelatedpairs(). We shuffled one of the sides of the test set to generate the unrelated pairs (unrel$).

We use the test set from WMT2013 (newstest2013) to simultaneously evaluate the dede, frfr, enen, and eses experiments; the last edition that includes these four languages. The test set contains 3K3K sentences translated into the four languages. As before, we shuffle one of the sides to obtain the test set with unrelated sentence pairs, but we could not generate the equivalent set with highly similar pairs.

Context Vectors in Multilingual NMT Systems

The NMT architecture used for the experiments is the encoder–decoder model with recurrent neural networks and attention mechanism described in Section 3, as implemented in Nematus. We use the sum of the context vector associated to every word (Eq. 1) at a specific point of the training as the representation of a source sentence ss:

This representation depends on the length of the sentence. However, we stick to this definition rather than using a mean over words because the length of the sentences is a feature one might take into account, since sentences with similar meaning tend to have similar lengths. Given sentence s1s_{1} represented by Cs1\mathbf{C}_{s_{1}} and sentence s2s_{2} represented by Cs2\mathbf{C}_{s_{2}}, we can estimate their similarity by means of the cosine measure:

By using this similarity measure we cancel the effect of the length of the sentence on the similarity between pairs but not on the representation of the sentence itself.We explored alternative sentence representations (sum vs mean) and similarity measures (cosine vs modified versions of weighted Jaccard similarity, and Kullback–Leibler and Jensen–Shannon divergences). Cosine over the mean resulted in the best performance as measured by the correlation with human judgements on similarity assessments.

Context vectors are high-dimensional structures: commonly used 1024-dimensional hidden layers lead to 2048-dimensional context vectors. In order to get a first impression on the behaviour of the embeddings, we project the vectors for a set of sentences into a 2D space using t-Distributed Stochastic Neighbour Embedding (t-SNE) .

Figure 1 shows 21 sentences extracted from the trial set of STS 2017 for this purpose and the relations between triplets. Some triplets are related semantically; e.g., a triplet with the element ‘‘Mandela’s condition has improved’’ is semantically related to the triplet with the element ‘‘Mandela’s condition has worsened over past 48 hours’’. In a real multilingual space, one would expect sentences within a triplet to lie together and sentences within related triplets to be close but, as Figure 2 shows, the range of behaviours may be diverse. The plot shows the evolution of the context vectors for these 21 sentences throughout the training (central panel), paying special attention to an early (left panel) and a late stage (right panel).

At the beginning of the training, enen and eses sentences in the same triplet (same colour) lie close together and even overlap for some triplets; e.g., t4t4 and t7t7. This is an effect of having a representation that depends on the length of the sentence: the elements in t4t4 and t7t7 not only share some vocabulary, but also have very similar lengths. Arabic sentences remain together, almost irrespective of their meaning. One has to take into account that enen and eses are closer between them than to arar. Meanwhile, arar is closer to eses than to enen. At this early training stage, the closer languages already cluster together (enen and eses) and sentences can be grouped according to their semantics, but the most distant language (arar) is not in the same stage yet. At this stage, pairs where both sentences are written in arar are considered more similar, even if they are semantically very different (also compared to semantically similar sentences across languages); sentence s9s9 is closer to s14s14 (another sentence in arar with similar length) than to s7s7 (a strict and longer translation of s9s9 into eses).

As training continues, arar sentences spread through the space and slowly tend to join their counterparts in the other languages. English and Spanish sentences also move apart towards a more general interlingua position. That is, there is a flow from near to overlapping locations for translations of the same sentence towards locations grouped by topic, irrespective of the language (e.g., see the evolution of the related triplets t6t6 and t7t7). This evolution must be considered if one wants to use context vectors as a semantic representation of a sentence: representations at different points of the training process might be useful for different tasks. For instance, as shown in the following subsections, using context vectors from a converged NMT training is beneficial to assess similarity, but one only needs to run some iterations to have appropriate vectors to identify parallel sentences.

However, not all the triplets show the expected behaviour. While at every iteration the sentences in the triples in t1t1 and t5t5 each move closer together, and therefore behave as expected, the sentences in t6t6 move further away from each other (notice that this triplet has the longest sentences and the highest length variation). A more systematic study is necessary in order to be able to draw strong conclusions. In the following sections we conduct such a study and draw conclusions quantitatively, rather than only qualitatively.

2 Source vs Source–Target Semantic Representations

The training of the ML-NMT systems involves one-to-many instances. That is, for the same source language L1 one has different examples of translations into L2, L3, or L4. A first question one can address given this setup is whether the interpretation of a source sentence learnt by the network depends on the language it is going to be translated into or not. In a truly interlingual space, such representations should be the same, or at least very close. To test this, we compute the cosine similarity between the representation of a source sentence ss when it is translated with the same engine into two different languages Lii and Ljj:

Sentence representations are extracted with engine S1-w for {arar, enen, eses} on subSTS2017 data and with engine S2-w-d1024 for {dede, enen, eses, frfr} on newstest2013. Afterwards, we compute the mean over all the sentences in a test set.

Table II shows the results. The similarities are close to 11 in all cases, a number that would indicate that the representations are fully equivalent, and are compatible with 11 within a 2σ2\sigma interval. Although the differences among languages and test sets are not significant at that level, some general trends are observed. Despite the fact that the similarity between instances of the same sentence is not 11, it is larger than the similarity between closely related sentences when translated into the same language (see Section 5.3); i.e. we can identify a sentence by a unique representation. Also notice that there is no difference when we translate into a language without any direct parallel data (zero-shot translation): system S2-w-d1024 had no data for eses–dede and frfr–dede, but the similarities involving these pairs (starred in Table II) are not statistically-significantly different from those involving eses–frfr and eses–enen, for example.

Finally, we can strengthen the correlation of the relatedness between languages and the closeness of the internal representations observed also via the first graphical analysis. The representation of an arar sentence when translated into enen or eses is almost the same (sim=0.97±0.05sim=0.97\pm 0.05), but the difference in the representation of an eses sentence when translated into arar or enen is the largest one (sim=0.91±0.05sim=0.91\pm 0.05) due to the disparity between arar and enen. The same effect is observed in {dede, frfr, enen, eses} at a lower degree when making the distinction between {frfr, eses} and {dede, enen} as two groups of ‘‘close’’ languages.

3 Representations throughout Training

During training, the network learns the most appropriate representation of words/sentences in order to be translated, so the embeddings themselves evolve over time. As seen in the graphical analysis (Section 5.1), it is interesting to follow this evolution and examine how sentences are grouped together depending on their language and semantics. Hence, we analyse in parallel an engine trained on lemmatised sentences (S1-l) and one trained on tokenised sentences (S1-w). The rationale is that the vocabulary in the lemmatised system is smaller and therefore can be better covered by the 60 K60\,K NMT fixed vocabulary during training. Still, the ambiguity becomes higher, which could damage the quality of the representations.

Table III shows the results. At the beginning of the training process, after having seen 4⋅1064\cdot 10^{6} sentences only, the results are still very much dependent on the language. Translations in arar–eses have a similarity of 0.81±0.040.81\pm 0.04, whereas translations in arar–enen have a similarity of 0.44±0.070.44\pm 0.07 (first row for system S1-lemmas). Perhaps for this reason monolingual pairs show higher similarity values than cross-language pairs, even for unrelated sentences (sim=0.70±0.09sim=0.70\pm 0.09 for arar and sim=0.73±0.09sim=0.73\pm 0.09 for enen). Nevertheless, within a language pair the system is already aware of the meaning of the sentences: cosine similarities are the highest for translations (tradtrad), slightly lower for semantically related sentences (semrelsemrel) and significantly lower for unrelated sentences (unrelunrel). The difference between the mean similarities obtained for translations and unrelated sentences,

shows that, already at this point, parallel sentences can be identified and located in the multilingual space, even though the similarity for translations is in general far from 11 and the similarity for unrelated sentences is far from . In the worst-case scenario, S1-lemmas for arar–enen, Δtr−ur=0.16±0.11\Delta_{\rm tr-ur}=0.16\pm 0.11, so translations are clearly distinguished at 1σ1\sigma level. In other words, if we look at the distance of one sentence to its translation and to all the unrelated sentences in the unrelunrel set, only in 1.6%1.6\% of the cases an unrelated sentence is closer or at the same distance as the translation. This number diminishes to 0.6%0.6\% in the best case scenario (S1-words for enen–eses). Also at this starting point, sentences lie closer together irrespective of their meaning in the lemmatised system than in the tokenised one. Similarities are always higher for S1-l than for its counterpart in S1-w. The separation between translations and unrelated sentences is always more important in the S1-w (Δtr−ur\Delta_{\rm tr-ur} is higher). This is true all along the training process, supporting the hypothesis that the ambiguity introduced by the lemmatisation damages the representativeness of the embeddings.

When the training process has covered 28⋅10628\cdot 10^{6} sentences, half an epoch for this system, the difference among languages diminishes. Now sentences lie closer together in the tokenised system than in the lemmatised one, irrespective of their meaning. From this point onwards, this trait is maintained. Although all similarities keep going down throughout the training, even for translations, Δtr−ur\Delta_{\rm tr-ur} remains almost constant. The maximum value for this difference is found after one epoch (∼56⋅106\sim 56\cdot 10^{6} sentences) for all the cross-language pairs in the tokenised system. In this case, Δtr−ur\Delta_{\rm tr-ur} is 0.34±0.120.34\pm 0.12 for arar–enen, 0.33±0.130.33\pm 0.13 for arar–eses and 0.43±0.120.43\pm 0.12 for enen–eses. Again, the distinction is the clearest for the closest language pair and diminishes when arar is involved, mainly because translations involving arar are more difficult to detect (the mean similarity between enen–eses translations is 0.74±0.060.74\pm 0.06; 0.61±0.080.61\pm 0.08 for arar–enen).

As Table IV shows, analogous conclusions can be drawn from the {dede, frfr, enen, eses} engine. The maximum distinction between related and unrelated sentences Δtr−ur\Delta_{\rm tr-ur} is found after ∼56⋅106\sim 56\cdot 10^{6} sentences, half an epoch in this case, even though the difference was well established at one third of an epoch. Δtr−ur\Delta_{\rm tr-ur} is 0.3±0.10.3\pm 0.1 when dede is involved (dede–enen, dede–eses, dede–frfr) and 0.4±0.10.4\pm 0.1 when not (enen–eses, enen–frfr, eses–frfr). The difference is mostly given by the similarity between translations, which is higher when dede is not concerned.

Notice that this optimal point does not correspond to the optimal point regarding translation quality. Figure 3 displays the progression of the BLEU score along training for the enen2eses translation. The dashed vertical line indicates the iteration where Δtr−ur\Delta_{\rm tr-ur} is maximum. At this time, the engine is still learning, as reflected by the fact that the translation quality is clearly increasing. Another interesting observation is that the expressiveness of the embeddings does not depend on their dimensionality. Context vectors with 1024 dimensions (S2-w-d512), 2048 dimensions (S2-w-d1024) and 4096 dimensions (S2-w-d2048), lead to similar figures for similarity values between pairs of sentences. At the beginning of the training, S2-w-d1024 gives slightly better representations than the other two systems, but this difference is narrowed when the training evolves. The training time almost doubles when doubling the dimensionality of the hidden layer, but this higher capacity does not result in a better description of the data. Indeed, 4096-dimensional vectors perform worse than the 1024-dimensional ones at all the training stages. However, translation quality does depend on the size of the hidden layer and, in our experiments, S2-w-d2048 performs better than the lower-dimensional systems.

4 Similarity Assessments

Up to now, we have mostly analysed how similar (tradtrad) and dissimilar (unrelunrel) sentences behave across languages and during training. The degree of similarity was left aside because the tradtrad and semrelsemrel test sets are too alike to draw statistically-significant conclusions in that setting. To do so, we evaluate the use of context vectors as a feature to assess similarities in the STS framework. In this case, we use all the available test sets for the 2017 evaluation campaign with sentence pairs ranging from completely unrelated sentences (score 0) to semantic equivalents (score 5). Only the subset of most similar sentences had been used in the earlier experiment (scores 4 and 5).

Table V shows the Pearson correlation between the predictions given by the context vectors of S1-w and S1-l and human assessments for five language pairs. Observing the evolution through training by taking a shot at four different points, the correlation increases with the number of iterations for all the language pairs and systems. In this fine-grained task, the internal representation improves in parallel to the translation quality. As before, the system with words is better than the one with lemmas with the only exception of arar–enen. A reason could be the low initial similarity for semantically equivalent sentences (tradtrad) for this pair with the S1-w system (0.26±0.100.26\pm 0.10). The initial point seems to be relevant for the final performance; i.e. the relative improvement from epoch to epoch for all language pairs is very similar, but the final performance seems to be conditioned to the quality of the initial representations. The performance in the monolingual tracks is always higher than in the cross-language ones, and the difference at the end of the training is proportional to the difference at the beginning. The study of how a proper initialisation of the input word embeddings could alleviate this disparity deserves future research.

The comparison with word vector embeddings obtained with the word2vec skip-gram model is specially interesting. We estimated 300 (WE-d300-nmt) and 1024 (WE-d1024-nmt) dimensional word embeddings with the same corpus used to train the NMT systems (adding monolingual corpora did not improve the results). When sentences belong to different languages, we translate them into enen and use the embeddings estimated for enen. As done with context vectors, the similarity between sentences is assessed by the cosine of the summed embeddings. Higher-dimensional word embeddings outperform the 300-dimensional ones in the task. Yet, even with the 1024-dimensional word embeddings, the performance is far from that obtained with context vectors —between 0.040.04 and 0.210.21 points lower (see last block of Table V).

Use Case: Parallel Sentence Extraction

The previous section showed how ML-NMT context vectors can be used as a representation to calculate sensitive similarities between sentences with the potential to distinguish translations from non-translations and even translations from pairs with similar meaning. Among other applications, we can use the representations learned when mapping parallel sentences —the NMT system training— to detect new parallel pairs. Now we use a semantic similarity measure based on the context vectors obtained with the NMT system of Section 5 to extract parallel sentences and study its performance compared to other measures. Our translation engine is the ML-NMT {dede, frfr, enen, eses} system described in Section 4. After the conclusions gathered in Section 5, we use system S2-w-d512 after half an epoch of training to extract the context vectors. This system gives the best trade-off between speed (low-dimensional vectors are extracted faster) and dissociation between translations and unrelated sentences, as this is the training point where the difference Δtr−ur\Delta_{\rm tr-ur} is maximum.

In order to perform a complete analysis, we consider five complementary measures to context vectors and test different scenarios. We borrow two well-known representations from cross-language information retrieval to account for syntactic features by means of cosine similarities: (i) character nn-grams with n=n= and (ii) pseudo-cognates. From a natural language point of view, cognates are ‘‘words that are similar across languages’’ . We relax the concept and consider as pseudo-cognates any words in two languages that share prefixes. To do so, tokens shorter than four characters are discarded, unless they contain non-alphabetical characters. The resulting tokens are cut down to four characters . The preprocessing consists only of casefolding and punctuation/diacritics removal. For the character nn-gram measure, we also remove spaces to better account for compounds in German. We also include general features at sentence level such as (iii) token and (iv) character counts, and (v) the length factor measure .

We test three different scenarios to observe the effect of context vectors when extracting sentence pairs and compare them against the other standard characterisations:

only the set of five complementary measures, and

For each scenario, we learn a binary classifier on annotated data. We use the dede–enen and frfr–enen training corpora provided for the shared task on identifying parallel sentences in comparable corpora at BUCC 2017 .https://comparable.limsi.fr/bucc2017/bucc2017-task.html This set contains 1.5 M1.5\,M sentences from Wikipedia and News Commentary from which 20 K20\,K are aligned sentence pairs. Negative indexes are manually added by randomly pairing up the same amount of non-matching pairs to build a balanced data set. We use 35 K35\,K instances from the full set for training and evaluating classifiers with 10-fold cross-validation, 4 K4\,K instances for training an ensemble of the best classifiers and 1 K1\,K instances for held-out testing purposes.

For ctxctx, where only the context vector similarities are considered, the problem can be reduced to finding a suitable decision threshold. To this end, similarity values between the lowest value among positive examples and the highest value among negative samples are incrementally increased by a step size of 0.005 and the threshold giving the highest accuracy on the training set is selected. With this methodology, we obtain a threshold t=0.43t=0.43 for dede–enen leading to an accuracy of 97.2%97.2\%, and 0.410.41 for frfr–enen with an accuracy of 97.4%97.4\%. These values are slightly lower than the ones reported in Table IV, but consistent with them. The thresholds in both cases depend on the language pair, but the fact that we are working with an interlingua representation makes the differences minimal. In such a case, one can estimate a joint threshold for the full training set in dede–enen and frfr–enen and later use this decision boundary for other language pairs. If we do the search on the joint datasets the best threshold is t=0.43t=0.43 leading to an accuracy of 97.2%97.2\% in the training set.

We have 7 and 8 features in compcomp and allall and employ supervised classifiers rather than a threshold estimation: support vector machines (SVM) with RBF kernel and gradient boosting (GB) on the deviance objective function with 10-fold cross-validation. A soft voting ensemble (Ens.) of SVM and GB is trained to obtain the final model.We use the Python scikit-learn package: http://scikit-learn.org

Table VI shows precision (P), recall (R) and F1 scores for the three scenarios. Notice that a greedy threshold search is better than any of the machine learning counterparts when only context vectors are used, but differences are not significant. The greedy search on the context vector similarities gives a better F1 on the held-out test set than an ensemble of SVM and GB operating only the set of additional features with almost no knowledge of semantics. As we argued in the previous section, translations and non-translations are clearly differentiated by a cosine similarity of the context vectors for these languages pairs, as the difference between the mean similarities of translations and unrelated texts is much higher than its uncertainty (Δtr−ur\Delta_{\rm tr-ur}=0.36±0.14=0.36\pm 0.14 for dede–enen, and 0.41±0.140.41\pm 0.14 for frfr–enen). This clear distinction in the similarities is translated into an F1=98.2%{}_{1}=98.2\% in the task of parallel sentence identification.

Due to its interlingual nature, our feature behaves equally well for both language pairs and improves in the multilingual setting (Table VI, joint columns). By contrast, the set of complementary features depends on the language pair and shows a performance drop for dede–enen. For this reason, the results in the multilingual setting are always worse than in the bilingual one. This fact is inherited in the allall scenario, where the classification for the joint corpus obtains F1=98.9%{}_{1}=98.9\%, which is lower than the one obtained for frfr–enen alone (F1=99.3%{}_{1}=99.3\%). Nevertheless, semantic and syntactic similarity features are complementary and the combination of all similarity measures slightly improves precision, recall and F1 in the multilingual setting. It is worth noting the high recall derived from the context vectors, which reaches 100%100\% for frfr–enen and falls to 98.1%98.1\% for the joint data, being still 6.5 points higher than for the compcomp features.

Conclusions

In this article we provide evidence of the interlingual nature of the context vectors generated by a multilingual neural machine translation system and study their power in the assessment of mono- and cross-language similarity. Comparisons with word vectors show that context vectors are able to capture better the semantics in the two settings.

The study addresses four main research questions, introduced in Section 1. Regarding RQ1, we investigate how the representation of a sentence varies in order to be accommodated to a particular target language and observe that the difference is negligible, even though it grows when we consider distant target languages, such as Arabic and English. Even in these cases, the representation of a sentence is unique enough as closely related sentences have a lower similarity than different instances of the same sentence. RQ2: The results also show that the context vectors are able to differentiate among sentences with identical, similar, and different meaning across different languages —Arabic, English, French, German, and Spanish. The difference between translations and non-translations can be established at least at 1σ1\sigma level for all the pairs. As a direct application, we identify parallel sentences in comparable corpora, obtaining F1=98.2%F_{1}=98.2\% on data of the shared task at BUCC 2017. The correlation of the cosine between context vectors with human judgements on continuous similarity assessments ranges in [0.4,0.8][0.4,0.8], always higher than the ones obtained for word vectors models: [0.3,0.6][0.3,0.6]. RQ3: The language dependence is not completely lost in the representations. In the latter experiment, correlations in the cross-language tasks are lower than in the monolingual ones, but in both cases related and unrelated sentence pairs are clearly distinguishable within the variance. RQ4: Our training-evolution experiments reveal that the first feature to locate a sentence in the multilingual space is its language but, after only ∼\sim4⋅1064\cdot 10^{6} training sentences, the model is already aware of the semantics. As the training evolves, the difference between translations and unrelated sentences grows till reaching a plateau when the system has been trained on ∼\sim40⋅10640\cdot 10^{6} sentences. Vectors at early training are therefore already adequate for identifying parallel sentences, whereas the optimal ones for fine-grained similarity assessments and translation require further training.

Given these conclusions, several research avenues are worth exploring in the future. The disparity in the performance of mono- and cross-language similarity assessment tasks triggers a question on how relevant the initialisation of the embeddings is. Could the results be improved with initialisations of the word embeddings other than random? The answer can be extended and exploited in other natural language processing tasks, in the same philosophy as , but in a multilingual setting. Additionally, similar studies using other NMT architectures could help in better understanding the insights of the learning.

Acknowledgments

This work was partially funded by the Leibniz Gemeinschaft via the SAW-2016-ZPID-2 project and by the European Union Horizon 2020 research and innovation programme under grant agreement 645452 (QT21). The research of A. Barrón-Cedeño is carried out in the framework of the Interactive sYstems for Answer Search project (IYAS) at QCRI. We thank the anonymous reviewers for their interesting insights that helped to improve this paper.

References