Generalized Data Augmentation for Low-Resource Translation

Mengzhou Xia, Xiang Kong, Antonios Anastasopoulos, Graham Neubig

Introduction

The task of Machine Translation (MT) for low resource languages (lrls) is notoriously hard due to the lack of the large parallel corpora needed to achieve adequate performance with current Neural Machine Translation (NMT) systems Koehn and Knowles (2017). A standard practice to improve training of models for an lrl of interest (e.g. Azerbaijani) is utilizing data from a related high-resource language (hrl, e.g. Turkish). Both transferring from hrl to lrl Zoph et al. (2016); Nguyen and Chiang (2017); Gu et al. (2018) and joint training on hrl and lrl parallel data Johnson et al. (2017); Neubig and Hu (2018) have shown to be effective techniques for low-resource NMT. Incorporating data from other languages can be viewed as one form data augmentation, and particularly large improvements can be expected when the hrl shares vocabulary or is syntactically similar with the lrl Lin et al. (2019). Simple joint training is still not ideal, though, considering that there will still be many words and possibly even syntactic structures that will not be shared between the most highly related languages. There are model-based methods that ameliorate the problem through more expressive source-side representations conducive to sharing Gu et al. (2018); Wang et al. (2019), but they add significant computational and implementation complexity.

In this paper, we examine how to better share information between related lrl and hrls through a framework of generalized data augmentation for low-resource MT. In our basic setting, we have access to parallel or monolingual data of an lrl of interest, its hrl, and the target language, which we will assume is English. We propose methods to create pseudo-parallel lrl data in this setting. As illustrated in Figure 1, we augment parallel data via two main methods: 1) back-translating from eng to lrl or hrl; 2) converting the hrl-eng dataset to a pseudo lrl-eng dataset.

In the first thread, we focus on creating new parallel sentences through back-translation. Back-translating from the target language to the source Sennrich et al. (2016) is a common practice in data augmentation, but has also been shown to be less effective in low-resource settings where it is hard to train a good back-translation model Currey et al. (2017). As a way to ameliorate this problem, we examine methods to instead translate from the target language to a highly-related hrl, which remains unexplored in the context of low-resource NMT. This pseudo-hrl-eng dataset can then be used for joint training with the lrl-eng dataset.

In the second thread, we focus on converting an hrl-eng dataset to a pseudo-lrl-to-eng dataset that better approximates the true lrl data. Converting between hrls and lrls also suffers from lack of resources, but because the lrl and hrl are related, this is an easier task that we argue can be done to some extent by simple (or unsupervised) methods.This sort of pseudo-corpus creation was examined in a different context of pivoting for SMT De Gispert and Marino (2006), but this was usually done with low-resource source-target language pairs with English as the pivot. In our proposed method, for the first step, we substitute hrl words on the source side of hrl parallel datasets with corresponding lrl words from an induced bilingual dictionary generated by mapping word embedding spaces Xing et al. (2015); Lample et al. (2018b). In the second step, we further attempt translate the pseudo-lrl sentences to be closer to lrl ones utilizing an unsupervised machine translation framework.

We conduct a thorough empirical evaluation of data augmentation methods for low-resource translation that take advantage of all accessible data, across four language pairs.

We explore two methods for translating between related languages: word-by-word substitution using an induced dictionary, and unsupervised machine translation that further uses this word-by-word substituted data as input. These methods improve over simple unsupervised translation from hrl to lrl by more than 2 to 10 BLEU points.

Our proposed data augmentation methods improve over standard supervised back-translation by 1.5 to 8 BLEU points, across all datasets, and an additional improvement of up to 1.1 BLEU points by augmenting from both eng monolingual data, as well as hrl-eng parallel data.

A Generalized Framework for Data Augmentation

In this section, we outline a generalized data augmentation framework for low-resource NMT.

Given an lrl of interest and its corresponding hrl, with the goal of translating the lrl to English, we usually have access to 1) a limited-sized lrl-eng parallel dataset {S\textscle,T\textscle}\{\mathcal{S}_{\textsc{le}},\mathcal{T}_{\textsc{le}}\}; 2) a relatively high-resource hrl-eng parallel dataset {S\textsche,T\textsche}\{\mathcal{S}_{\textsc{he}},\mathcal{T}_{\textsc{he}}\}; 3) a limited-sized lrl-hrl parallel dataset {S\textschl,T\textschl}\{\mathcal{S}_{\textsc{hl}},\mathcal{T}_{\textsc{hl}}\}; 4) large monolingual datasets in lrl M\textscl\mathcal{M}_{\textsc{l}}, hrl M\textsch\mathcal{M}_{\textsc{h}} and English M\textsce\mathcal{M}_{\textsc{e}}.

2 Augmentation from English

The first two options for data augmentation that we explore are typical back-translation approaches:

eng-lrl We train an eng-lrl system and back-translate English monolingual data to lrl, denoted by \{\hat{\mathcal{S}}_{\textsc{e\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}},\mathcal{M}_{\textsc{e}}\}.

eng-hrl We train an eng-hrl system and back-translate English monolingual data to hrl, denoted by \{\hat{\mathcal{S}}_{\textsc{e\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}h}},\mathcal{M}_{\textsc{e}}\}.

Since we have access to lrl-eng and hrl-eng parallel datasets, we can train these back-translation systems Sennrich et al. (2016) in a supervised fashion. The first option is the common practice for data augmentation. However, in a low-resource scenario, the created lrl data can be of very low quality due to the limited size of training data, which in turn could deteriorate the lrl ⁣)\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}eng translation performance. As we show in Section 5, this is indeed the case.

The second direction, using hrl back-translated data for lrl ⁣)\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}eng translation, has not been explored in previous work. However, we suggest that in low-resource scenarios it has potential to be more effective than the first option because the quality of the generated hrl data will be higher, and the hrl is close enough to the lrl that joint training of a model on both languages will likely have a positive effect.

3 Augmentation via Pivoting

Using hrl-eng data improves lrl-eng translation because (1) adding extra eng data improves the target-side language model, (2) it is possible to share vocabulary (or subwords) between languages, and (3) because the syntactically similar hrl and lrl can jointly learn parameters of the encoder. However, regardless of how close these related languages might be, there still is a mismatch between the vocabulary, and perhaps syntax, of the hrl and lrl. However, translating between hrl and lrl should be an easier task than translating from English, and we argue that this can be achieved by simple methods.

Hence, we propose ``Augmentation via Pivoting" where we create an lrl-eng dataset by translating the source side of hrl-eng data, into the lrl. There are again two ways in which we can construct a new lrl-eng dataset:

hrl-lrl We assume access to an hrl-eng dataset. We then train an hrl-lrl system and convert the hrl side of S\textsche\mathcal{S}_{\textsc{he}} to lrl, creating a \{\hat{\mathcal{S}}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}},\mathcal{T}_{\textsc{he}}\} dataset.

eng-hrl-lrl Exactly as before, except that the hrl-eng dataset is the result of back-translation. That means that we have first converted English monolingual data M\textsce\mathcal{M}_{\textsc{e}} to \hat{\mathcal{S}}_{\textsc{e\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}h}}, and then we convert those to the lrl, creating a dataset \{\hat{\mathcal{S}}_{\textsc{e\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}hh\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}},\mathcal{M}_{\textsc{e}}\}.

Given a lrl-hrl dataset {S\textsclh,T\textsclh}\{\mathcal{S}_{\textsc{lh}},\mathcal{T}_{\textsc{lh}}\} one could also train supervised back-translation systems. But we still face the same problem of data scarcity, leading to poor quality of the augmented datasets. Based on the fact that an lrl and its corresponding hrl can be similar in morphology and word order, in the following sections, we propose methods to convert hrl to lrl for data augmentation in a more reliable way.

LRL-HRL Translation Methods

In this section, we introduce two methods for converting hrl to lrl for data augmentation.

Mikolov et al. (2013) show that the word embedding spaces share similar innate structure over different languages, making it possible to induce bilingual dictionaries with a limited amount of or even without parallel data Xing et al. (2015); Zhang et al. (2017); Lample et al. (2018b). Although the capacity of these methods is naturally constrained by the intrinsic properties of the two mapped languages, it's more likely to create a high-quality bilingual dictionary for two highly-related languages. Given the induced dictionary, we can substitute hrl words with lrl ones and construct a word-by-word translated pseudo-lrl corpus.

We use a supervised method to obtain a bilingual dictionary between the two highly-related languages. Following Xing et al. (2015), we formulate the task of finding the optimal mapping between the source and target word embedding spaces as the Procrustes problem Schönemann (1966), which can be solved by singular value decomposition (SVD):

where XX and YY are the source and target word embedding spaces respectively.

As a seed dictionary to provide supervision, we simply exploit identical words from the two languages. With the learned mapping WW, we compute the distance between mapped source and target words with the CSLS similarity measure Lample et al. (2018b). Moreover, to ensure the quality of the dictionary, a word pair is only added to the dictionary if both words are each other's closest neighbors. Adding an lrl word to the dictionary for every hrl word results in relatively poor performance due to noise as shown in Section 5.3.

Corpus Construction

Given an hrl-eng {S\textsche,T\textsche}\{\mathcal{S}_{\textsc{he}},\mathcal{T}_{\textsc{he}}\} or a back-translated \{\hat{\mathcal{S}}_{\textsc{e\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}h}},\mathcal{M}_{\textsc{e}}\} dataset, we substitute the words in S\textsche\mathcal{S}_{\textsc{he}} with the corresponding lrl ones using our induced dictionary. Words not in the dictionary are left untouched. By injecting lrl words, we convert the original or augmented hrl data into pseudo-lrl, which explicitly increases lexical overlap between the concatenated lrl and hrl data. The created datasets are denoted by \{\hat{\mathcal{S}}^{w}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}e}},\mathcal{T}_{\textsc{he}}\} and \{\hat{\mathcal{S}}^{w}_{\textsc{e\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}hh\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}},\mathcal{M}_{\textsc{e}}\} where ww denotes augmentation with word substitution.

2 Augmentation with Unsupervised MT

Although we assume lrl and hrl to be similar with regards to word morphology and word order, the simple word-by-word augmentation process will almost certainly be insufficient to completely replicate actual lrl data. A natural next step is to further convert the pseudo-lrl data into a version closer to the real lrl. In order to achieve this in our limited-resource setting, we propose to use unsupervised machine translation (umt).

Unsupervised Neural Machine Translation Artetxe et al. (2018); Lample et al. (2018a, c) makes it possible to translate between languages without parallel data. This is done by coupling denoising auto-encoding, iterative back-translation, and shared representations of both encoders and decoders, making it possible for the model to extend the initial naive word-to-word mapping into learning to translate longer sentences.

Initial studies of umt have focused on data-rich, morphologically simple languages like English and French. Applying the umt framework to low-resource and morphologically rich languages is largely unexplored, with the exception of Neubig and Hu (2018) and Guzmán et al. (2019), showing that umt performs exceptionally poorly between dissimilar language pairs with BLEU scores lower than 1. The problem is naturally harder for morphologically rich lrls due to two reasons. First, morphologically rich languages have a higher proportions of infrequent words Chahuneau et al. (2013). Second, even though still larger than the respective parallel datasets, the size of monolingual datasets in these languages is much smaller compared to hrls.

Modified Initialization

As pointed out in Lample et al. (2018c), a good initialization plays a critical role in training nmt in an unsupervised fashion. Previously explored initialization methods include: 1) word-for-word translation with an induced dictionary to create synthetic sentence pairs for initial training Lample et al. (2018a); Artetxe et al. (2018); 2) joint Byte-Pair-Encoding (BPE) for both the source and target corpus sides as a pre-processing step. While the first method intends to give a reasonable prior for parameter search, the second method simply forces the source and target languages to share the same subword vocabulary, which has been shown to be effective for translation between highly related languages.

Inspired by these two methods, we propose a new initialization method that uses our word substitution strategy (§3.1). Our initialization is comprised of a sequence of three steps:

First, we use an induced dictionary to substitute hrl words in M\textsch\mathcal{M}_{\textsc{h}} to lrl ones, producing a pseudo-lrl monolingual dataset M^\textscl\hat{\mathcal{M}}_{\textsc{l}}.

Second, we learn a joint word segmentation model on both M\textscl\mathcal{M}_{\textsc{l}} and M^\textscl\hat{\mathcal{M}}_{\textsc{l}} and apply it to both datasets.

Third, we train a NMT model in an unsupervised fashion between M\textscl\mathcal{M}_{\textsc{l}} and M^\textscl\hat{\mathcal{M}}_{\textsc{l}}. The training objective L\mathcal{L} is a weighted sum of two loss terms for denoising auto-encoding and iterative back-translation:

where u∗u^{*} denotes translations obtained with greedy decoding, CC denotes a noisy manipulation over input including dropping and swapping words randomly, λ1\lambda_{1} and λ2\lambda_{2} denotes the weight of language modeling and back translation respectively.

In our method, we do not use any synthetic parallel data for initialization, expecting the model to learn the mappings between a true lrl distribution and a pseudo-lrl distribution. This takes advantage of the fact that the pseudo-lrl is naturally closer to the true lrl than the hrl is, as the injected lrl words increase vocabulary overlap.

Corpus Construction

3 Why Pivot for Back-Translation?

Pivoting through an hrl in order to convert English to lrl will be a better option compared to directly translating eng to lrl under the following three conditions: 1) hrl and lrl are related enough to allow for the induction of a high-quality bilingual dictionary; 2) There exists a relatively high-resource hrl-eng dataset; 3) A high-quality lrl-eng dictionary is hard to acquire due to data scarcity or morphological distance.

Essentially, the direct eng ⁣)\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}lrl back-translation may suffer from both data scarcity and morphological differences between the two languages. Our proposal breaks the process into two easier steps: eng ⁣)\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}hrl translation is easier due to the availability of data, and hrl ⁣)\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}lrl translation is easier because the two languages are related.

A good example is the agglutinative language of Azerbaijiani, where each word may consist of several morphemes and each morpheme could possibly map to an English word itself. Correspondences to (also agglutinative) Turkish, however, are easier to uncover. To give a concrete example, the Azerbijiani word ``düşünclrim'' can be fairly easily aligned to the Turkish word ``düşüncelerim'' while in English it corresponds to the phrase ``my thoughts'', which is unlikely to be perfectly aligned.

Experimental Setup

We use the multilingual TED corpus Qi et al. (2018) as a test-bed for evaluating the efficacy of each augmentation method. We conduct extensive experiments over four low-resource languages: Azerbaijani (aze), Belarusian (bel), Galician (glg), and Slovak (slk), along with their highly related languages Turkish (tur), Russian (rus), Portuguese (por), and Czech (ces) respectively. We also have small-sized lrl-hrl parallel datasets, and we download Wikipedia dumps to acquire monolingual datasets for all languages.

The statistics of the parallel datasets are shown in Table 1. For aze, bel and glg, we use all available Wikipedia data, while for the rest of the languages we sample a similar-sized corpus. We sample 2M/200K English sentences from Wikipedia data, which are used for baseline umt training and augmentation from English respectively.

2 Pre-processing

We train a joint sentencepiecehttps://github.com/google/sentencepiece model for each lrl-hrl pair by concatenating the monolingual corpora of the two languages. The segmentation model for English is trained on English monolingual data only. We set the vocabulary size for each model to 20K. All data are then segmented by their respective segmentation model.

We use FastTexthttps://github.com/facebookresearch/fastText to train word embeddings using M\textscl\mathcal{M}_{\textsc{l}} and M\textsch\mathcal{M}_{\textsc{h}} with a dimension of 256 (used for the dictionary induction step). We also pre-train subword level embeddings on the segmented M\textscl\mathcal{M}_{\textsc{l}}, M^\textscl\hat{\mathcal{M}}_{\textsc{l}} and M\textsch\mathcal{M}_{\textsc{h}} with the same dimension.

3 Model Architecture

We use the self-attention Transformer model Vaswani et al. (2017). We adapt the implementation from the open-source translation toolkit OpenNMT Klein et al. (2017). Both encoder and decoder consist of 4 layers, with the word embedding and hidden unit dimensions set to 256. We tuned on multiple settings to find the optimal parameters for our datasets. We use a batch size of 8096 tokens.

Unsupervised NMT

We train unsupervised Transformer models with the UnsupervisedMT toolkit.https://github.com/facebookresearch/UnsupervisedMT Layer sizes and dimensions are the same as in the supervised NMT model. The parameters of the first three layers of the encoder and the decoder are shared. The embedding layers are initialized with the pre-trained subword embeddings from monolingual data. We set the weight parameters for autodenoising language modeling and iterative back translation as λ1=1\lambda_{1}=1 and λ2=1\lambda_{2}=1.

4 Training and Model Selection

After data augmentation, we follow the pre-train and fine tune paradigm for learning Zoph et al. (2016); Nguyen and Chiang (2017). We first train a base NMT model on the concatenation of {S\textscle,T\textscle}\{\mathcal{S}_{\textsc{le}},\mathcal{T}_{\textsc{le}}\} and {S\textsche,T\textsche}\{\mathcal{S}_{\textsc{he}},\mathcal{T}_{\textsc{he}}\}. Then we adopt the mixed fine-tuning strategy of Chu et al. (2017), fine-tuning the base model on the concatenation of the base and augmented datasets. For each setting, we perform a sufficient number of updates to reach convergence in terms of development perplexity.

We use the performance on the development sets (as provided by the TED corpus) as our criterion for selecting the best model, both for augmentation and final model training.

Results and Analysis

A collection of our results with the baseline and our proposed methods is shown in Table 2.

The performance of the base supervised model (row 1) varies from 11.8 to 29.5 BLEU points. Generally, the more distant the source language is from English, the worse the performance. A standard unsupervised MT model (row 2) achieves extremely low scores, confirming the results of Guzmán et al. (2019), indicating the difficulties of directly translating between lrland eng in an unsupervised fashion. Rows 3 and 4 show that standard supervised back-translation from English at best yields very modest improvements. Notable is the exception of slk-eng, which has more parallel data for training than other settings. In the case of bel and glg, it even leads to worse performance. Across all four languages, supervised back-translation into the hrl helps more than into the lrl; data is insufficient for training a good lrl-eng MT model.

2 Back-translation from hrl

Rows 5–9 show the results when we create data using the hrl side of an hrl-eng dataset. Both the low-resource supervised (row 5) and vanilla unsupervised (row 6) hrl ⁣)\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}eng translation do not lead to significant improvements. On the other hand, our simple word substitution approach (row 7) and the modified umt approach (row 8) lead to improvements across the board: +3.0 BLEU points in aze, +7.8 for bel, +2.3 for glg, +1.1 for slk. These results are significant, demonstrating that the quality of the back-translated data is indeed important.

In addition, we find that combining the datasets produced by our word substitution and umt models provide an additional improvement in all cases (row 9). Interestingly, this happens despite the fact that the eng data are the exact same between rows 5–9.

eng-hrl-lrl

We also show that even in the absence of parallel hrl-lrl data, our pivoting method is still valuable. Rows 10 and 11 in Table 2 show the translation accuracy when the augmented data are the result of our two-step pivot back-translation. In both cases, monolingual eng is first translated into hrl and then into lrl with either just word substitution (row 10) or modified umt (row 11). Although these results are slightly worse than our one-step augmentation of a parallel hrl-lrl dataset, they still outperform the baseline standard back-translation (rows 3 and 4). An interesting note is that in this setting, word substitution is clearly preferable to umt for the second translation pivoting step, which we explain in §5.3.

Combinations

We obtain our best results by combining the two sources of data augmentation. Row 12 shows the result of using our simple word substitution technique on the hrl side of both a parallel and an artificially created (back-translated) hrl-eng dataset. In this setting, we further improve not only the encoder side of our model, as before, but we also aid the decoder's language modeling capabilities by providing eng data from two distinct resources. This leads to improvements of 3.6 to 8.2 BLEU points over the base model and 0.3 to 2.1 over our best results from hrl-eng augmentation.

Finally, row 13 shows our attempt to obtain further gains by combining the datasets from both word substitution and umt, as we did in setting 7. This leads to a small improvement of 0.2 BLEU points in aze, but also to a slight degradation on the other three datasets.

We also compare the results of our augmentation methods with other state-of-the-art methods that either perform improvements to modeling to improve the ability to do parameter sharing Wang et al. (2019), or train on many different target languages simultaneously Aharoni et al. (2019). The results demonstrate that the simple data augmentation strategies presented here improve significantly over these previous methods.

3 Analysis

In this section we focus on the quality of hrl ⁣)\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}lrl translation, showing that our better m-umt initialization method leads to significant improvements compared to standard umt.

We use the dev sets of the hrl-lrl datasets to examine the performance of m-umt between related languages. We calculate the pivot BLEUWe will refer to pivot BLEU in order to avoid confusion with translation BLEU scores from the previous sections. score on the lrl side of each created dataset (S\textschl\mathcal{S}_{\textsc{hl}}, \hat{\mathcal{S}}^{w}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}}, \hat{\mathcal{S}}^{u}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}}, \hat{\mathcal{S}}^{m}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}}). In Figure 2 we plot pivot hrl-lrl BLEU scores against the translation lrl-eng BLEU ones. First, we observe that across all datasets, the pivot BLEU of our m-umt method is higher than standard umt (the squares are all further right than their corresponding stars). Vanilla umt's scores are 2 to 10 BLEU points worse than the m-umt ones. This means that umt across related languages significantly benefits from initializing with our simple word substitution method.

Second, as illustrated in Figure 2, the pivot BLEU score and the translation BLEU are imperfectly correlated; even though m-umt reaches the highest pivot BLEU, the resulting translation BLEU is comparable to using the simple word substitution method (rows 7 and 8 in Table 2). The reason is that the quality of \{\hat{\mathcal{S}}^{m}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}}\ ,\ \mathcal{T}_{\textsc{he}}\} is naturally restricted by the \{\hat{\mathcal{S}}^{w}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}}\ ,\ \mathcal{T}_{\textsc{he}}\}, whose quality is in turn restricted by the induced dictionary. However, by combining the augmented datasets from these two methods, we consistently improve the translation performance over using only word substitution augmentation (compare Table 2 rows 7 and 9). This suggests that the two augmented sets improve lrl-eng translation in an orthogonal way.

Additionally, we observe that augmentation from back-translated hrl data leads to generally worse results than augmentation from original hrl data (compare rows 7,8 with rows 10,11 in Table 2). We believe this to be the result of noise in the back-translated hrl, which is then compounded by further errors from the induced dictionary. Therefore, we suggest that the simple word substitution method should be preferred for the second pivoting step when augmenting back-translated hrl data.

Table 3 provides an example conversion of an hrl sentence to pseudo-lrl with the word substitution strategy, and its translation with m-umt. From S\textsche\mathcal{S}_{\textsc{he}} to \hat{\mathcal{S}}^{w}_{\textsc{h\mathrel{\hbox{\rule[-0.2pt]{3.0pt}{0.4pt}}\mkern-4.0mu\hbox{\char 41\relax}}l}}, the word substitution strategy achieves very high unigram scores (0.50 in this case), largely narrowing the gap between two languages. The m-umt model then edits the pseudo-lrl sentence to convert all its words to lrl.

Next, we quantitatively evaluate how our pivoting augmentation methods increase rare word coverage and the correlation with lrl-eng translation quality. For each word in the tested set, we define a word as ``rare'' if it is in the training set's lowest 10th frequency percentile. This is particularly true for lrl test set words when using concatenated hrl-lrl training data, as the lrl data will be smaller. We further define rare words to be ``addressed'' if after adding augmented data the rare word is not in the lowest 10th frequency percentile anymore. Then, we define the ``address rate'' of a test dataset as the ratio of the number of addressed words to the number of rare words. The address rate of each method, along with the corresponding translation BLEU score is shown in Figure 3. As indicated by the Pearson correlation coefficients, these two metrics are highly correlated, indicating that our augmentation methods significantly mitigate problems caused by rare words, improving MT quality as a result.

Dictionary Induction

We conduct experiments to compare two methods of dictionary induction from the mapped word embedding spaces: 1) Unidirectional: For each hrl word, we collect its closest lrl word to be added to the dictionary; 2) Bidirectional: We only add word pairs the two words of which are each other's closest neighbor to the dictionary.

In order to know how many lrl words are injected into the hrl corpus, we show the number of injected unique word types, number of injected words, and the corresponding BLEU score of models trained with bidirectional and unidirectional word induction in Table 4. It can be seen that the ratio of word numbers is higher than that of word types between bidirectional and unidirectional word induction, indicating that the injected words using the bidirectional method are of relatively high frequency. The BLEU scores show that bidirectional word induction performs better than unidirectional induction in most cases (except bel). One explanation could be that adding each word's closest neighbor as a pair into the dictionary introduces additional noise that might harm the low-resource translation to some extent.

Related Work

Our work is related to multilingual and unsupervised translation, bilingual dictionary induction, as well as approaches for triangulation (pivoting).

In a low-resource MT scenario, multilingual training that aims at sharing parameters by leveraging parallel datasets of multiple languages is a common practice. Some works target learning a universal representation for all languages either by leveraging semantic sharing between mapped word embeddings Gu et al. (2018) or by using character n-gram embeddings Wang et al. (2019) optimizing subword sharing. More related with data augmentation, Nishimura et al. (2018) fill in missing data with a multi-source setting to boost multilingual translation.

Unsupervised machine translation enables training NMT models without parallel data Artetxe et al. (2018); Lample et al. (2018a, c). Recently, multiple methods have been proposed to further improve the framework. By incorporating a statistical MT system as posterior regularization, Ren et al. (2019) achieved state-of-the-art for en-fr and en-de MT. Besides MT, the framework has also been applied to other unsupervised tasks like non-parallel style transfer Subramanian et al. (2019); Zhang et al. (2018).

Bilingual dictionaries learned in both supervised and unsupervised ways have been used in low-resource settings for tasks such as named entity recognition Xie et al. (2018) or information retrieval Litschko et al. (2018). Hassan et al. (2017) synthesized data with word embeddings for spoken dialect translation, with a process that requires a lrl-eng as well as a hrl-lrl dictionary, while our work only uses a hrl-lrl dictionary.

Bridging source and target languages through a pivot language was originally proposed for phrase-based MT De Gispert and Marino (2006); Cohn and Lapata (2007). It was later adapted for Neural MT Levinboim and Chiang (2015), and Cheng et al. (2017) proposed joint training for pivot-based NMT. Chen et al. (2017) proposed to use an existing pivot-target NMT model to guide the training of source-target model. Lakew et al. (2018) proposed an iterative procedure to realize zero-shot translation by pivoting on a third language.

Conclusion

We propose a generalized data augmentation framework for low-resource translation, making best use of all available resources. We propose an effective two-step pivoting augmentation method to convert hrl parallel data to lrl. In future work, we will explore methods for controlling the induced dictionary quality to improve word substitution as well as m-umt. We will also attempt to create an end-to-end framework by jointly training m-umt pivoting system and low-resource translation system in an iterative fashion in order to leverage more versions of augmented data.

Acknowledgements

The authors thank Junjie Hu and Xinyi Wang for discussions on the paper. This material is based upon work supported in part by the Defense Advanced Research Projects Agency Information Innovation Office (I2O) Low Resource Languages for Emergent Incidents (LORELEI) program under Contract No. HR0011-15-C0114 and the National Science Foundation under grant 1761548. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation here on.

References