Extreme Adaptation for Personalized Neural Machine Translation

Paul Michel, Graham Neubig

Introduction

The production of language varies depending on the speaker or author, be it to reflect personal traits (e.g. job, gender, role, dialect) or the topics that tend to be discussed (e.g. technology, law, religion). Current Neural Machine Translation (NMT) systems do not incorporate any explicit information about the speaker, and this forces the model to learn these traits implicitly. This is a difficult and indirect way to capture inter-personal variations, and in some cases it is impossible without external context (Table 1, Mirkin et al. (2015)).

Recent work has incorporated side information about the author such as personality Mirkin et al. (2015), gender Rabinovich et al. (2017) or politeness Sennrich et al. (2016a), but these methods can only handle phenomena where there are explicit labels for the traits. Our work investigates how we can efficiently model speaker-related variations to improve NMT models.

In particular, we are interested in improving our NMT system given few training examples for any particular speaker. We propose to approach this task as a domain adaptation problem with an extremely large number of domains and little data for each domain, a setting where we may expect traditional approaches to domain adaptation that adjust all model parameters to be sub-optimal (§2). Our proposed solution involves modeling the speaker-specific variations as an additional bias vector in the softmax layer, where we either learn this bias directly, or through a factored model that treats each user as a mixture of a few prototypical bias vectors (§3).

We construct a new dataset of Speaker Annotated TED talks (SATED, §4) to validate our approach. Adaptation experiments (§5) show that explicitly incorporating speaker information into the model improves translation quality and accuracy with respect to speaker traits.Data/code publicly available at http://www.cs.cmu.edu/~pmichel1/sated/ and https://github.com/neulab/extreme-adaptation-for-personalized-translation respectively.

Problem Formulation and Baselines

In the rest of this paper, we refer to the person producing the source sentence (speaker, author, etc…) generically as the speaker. We denote as S\mathcal{S} the set of all speakers.

The usual objective of NMT is to find parameters θ\theta of the conditional distribution p(y∣x;θ)p(y\mid x;\theta) to maximize the empirical likelihood. We argue that personal variations in language warrant decomposing the empirical distribution into ∣S∣|\mathcal{S}| speaker specific domains Ds\mathcal{D}_{s} and learning a different set of parameters θs\theta_{s} for each. This setting exhibits specific traits that set it apart from common domain adaptation settings:

The number of speakers is very large. Our particular setting deals with ∣S∣≈1800|\mathcal{S}|\approx 1800 but our approaches should be able to accommodate orders of magnitude more speakers.

There is very little data (even monolingual, let alone bilingual or parallel) for each speaker, compared to millions of sentences usually used in NMT.

As a consequence of 1, we can assume that many speakers share similar characteristics such as gender, social status, and as such may have similar associated domains.Note that the speakers are still unique, and many might use very specific words (e.g. the name of their company or of a specific medical procedure that they are an expert on).

All of our experiments are based on a standard neural sequence to sequence model. We use one layer LSTMs as the encoder and decoder and the concat attention mechanism described in Luong and Manning (2015). We share the parameters in the embedding and softmax matrix of the decoder as proposed in Press and Wolf (2017). All the layers have dimension 512 except for the attention layer (dimension 256). To make our baseline competitive, we apply several regularization techniques such as dropout Srivastava et al. (2014) in the output layer and within the LSTM (using the variant presented in Gal and Ghahramani, 2016). We also drop words in the target sentence with probability 0.1 according to Iyyer et al. (2015) and implement label smoothing as proposed in Szegedy et al. (2016) with coefficient 0.10.1. Appendix A provides a more thorough description of the baseline model.

2 Baseline adaptation strategy

As mentioned in §2, our goal is to learn a separate conditional distribution p(y∣x,s)p(y\mid x,s) and parametrization θs\theta_{s} to improve translation for speaker ss. The usual way of adapting from general domain parameters θ\theta to θs\theta_{s} is to retrain the full model on the domain specific data Luong and Manning (2015). Naively applying this approach in the context of personalizing a model for each speaker however has two main drawbacks:

Maintaining a set of model parameters for each speaker is expensive. For example, the model in §2.1 has ≈\approx47M parameters when the vocabulary size is 40k, as is the case in our experiments in §5. Assuming each parameter is stored as a 32bit float, every speaker-specific model costs ≈\approx188MB. In a production environment with thousands to billions of speakers, this is impractical.

Overfitting

Training each speaker model with very little data is a challenge, necessitating careful and heavy regularization Miceli Barone et al. (2017) and an early stopping procedure.

3 Domain Token

A more efficient domain adaptation technique is the domain token idea used in Sennrich et al. (2016a); Chu et al. (2017): introduce an additional token marking the domain in the source and/or the target sentence. In experiments, we add a token indicating the speaker at the start of the target sentence for each speaker. We refer to this method as the spk_token method in the following.

Note that in this case there is now only an embedding vector (of dimension 512 in our experiments) for each speaker. However, the resulting domain embedding are non-trivial to interpret (i.e. it is not clear what they tell us about the domain or speaker itself).

Speaker-specific Vocabulary Bias

In NMT models, the final choice of which word to use in the next step tt of translation is generally performed by the following softmax equation

where oto_{t} is predicted in a context-sensitive manner by the NMT system and ETE_{T} and bTb_{T} are the weight matrix and bias vector parameters respectively. Importantly, bTb_{T} governs the overall likelihood that the NMT model will choose particular vocabulary. In this section, we describe our proposed methods for making this bias term speaker-specific, which provides an efficient way to allow for speaker-specific vocabulary choice.Notably, while this limits the model to only handling word choice and does not explicitly allow it to model syntactic variations, favoring certain words over others can indirectly favor certain phenomena (e.g. favoring passive speech by increasing the probability of auxiliaries).

We first propose to learn speaker-specific parameters for the bias term in the output softmax only. This means changing Eq. 1 to

for speaker ss. This only requires learning and storing a vector equal to the size of the vocabulary, which is a mere 0.09% of the parameters in the full model in our experiments. In effect, this greatly reducing the parameter cost and concerns of overfitting cited in §2.2. This model is also easy to interpret as each coordinate of the bias vector corresponds to a log-probability on the target vocabulary. We refer to this variant as full_bias.

2 Factored speaker bias

The biases for a set of speakers S\mathcal{S} on a vocabulary V\mathcal{V} can be represented as a matrix:

where each row of BB is one speaker bias bsb_{s}. In this formulation, the ∣S∣|\mathcal{S}| rows are still linearly independent, meaning that BB is high rank. In practical terms, this means that we cannot share information among users about how their vocabulary selection co-varies, which is likely sub-ideal given that speakers share common characteristics.

Thus, we propose another parametrization of the speaker bias, fact_bias, where the BB matrix is factored according to:

We provide a graphical summary of our proposed approaches in figure 1.

Speaker Annotated TED Talks Dataset

In order to evaluate the effectiveness of our proposed methods, we construct a new dataset, Speaker Annotated TED (SATED) based on TED talks,https://www.ted.com with three language pairs, English-French (en-fr), English-German (en-de) and English-Spanish (en-es) and speaker annotation.

The dataset consists of transcripts directly collected from https://www.ted.com/talks, and contains roughly 271K sentences in each language distributed among 2324 talks. We pre-process the data by removing sentences that don’t have any translation or are longer than 60 words, lowercasing, and tokenizing (using the Moses tokenizer Koehn et al. (2007)).

Some talks are partially or not translated in some of the languages (in particular there are fewer translations in German than in French or Spanish), we therefore remove any talk with less than 10 translated sentences in each language pair.

The data is then partitioned into training, validation and test sets. We split the corpus such that the test and validation split each contain 2 sentence pairs from each talk, thus ensuring that all talks are present in every split. Each sentence pair is annotated with the name of the talk and the speaker. Table 2 lists statistics on the three language pairs.

This data is made available under the Creative Commons license, Attribution-Non Commercial-No Derivatives (or the CC BY-NC-ND 4.0 International, https://creativecommons.org/licenses/by-nc-nd/4.0/legalcode), all credit for the content goes to the TED organization and the respective authors of the talks. The data itself can be found at http://www.cs.cmu.edu/~pmichel1/sated/.

Experiments

We run a set of experiments to validate the ability of our proposed approach to model speaker-induced variations in translation.

We test three models base (a baseline ignoring speaker labels), full_bias and fact_bias. During training, we limit our vocabulary to the 40,000 most frequent words. Additionally, we discard any word appearing less than 2 times. Any word that doesn’t satisfy those conditions is replaced with an UNK token.Recent NMT systems also commonly use sub-word units Sennrich et al. (2016b). This may influence on the result, either negatively (less direct control over high-frequency words) or positively (more capacity to adapt to high-frequency words). We leave a careful examination of these effects for future work.

All our models are implemented with the DyNet Neubig et al. (2017) framework, and unless specified we use the default settings therein. We refer to appendix B for a detailed explanation of the training process. We translate the test set using beam search with beam size 5.

2 Does explicitly modeling speaker-related variation improve translation quality?

Table 3 shows final test scores for each model with statistical significance measured with paired bootstrap resampling Koehn (2004). As shown in the table, both proposed methods give significant improvements in BLEU score, with the biggest gains in English to French (+0.99+0.99) and smaller gains in German and Spanish (+0.74+0.74 and +0.40+0.40 respectively). Reducing the number of parameters with fact_bias gives slightly better (en-fr) or worse (en-de) BLEU score, but in those cases the results are still significantly better than the baseline.

However, BLEU is not a perfect evaluation metric. In particular, we are interested in evaluating how much of the personal traits of each speaker our models capture. To gain more insight into this aspect of the MT results, we devise a simple experiment. For every language pair, we train a classifier (continuous bag-of-n-grams; details in Appendix C) to predict the author of each sentence on the target language part of the training set. We then evaluate the classifier on the ground truth and the outputs from our 3 models (base, full_bias and fact_bias).

The results are reported in Figure 2. As can be seen from the figure, it is easier to predict the author of a sentence from the output of speaker-specific models than from the baseline. This demonstrates that explicitly incorporating information about the author of a sentence allows for better transfer of personal traits during translations, although the difference from the ground truth demonstrates that this problem is still far from solved. Appendix D shows qualitative examples of our model improving over the baseline.

3 Further experiments on the Europarl corpus

One of the quirks of the TED talks is that the speaker annotation correlates with the topic of their talk to a high degree. Although the topics that a speaker talks about can be considered as a manifestation of speaker traits, we also perform a control experiment on a different dataset to verify that our model is indeed learning more than just topical information. Specifically, we train our models on a speaker annotated version of the Europarl corpus Rabinovich et al. (2017), on the en-de language pairavailable here: https://www.kaggle.com/ellarabi/europarl-annotated-for-speaker-gender-and-age/version/1.

We use roughly the same training procedure as the one described in §5.1, with a random train/dev/test split since none is provided in the original dataset. Note that in this case, the number of speakers is much lower (747) whereas the total size of the dataset is bigger (≈\approx300k).

We report the results in table 4. Although the difference is less salient than in the case of SATED, our factored bias model still performs significantly better than the baseline (+0.83+0.83 BLEU). This suggests that even outside the context of TED talks, our proposed method is capable of improvements over a speaker-agnostic model.

Related work

Domain adaptation techniques for MT often rely on data selection Moore and Lewis (2010); Li et al. (2010); Chen et al. (2017); Wang et al. (2017), tuning Luong and Manning (2015); Miceli Barone et al. (2017), or adding domain tags to NMT input Chu et al. (2017). There are also methods that fine-tune parameters of the model on each sentence in the test set Li et al. (2016), and methods that adapt based on human post-edits Turchi et al. (2017), although these follow our baseline adaptation strategy of tuning all parameters. There are also partial update methods for transfer learning, albeit for the very different task of transfer between language pairs Zoph et al. (2016).

Pioneering work by Mima et al. (1997) introduced ways to incorporate information about speaker role, rank, gender, and dialog domain for rule based MT systems. In the context of data-driven systems, previous work has treated specific traits such as politeness or gender as a “domain” in domain adaptation models and applied adaptation techniques such as adding a “politeness tag” to moderate politeness Sennrich et al. (2016a), or doing data selection to create gender-specific corpora for training Rabinovich et al. (2017). The aforementioned methods differ from ours in that they require explicit signal (gender, politeness…) for which labeling (manual or automatic) is needed, and also handle a limited number of “domains” (≈2\approx 2), where our method only requires annotation of the speaker, and must scale to a much larger number of “domains” (≈1,800\approx 1,800).

Conclusion

In this paper, we have explained and motivated the challenge of modeling the speaker explicitly in NMT systems, then proposed two models to do so in a parameter-efficient way. We cast this problem as an extreme form of domain adaptation and showed that, even when adapting a small proportion of parameters (the softmax bias, <0.1%<0.1\% of all parameters), allowed the model to better reflect personal linguistic variations through translation.

We further showed that the number of parameters specific to any person could be reduced to as low as 10 while still retaining better scores than a baseline for some language pairs, making it viable in a real world application with potentially millions of different users.

Acknowledgements

The authors give their thanks the anonymous reviewers for their useful feedback which helped make this paper what it is, as well as the members of Neulab who helped proof read this paper and provided constructive criticism. This work was supported by a Google Faculty Research Award 2016 on Machine Translation.

References

Appendix A Detailed model description

Encoder

Our encoder is a one layer bidirectional LSTM with dimension dh=512d_{h}=512. For a source sentence e=e1,…,e∣e∣e=e_{1},\dots,e_{|e|} the concatenated output of the encoder is thus of shape ∣e∣×2dh|e|\times 2d_{h}.

Attention

We use a multilayer perceptron attention mechanism: given a query hth_{t} at step tt of decoding and encodings x1,…,x∣e∣x_{1},\ldots,x_{|e|}, the context vector ctc_{t} is computed according to:

where Va,Wa,Wah,baV_{a},W_{a},W_{ah},b_{a} are learned parameters. We choose da=256d_{a}=256 as the dimension of the intermediate layer.

Decoder

The decoder is a single layer LSTM of dimension dh=512d_{h}=512. At each timestep tt, it takes as input the previous word embedding wt−1w_{t-1} and the previous context ct−1c_{t-1}. Its output hth_{t} is used to compute the next context vector ctc_{t} and the distribution over the next possible target words wtw_{t}:

where Wo∗,bo,bTW_{o*},b_{o},b_{T} are learned parameters, ETE_{T} is the target word embedding matrix and ETwt−1E_{Tw_{t-1}} is the embedding of the previous target word.

Learning paradigm

We employ several techniques to improve training. First, we are using the same parameters for the target word embeddings and the weights of the softmax matrix Press and Wolf (2017). This reduces the number of total parameters and in practice this gave slightly better BLEU scores.

We apply dropout Srivastava et al. (2014) between the output layer and the softmax layer, as well as within the LSTM (using the variant presented in Gal and Ghahramani (2016)). We also drop words in the target sentence with probability 0.1 according to Iyyer et al. (2015). Intuitively, this forces the decoder to use the conditional information.

In addition to this, we implement label smoothing as proposed in Szegedy et al. (2016) with a smoothing coefficient 0.1. We noticed improvements of up to 1 BLEU point with this additional regularization term.

Appendix B Training process

We first train each model using the Adam optimizer Kingma and Ba (2014) with learning rate 0.001 (we clip the gradient norm to 1). The data is split into batches of size 32 where every source sentence has the same length. We evaluate the validation perplexity after each epoch. Whenever the perplexity doesn’t improve, we restart the optimizer with a smaller learning rate from the previous best model Denkowski and Neubig (2017). Training is stopped when the perplexity doesn’t go down for 3 epochs. We then perform a tuning step: we restart training with the same hyper-parameters except for using simple stochastic gradient descent and gradient clipping at a norm of 0.1, which improved the validation BLEU by 0.3-0.9 points.

Appendix C User classifier

In our analysis, we use a classifier to estimate which user wrote each output, which we describe more in this section.

The model uses a continuous bag of nn-grams where vn-gramv_{\text{n-gram}} is a parameter vector for a paticular n-gram and the probability of speaker ss for sentence ff is given by:

The size of hidden vectors is 128. We limit n-grams to unigrams and bigrams. We estimate the parameters with Adam and a batch size of 32 for 50 epochs.

Appendix D Qualitative examples

Table 5 shows examples where our full_bias/fact_bias model helped translation by favoring certain words as opposed to the baseline in en-fr.