Dictionary-based Phrase-level Prompting of Large Language Models for Machine Translation

Marjan Ghazvininejad, Hila Gonen, Luke Zettlemoyer

Introduction

Large language models (LLMs) can be prompted to perform very high quality machine translation (MT), even though they were not explicitly trained for this task Brown et al. (2020); Lin et al. (2021); Zhang et al. (2022); Scao et al. (2022). To do this, we simply ask them to complete a string Translate the following sentence to English or French etc. following by an input sentence. However, despite the fact that they are trained on massive corpora, these models can struggle to correctly translate rare words, which are common in low resource or domain transfer scenarios. In this paper, we show that the controlability that comes with prompting LLMs extends to the word level, allowing us to not only specify the overall task but also provide hints or options about individual choices the model should consider when performing the task.

More specifically, we assume access to prior knowledge from bilingual dictionaries for machine translation, and show how to prompt LLMs with hints that specify a set of possible translation options for specific input words. As seen in Figure 1, this involves simply appending a string such as in this context, the word ‘‘membatasi’’ means ‘‘limiting,’’ ‘‘restrict,’’ ‘‘limit.’’ to the end of the typical MT prompt, for every word that we have hints for in the lexicon. This approach is inspired by supervised machine translation models that have successfully used dictionaries to improve translation Zhang and Zong (2016); Arthur et al. (2016); Zhong and Chiang (2020), but crucially we show how to incorporate the knowledge in a zero-shot fashion with no model training. It provides a simple method to not only specify the task in the prompt, but also provide background knowledge that might be useful for helping the model complete the task, without enforcing any hard constraints on how the model uses this knowledge.

Extensive experiments show that this approach, which we call DiPMT (Dictionary-based Prompting for Machine Translation), we are able to achieve significant gains for low-resource translation from and into English across multiple languages and language models. Additionally, we experiment also with out-of-domain translation, where we automatically extract dictionaries from the training data Shi et al. (2021), and show improvements of up to 13 Bleu points in this setting. As a next step into understanding the abilities of DiPMT better, we analyze and explore the behaviour and the limits of the model, and measure the extent to which DiPMT is controlling the final translation in practice.

Our contributions in this work are the following: (a) we propose DiPMT, a novel method for incorporating bilingual dictionary information in prompting-based machine translation (Section 2); (b) we show empirically that DiPMT improves translation significantly both for low-resource and out-of-domain translation (Section 4); (c) we analyze DiPMT and provide insights about the benefits and limitations of the method (Section 5); (d) we explore the extent to which DiPMT controls the final translation in practice (Section 6).

Machine Translation via Prompting with Dictionary

Prompting language models for translation assumes that the model is pretrained on enough training data in both the source and the target languages. However, that is often not the case, especially for low-resource languages and more so when the model is trained primarily on English data. Another challenge that affects translation quality is that of out-of-domain data. When asked to translate data that is out-of-domain with respect to the pretraining data, the model struggles to output high quality translations Koehn and Knowles (2017).

To alleviate these challenges, we present a method for incorporating dictionary knowledge into prompting-based MT. Dictionaries are relatively easy to obtain, even for low resource languages, making them appealing candidates for external source of translation Zhang and Zong (2016); Arthur et al. (2016); Hämäläinen and Alnajjar (2019); Zhong and Chiang (2020). Our method is simple, easy to use, and applies to any language pair where a dictionary exists.

Our method, DiPMT uses prompting-based translation and incorporates dictionary information into the prompt directly. Given a source sentence, we look for dictionary entries for the words included in it, and add them to the prompt directly: after asking for the translation of the full sentence, we add an additional part to the prompt that lists the possible translations for specific words. The prompt for each sentence has three parts as demonstrated in Figure 1:

(1) the source sentence: ‘‘Translate the following sentence to English: ’’; (2) the dictionary translations: ‘‘In this context, the word X means A; the word Y means B,C,D.’’; (3) Asking for the translation to the target language: ‘‘The full translation to English is:’’

To make the model familiar with the specific framework we choose, we add kk demonstrations before prompting for the current instance. These demonstrations include the translation of the source sentence in their third part.

The dictionaries we use are word-level dictionaries. Extending DiPMT to incorporate phrase-level translations should be trivial, and we leave this exploration to future work.

DiPMT is simple and straight-forward to use, and also effective in improving translation quality, as we show in Section 4.

Experimental Setup

We experiment with two large scale language models: OPT Zhang et al. (2022) which is mostly trained on English data, and Bloom Scao et al. (2022), a multilingual language model.

We use the publicly available checkpoint of the OPT-175B, a decoder-only causal language model, with 96 layers and a hidden dimension of 12,288. The training corpora contains predominantly English text with a small amount of non-English data that has skipped the filtering process. Yet, the model exhibits good performance on few-shot translation, especially when translating into English (see Table 3, second column, for reference).

{‌

Bloom To assess the effectiveness of our proposed method on multilingual language models, we also use the publicly available Bloom-176B checkpoint. Bloom is trained on the ROOTS corpus Laurençon et al. (2022) consisting of 498 Hugging Face datasets Lhoest et al. (2021) involving 46 natural languages and 13 programming languages. Bloom is also a decoder-only transformer language model, with 70 layers and a hidden dimension of 14,336.

2 Datasets and Evaluation Metrics

For in-domain evaluation, we use Flores-101 Goyal et al. (2022), which contains 3,001 sentences taken from English Wikipedia on various topics and domains. These sentences have been professionally translated into 101 different languages. As we aim to focus on low-resource languages (with respect to the model), we select 10 languages on which OPT performs moderately wellBased on a random selection of 60 examples from the development set (10−3010-30 BLEU points) in a four-shot machine translation setting, from and into English. These languages are listed in Table 1. We refrain from using languages on which OPT performs poorly (<10<10 BLEU points) since we assume that the performance for those ones is too low to expect reasonable translations even when incorporating external information.

For out-of-domain evaluation we use De−EnDe-En data from Aharoni and Goldberg (2020), covering the following domains: Medical, Law, IT, and Koran. The dataset statistics are presented in table 2. All sentences longer than 250 tokens and sentence pairs with a source/target length ratio of more than 1.5 are removed from the training sets. We evaluate the detokenized length generated by the model using sacreBLEU Post (2018).https://github.com/mjpost/sacrebleu

3 Dictionaries

For in-domain translation, we use the ground-truth bilingual dictionaries provided in Conneau et al. (2017).https://github.com/facebookresearch/MUSE##ground-truth-bilingual-dictionaries These dictionaries were built using Meta’s internal translation tool and were designed to handle polysemy of words. For out-of-domain translation we build dictionaries based on the training data, using the method suggested by Shi et al. (2021). The process is explained in detail in Section 4.2.

4 Prompting Formulation

Given an instance to translate, we prepend 4 demonstrations to it. Each demonstration consists of 3 parts: (a) the source sentence (along with the translation instruction and the target language): ‘‘Translate the following sentence to English: ’’; (b) the dictionary-based word-level translations: ‘‘In this context, the word X means A; the word Y means B,C,D.’’; (c) the translation to the target language: ‘‘The full translation to English is: .’’ For the current instance to be translated, part (c) does not include the translation itself, and the model is expected to generate the translation into the target language. See a full example in Figure 1. To extract the dictionary hints, we look up each of the source words in the dictionary and if an exact match is found,We also consider matching after lemmatization (using Stanza Qi et al. (2020)), but since our preliminary experiments find no meaningful difference between these methods, we choose exact match. we provide the respective dictionary translation(s) as the hint(s). We do not provide any hints for the 500 most frequent source words (based on the development set) as we assume that those are easier for the model to learn and might incorporate noise into the model. In some rare cases where we have more than 3 possible translations for a source word, we choose 3 of them randomly.

For baselines, we use the same prompt format but without providing the dictionary-based word-level translations. Here, each demonstration consists of only two parts: (a) the source sentence (along with the translation instruction and the target language); and (b) the translation to the target language. See Figure 1 for an example – for the baseline we omit the blue parts (dictionary-based translations). To ensure a fair comparison, we select the demonstrations randomly from the development set, and consistently use the same demonstrations for all test sentences and across all models.

Improving Translation Performance

In this section, we show the effect of DiPMT in the two challenging translation settings: low-resource MT and out-of-domain MT. In both of them we expect to see significant improvements as external information has the potential to fill the gap of missing relevant pretraining data.

In Table 3, we report the results for low-resource languages of DiPMT vs. the baseline. As described in Section 3.2, we select 10 languages on which OPT performs moderately well as a proxy for low-resource languages. We experiment with OPT and BLOOM for both translation directions (from and into English) across the 10 languages. On average, we gain an improvement of 0.9 BLEU points with OPT and 1.1 BLEU points with BLOOM. Interestingly, OPT performance improves more when translating from English, while Bloom performance improves more when translating into English.

2 Out-of-domain MT

We also study how DiPMT performs in out-of-domain prompting-based translation. This setting is especially interesting since incorporating external information has the potential to result in significant gains when dealing with out-of-domain data. In these experiments, we translate medical, law, Koran, and IT texts. Despite the possibility that LLMs are trained also on similar domains, we get that the translation quality in these domains is still lacking, as can be seen in the baseline results in Table 4 (second raw). This is probably the case since LLMs are less likely to have observed sufficient monolingual data for these specialized domains, in contrast to Wikipedia-style data, for example. Sentences from these domains might require translating rare technical terms and idiosyncrasies which present unique challenges even for supervised neural MT models that are well-trained and suited explicitly for translation Koehn and Knowles (2017). In this setting, we do not assume that a comprehensive and accurate dictionary specific to a particular domain is available. Such dictionaries can be difficult to obtain, and in some cases may not even exist. Instead, we assume that there is some parallel data available for each domain, and we use this data to create a domain-specific dictionary. To extract the domain-specific dictionary, we use a combination of word alignment and a fully unsupervised bilingual lexicon induction method. Specifically, we first run the SimAlign algorithm Sabet et al. (2020) on the parallel data, and then apply the method proposed by Shi et al. (2021). Shi et al. (2021) propose a to estimate p(s,t)p(s,t) via a smoothed matched ratio between source word ss and target word tt in their parallel data and align a source word ss to the target word tt with the highest p(s,t)p(s,t). For more information, please refer to the paper. Since this method does not take into account the fact that words can have multiple meanings and translates each source word to only one target word, we modify the algorithm to consider all target words tt that have a probability p(s,t)≥λp(s,t)\geq\lambda for each source word ss.Based on our result on development set, we choose λ=0.1\lambda=0.1. We compare DiPMT with the baseline model (based on BloomWe choose Bloom for this experiment as our initial experiments indicate that it is more effective for De-En compared to OPT.), and with two other models: RePP Sun et al. (2022) and kNN-MT Khandelwal et al. (2020). RePP is a supervised machine translation system that utilizes bilingual phrase-level translation at test time, improving translation quality during inference through the use of retrieved phrase-level cues. kNN-MT is also based on a supervised MT model and uses the in-domain data for retrieval during inference. Then, the MT output is interpolated with the distributions of the retrieved tokens. The results are listed in Table 4. The improvement of DiPMT over the baseline is striking – we get that DiPMT outperforms the baseline by 9.49.4 BLEU points on average. Additionally, DiPMT also outperforms RePP by a large margin – 3.33.3 BLEU points on average. As for kNN-MT – DiPMT outperforms it for the KORAN domain. The superiority of kNN-MT over DiPMT for the other domains is likely due to a combination of two key-properties: (a) kNN-MT is based on retrieval from a corpus. This allows for full style change when needed, which is useful for these domains, as some of them are very patten based (e.g. medical). However, our model is focused on more limited word-level alternations; (b) kNN-MT is much more expensive to run than DiPMT.

Analysis

In this section we present a detailed analysis of DiPMT. We provide different hints in several different settings, as explained below, and analyze the resulting outputs. Naturally, not all source word types have a match in the dictionary. We study the effect of word type coverage in section 5.1. Additionally, not all respective translations from the dictionary appear in the target reference. In Section 5.2 we present an oracle experiment where we only provide gold hints, i.e., hints that are included in the reference translation, to get a sense of the upper bound of DiPMT. Finally, to get a better understanding of the generated output and model behavior, we look at a selection of input/output examples in Section 5.3.

Not all source words have a corresponding entry in the dictionary. The statistics for word token and word type coverage for each language pair in our dictionary are presented in Table 5. We report both tokens and types in order to get the full picture of word coverage – word type coverage is the most representative, since once we have a specific word in the dictionary, it will be covered for all instances. Word token coverage is relevant in order to estimate the percent of word instances that are covered in practice. To examine the relationship between BLEU improvement and source word type coverage, we start with the original dictionary and gradually remove random entries to obtain a dictionary version with lower coverage rates. We use a range of coverage rates, from 0% (no dictionary used) to the full coverage rate (typically around 35%), with steps of 5%. We then run our method with these modified versions of the dictionary and report the results. We perform this experiment with two language pairs (Eng↔\leftrightarrowInd and Eng↔\leftrightarrowMsa) on OPT and present the results in Figure 2. The model performs better than the baseline when the word type coverage is above a certain threshold (20%20\% for Ind→\rightarrowEng and 5−10%5-10\% for other pairs). Additionally, performance consistently improves as the coverage rate increases. This shows that higher dictionary coverage leads to better results, supporting the utility of DiPMT and suggesting that improved dictionary learning is an interesting direction for future work.

2 Gold Hints

To get a better understanding of the limits of DiPMT, we experiment with using only gold hints from the dictionary, i.e., hints that are included in the reference translation. This experiment provides us with the expected upper bound performance of incorporating dictionary entries into the prompts. For each source word with a dictionary match, we provide the model only with a single translation – the one that appears in the reference target sentence, if at all.In rare cases where several dictionary translations to the same token are included in the reference, we do not provide a hint for that word. This is in contrast to the usual use of DiPMT where we provide all the available dictionary translations.Table 6 presents the results. It shows that if DiPMT has access to the oracle translations, it can improve the BLEU score by another a 1.1 points on average for OPT.

3 Output Analysis

In this section we select a few representative input/output examples to demonstrate the strengths and weaknesses of DiPMT. Figure 3 exemplifies the way DiPMT helps the model generate the correct translation. Such examples are common and are the source of the BLEU improvement we see in Section 4. In this example the hints in boldface for the three words "serangga" (insects), "melipat" (fold), "capung" (dragonflies) have helped the system generate a better translation.

The next examples show failure cases of DiPMT: the model’s output initially matches the reference, but when provided with the hints, some tokens are mistranslated either because the desired translation is not included in the hints (Figure 4) or it is included but the model picks some other provided hint (Figure 5). In both of these examples the generated output is still a correct translation but there is a style mismatch between the new output and the reference.

The model often makes use of the hints to generate the output. The example in Figure 6 shows a relatively rare case where the baseline output is incorrect and the hints are not used in the new output even though they match the translation in the reference. In this example the model chooses to ignore the hints even though they could help in improving the translation.

Controllability

We now turn to study the extent to which DiPMT controls the final translation. In the previous section we observe that incorporating hints indeed helps the model generate better translations, but now we are interested to explore the actual direct effect the hints have on the translation. To evaluate the level of controllability of DiPMT we look at the percentage of times the system uses a suggested hint in the generated output. This should, of course, be compared with the respective percentage for the baseline, since the model might use a suggested hint even in baseline setting, when the hint is not provided. In case multiple translation hints are available for a single source token – if any of the hints appears in the target sentence, we consider it as hit. Multiple hits for a single source token are counted as one. We compare these results with two control experiments: in random hint, we choose a single dictionary translations (if available) for each word at random and provide it as the only hint for that word; in false hint, we shuffle the target side of the dictionary to obtain a false dictionary, and proceed similarly. The results are presented in Table 7. As expected, we get that DiPMT is able to control the final input, with significant gap over the percentage of the baseline (dictionary hint). The highest controllability level is achieved when providing only the gold hint. Surprisingly, for the false hint setting, we get almost no change with respect to the baseline, suggesting that our method is resilient to false word-level hints.

Related Work

There have been relatively few studies on prompting language models for machine translation. Most research in this area has focused on testing the machine translation ability of large language models using simple prompts like {\{Source text}={\}=\{Target text}\} or Translate to {\{language_name}:{\}:\{text}\} Brown et al. (2020); Lin et al. (2021); Zhang et al. (2022); Scao et al. (2022); Garcia and Firat (2022). Reynolds and McDonell (2021) experiment with different prompt templates, and Garcia and Firat (2022) explore the use of prompts for controlling various aspects of the formality or specific dialect of the output. There is also a line of work (Agrawal et al. 2022; Vilar et al. 2022) that concentrates on choosing good in-context examples for machine translation. This is in parallel with our work, which utilizes dictionary for better translation.

Using Dictionary in MT

Several researchers have investigated the use of dictionaries in supervised machine translation. Zhang and Zong (2016) propose a method that combines NMT with a bilingual dictionary containing rare or unseen words in the bilingual training data. Arthur et al. (2016) present a method for improving the translation of low-frequency words in NMT by augmenting the system with discrete translation lexicons and using the attention vector to select relevant lexical probabilities. Zhong and Chiang (2020) propose a similar method for using bilingual and monolingual dictionaries to improve NMT. In related research, Hämäläinen and Alnajjar (2019) use a dictionary to build synthetic parallel data for training a stronger NMT system. Our work is also related to lexical constraints in MT, which can be classified as hard constraints Hokamp and Liu (2017); Post and Vilar (2018) or soft constraints Song et al. (2019); Dinu et al. (2019); Chen et al. (2021).

Domain Adaptation for MT

Previous efforts have been made to enhance the performance of pre-trained NMT models using out-of-domain bilingual or monolingual datasets. Our method, similar to prevoius works Khandelwal et al. (2020); Zheng et al. (2021); Agrawal et al. (2022), adapts to the new domain during inference time and does not require additional training. The work most similar to ours is Sun et al. (2022), which trains a system that can use bilingual phrase-level translation at test time and improves translation quality during inference by constructing a bilingual phrase-level database and using retrieved phrase-level cues.

Conclusion

We propose a method for incorporating bilingual dictionary information into prompting-based MT by explicitly adding possible word-level translations into the prompt. We show that our method, DiPMT, improves translation quality for low-resource languages. Additionally, we tackle out-of-domain translation by creating dictionaries in an unsupervised manner and using them with our method to get impressive translation improvements. Our analyses of DiPMT demonstrate the benefits and limitations of the method and quantify the level of controllability the method has on the final translation under different conditions.

Список литературы