Prompting Large Language Model for Machine Translation: A Case Study

Biao Zhang, Barry Haddow, Alexandra Birch

Introduction

Large language models (LLMs) pretrained on massive unlabeled corpora have shown impressive emergent abilities under model scaling which enable prompting for downstream applications Brown et al. (2020); Kaplan et al. (2020); Wei et al. (2022b); Zhang et al. (2022a); Chowdhery et al. (2022). Different from task-specific finetuning, prompting constructs task-specific prompts by rephrasing test examples with descriptive task instructions and executes the task by feeding prompts to LLMs directly. It can be further enhanced through in-context learning by providing a few labeled examples (or prompt examples) as a demonstration Brown et al. (2020). As a new paradigm, prompting LLMs has achieved state-of-the-art performance over a range of natural language processing (NLP) tasks Chung et al. (2022); Goyal et al. (2022); Wei et al. (2022c); Chowdhery et al. (2022).

In this paper, we focus on prompting LLMs for machine translation (MT). MT represents a complex task requiring transforming a source input into its semantically equivalent target output in a different language, which combines sequence understanding and generation. It offers a unique platform to assess the cross-lingual generation capability of LLMs, and the assessment may shed light on pretraining/finetuning algorithm design for achieving universal LLMs Chowdhery et al. (2022). While a few studies have reported translation results Brown et al. (2020); Chowdhery et al. (2022), a systematic study on how prompting works for MT is still missing in the literature.

We aim at filling this gap by thoroughly examining different prompting setups using the recently released GLM-130B Zeng et al. (2022), particularly concerning three aspects: the prompting strategy, the use of unlabeled/monolingual data, and the feasibility of transfer learning. Prompting has shown varying sensitivity to the choice of prompt templates and examples Zhao et al. (2021). For MT, prior studies adopted different templates Brown et al. (2020); Wei et al. (2022a); Chowdhery et al. (2022), and we reevaluate them to figure out the optimal one. We further design a set of features for prompt examples and explore which one(s) could explain the prompting performance, according to which we develop the example selection strategy.

Since leveraging monolingual data to improve MT has long been of interest, we would like to determine whether and how such data can be used in prompt example construction. We make a step in this direction by studying the effect of data augmentation using back-/forward-translation Sennrich et al. (2016b); Zhang and Zong (2016) via zero-shot prompting. In addition, neural MT and pretrained LLMs have shown encouraging transfer abilities Devlin et al. (2019); Arivazhagan et al. (2019); Zhang et al. (2020); Xue et al. (2021) but transfer learning for prompting has received little attention. Whether prompt examples are transferable across different settings, such as from one domain/language pair to another and from sentence-level examples to document-level translation, is yet to be addressed.

We address the above concerns with GLM-130B as the testbed and conduct extensive experiments on FLORES and WMT evaluation sets. We mainly study translation for three languages: English, German and Chinese. We also provide a quantitative and qualitative analysis to disclose problems when prompting for MT, which might offer insights for future study. Our main findings are listed as below:

Prompting performance varies greatly across templates, and language-specific templates mainly work when translating into languages LLMs are pretrained on. An English template in a simple form works best for MT.

Several features of prompt examples, such as sequence length, language model score, and semantic similarity, correlate significantly with its prompting performance while the correlation strength is weak in general. Selecting examples based on these features can outperform the random strategy, but not consistently.

Using monolingual examples for prompting hurts translation. By contrast, constructing pseudo parallel examples via back-/forward-translation is a good option. Back-translation performs better and is more robust.

Prompting shows some degree of transferability. Using demonstrations from other settings can improve translation over the zero-shot counterpart, while the superiority of a demonstration in one setting can hardly generalize to another.

Prompting for MT still suffers from copying, mistranslation of entities, hallucination, inferior direct non-English translation, and prompt trap where translating the prompt itself via prompting becomes non-trivial.

Setup

Given a pretrained and fixed LLM L\mathcal{L}, MT prompting first converts each test input XX to a prompt according to a template T\mathcal{T} and then generate the translation YY by feeding the prompt to L\mathcal{L}. In this study, we consider zero-shot and few-shot prompting for translation.

Zero-shot prompting only has access to the test input XX, while few-shot prompting assumes that a few extra labeled examples (or prompt/demonstration examples) DP={Xi′,Yi′}i=1K\mathcal{D}^{P}=\{X_{i}^{\prime},Y_{i}^{\prime}\}_{i=1}^{K} are available and can be used as a demonstration. Particularly, we adopt the following template for zero-shot prompting based on the results in Section 3:

where [src] and [tgt] denote test language(s), i.e., the source and target language name of the test language pair, respectively. For few-show prompting, we concatenate the given prompt examples:

where [psrc] and [ptgt] denote prompt language(s), i.e., the source and target language name of the prompt example, respectively. By default, prompt examples and test data are in the same language pair. However, when considering cross-lingual transfer for prompting, prompt examples might be in a different language pair.

We also explore template language, which denotes the language in which the template is expressed. For example, the Chinese template “ 中文:XX 英文: ” represents the Chinese counterpart of the following English template “Chinese: XX English: ”.

Setting

We experiment with GLM-130B, a LLM with 130B parameters pretrained on Chinese and English monolingual corpora, which was reported to outperform GPT-3 and OPT-175B on several NLP tasks Zeng et al. (2022). Note GLM-130B is a raw LLM without any further finetuning. We use its INT4-quantized version, which is more affordable and suffers little performance degradation. We adopt beam search for MT with a beam size of 2, and perform experiments with 4 RTX 3090 and A100-40G GPUs.

We work on three languages: English (En), German (De), and Chinese (Zh). We perform major analysis on FLORES (Wiki domain, En-De-Zh, NLLB Team et al., 2022) and WMT21 (News domain, En-De, En-Zh, Akhbardeh et al., 2021), and also report results on Multi-Domain (IT, Law and Medical domain, De-En, Aharoni and Goldberg, 2020) to examine domain robustness and transfer ability, and PDC (News domain, Zh→\rightarrowEn, Sun et al., 2022) for document-level translation. To understand the relation between prompt examples and their prompting performance, we construct an Ablation set for Wiki, WMT and Multi-Domain (IT and Medical) based on the dev set of FLORES, WMT21 and Multi-Domain, separately, where we randomly sample 100 instances as the ablation test set and use the rest as the default example selection pool. To distinguish, we will refer to the official dev and test set as Full set. Detailed statistics are listed in Table 9, Appendix.

We evaluate translation performance using both a surface-based metric, detokenized BLEU↑\uparrow from SacreBLEU Post (2018), and a model-based metric, COMET↑\uparrow from unbabel-comet with wmt20-comet-da Rei et al. (2020).

Prompting Strategy for MT

To perform MT, prompting needs to cast the translation problem into a language modeling problem via the prompt. Thus, the format of the prompt, including its wording, directly affects how LLM understands the task and its behavior. For MT, we are interested in the following research questions:

Which template should we use for MT prompting? And what language for the template?

Does demonstration matter for MT prompting? How to select optimal prompt examples?

We address them through extensive experiments on Wiki Ablation sets.

We start with zero-shot prompting and explore the effect of different templates. Depending on how to describe MT and partially inspired by prior studies Brown et al. (2020); Chowdhery et al. (2022); Wei et al. (2022a), we compare 6 templates and evaluate them on the Wiki Ablation sets covering 6 language pairs (En↔\leftrightarrowDe, En↔\leftrightarrowZh, De↔\leftrightarrowZh). Table 1 shows the results (we list detailed results in Table 10, Appendix). The template affects zero-shot quality substantially, and the simple template Ⓐ in English specifying just the source and target language name achieves the best overall results. In follow-up experiments, we thus focus on template Ⓐ.

Language-specific template delivers mixed results.

Table 1 also shows the prompting results of German and Chinese templates, which often largely underperform their English counterparts. Since German is not a major pretraining language in GLM-130B, a German template degenerates the translation substantially. By contrast, a Chinese template yields improved performance when translating into Chinese (see Table 10). Still, an English template works best on average.

The preference of GLM-130B to English template also shows that the level of language understanding and cross-lingual ability in GLM-130B varies across languages, even though it’s pretrained on the same amount of monolingual Chinese and English tokens. This might be caused by the fact that English is used more globally than Chinese, but might also suggest that improving the language understanding of LLM requires more advanced training algorithms beyond scaling training data.

Using more prompt examples for demonstration improves translation significantly on average.

We next study few-shot prompting following the template Ⓐ but in format (2) with KK varying from 1 to 20. We evaluate multiple demonstrations for each KK via random sampling to reduce data biases. Figure 1 shows that the more examples used, the better average performance (more results are shown in Figure 5, Appendix), albeit at the cost of using more GPU memory and increasing the inference time per token as in Figure 3.

The performance of demonstration is not stable.

However, we also see high performance variance under the same KK. It’s possible that a demonstration with 5 examples outperforms its 10 or 20 counterpart. Figure 1 also shows that 1-shot prompting underperforms zero-shot prompting in many cases, even on average. This echoes with previous findings on other NLP tasks Zhao et al. (2021); Liu et al. (2022) and also highlights the significance of developing effective example selection strategies.

Note that few-shot prompting greatly improves translation into Chinese. The reason based on our manual analysis is that the zero-shot baseline tends to translate into traditional Chinese with messy codes, where prompt examples help (the reference text is always simplified Chinese).

Several features correlate with prompting performance significantly yet weakly.

We thus turn to explore example selection for prompting. Our idea is to extract a couple of diverse features from demonstration and examine whether any of them are informative enough to be used as an indicator for the selection. In this study, we simplify our analysis by focusing on 1-shot prompting, which ignores the ordering of prompt examples (we leave few-shot analysis to future). Particularly, we extract and analyze 7 features of a demonstration:

GLM-130B-based, length-normalized log likelihood of the demonstration;

translation quality of the prompt example from COMET QE model wmt20-comet-qe-da Rei et al. (2020);

semantic score based on the cosine similarity of the demonstration’s source and target sentence embeddings from LASER2 Heffernan et al. (2022);

similarity to the input that averages over SemScores between the test input and the demonstration’s source;

similar to CaseSemScore-Src but compares to demonstration’s target;

We sample multiple demonstrations randomly and inspect the Spearman’s correlation between feature values and prompting performance. We consider high-quality and low-quality pool for sampling.

Table 2 summarizes the results and Figure 2 illustrates the relation between COMET and LMScore (more results are given in Table 11 and Figures 6, 7, Appendix). With the high-quality pool, different demonstrations yield similar translation results (see blue points) despite their feature values varying greatly. Several features show insignificant and inconsistent correlation, particularly for De→\rightarrowEn and Zh→\rightarrowEn. This suggests developing selection policy for high-quality example pool is non-trivial.

After mixing with demonstrations from the low-quality pool, the significance gets strengthened. LMScore and CaseSemScore-Tgt shows the highest correlation on average followed by TLength and SemScore. MTScore behaves much worse which might be caused by its instability on sentence-level evaluation Moghe et al. (2022). However, we didn’t see significant difference in terms of Spearman’s ρ\rho between input-relevant and input-agnostic features Agrawal et al. (2022), neither among surface-based, LLM-based or semantic-based features. Surprisingly, the simple feature, S/TLength, yields reasonably high correlation. We argue that long examples could offer LLM with more signals about the task’s input and output space. This finding suggests that researchers should select long unlabeled sentences for annotation to improve prompting. Yet, most Spearman’s ρ\rhos are much smaller than 0.5, indicating a weak/fragile relation.

In general, selecting prompt examples of high translation quality, high semantic similarity, high LLM likelihood, long sequence length and high similarity to test inputs are all preferable strategies. Unfortunately, none of them can guarantee optimal translation performance.

Using prompt examples selected based on the proposed features yields improved performance.

We next verify the above findings on the Full sets. We explore selection strategies based on SemScore, LMScore and TLength (i.e. use top-ranked examples) as they show high average correlation. We didn’t analyze CaseSemScore-Tgt as it’s more complicated and doesn’t make significant difference. Note we excluded too long (more than 100 tokens) or too short (less than 10 tokens) examples during selection. We also consider 5-shot prompting, where we concatenate top-ranked 5 examples in an ascending order Liu et al. (2022).

Table 3 shows that, with high-quality pool, adopting the feature-based strategy is likely to outperform the random baseline, and the SemScore-based strategy performs well across different settings (detailed results are available in Table 13 and 14, Appendix). These strategies also generalize to 5-shot prompting to some extent. For selection from low-quality pool, we propose a combined strategy: we first choose top-11K examples according to SemScore to filter out poor examples, the top-1K of which are also dropped as they tend to be uninformative (see Table 12 in Appendix); then we re-rank the rest with LMScore and retain top-1K examples, upon which we further apply the TLength-based strategy. In Table 3, this combined strategy outperforms the random one by varying degrees.

Monolingual Data for Prompting

A longstanding concern in MT is how to utilize unlabeled data to improve translation. While prompting enables few-shot learning reducing the data requirement, exploring whether demonstration could benefit from monolingual examples is still valuable, both for MT study and for understanding of the role of demonstration in prompting.

Min et al. (2022) argue that the key role of demonstration lies in its support of the input space, the label space and the prompt format, rather than the genuineness of the examples. They found that randomly replacing labels in demonstration barely hurts performance on classification tasks. We reexamine this argument in the context of MT by studying the following three prompting settings: 1) random examples constructing sentence pairs from monolingual sources and targets randomly; 2) source/target example only using monolingual source/target alone for prompting.

Figure 4 (top) shows a totally different story (see Figures 8 and 9 in Appendix for more results): monolingual example-based demonstration almost always hurts translation, and the more examples used, the more degeneration yielded. Using random examples misleads the prompting and performs the worst in general; compared to target-only examples, using source examples yields slightly better results except translating into Chinese. This indicates that the genuine source-target mapping should be retained in the demonstration, and also indicates that MT features unique challenges which deserves more attention when studying prompting.

Pseudo parallel examples by forward-/back-translation benefits prompting.

Inspired by data augmentation in MT Sennrich et al. (2016b); Zhang and Zong (2016), we next resort to constructing pseudo parallel data. We first adopt GLM-130B to translate the source or target examples via zero-shot prompting, and then use the generated parallel examples as demonstration. Despite low quality, Figure 4 (bottom) shows that this is an effective way to improve prompting, and using more examples often produces better results. We also observe that back-translation (i.e. translating target monolingual examples) performs better and behaves more robustly than forward-translation (i.e. translating source examples instead), which even approaches prompting with real parallel examples.

Transfer Learning for Prompting

After obtaining a performant demonstration, we are interested in to what extent its capability could be transferred across different settings, especially from one domain/language pair to another and from sentence-level to document-level translation. While previous studies demonstrate the feasibility with continuous prompts on classification tasks Wang et al. (2021), transfer for hard prompting on MT has never been investigated.

Assume that demonstrations D1D_{1} and D2D_{2} are selected in setting S1S_{1} and that D1D_{1} performs better (i.e. D1>D2D_{1}>D_{2}), We have the following research questions:

Could we also expect D1>D2D_{1}>D_{2} in setting S2S_{2}?

Whether using demonstrations from S1S_{1} could outperform zero-shot prompting in S2S_{2}?

We next study these questions through experiments with 1-shot prompting.

If the ranking D1>D2D_{1}>D_{2} holds across settings, the results of the same set of demonstrations in different settings should show high and significant Spearman’s correlation. Unfortunately, the correlations in Table 4 and 5 are very weak and often insignificant (more results are given in Table 15, 16, and 17), even for the same language pairs in different directions (Reversed) and for similar domains (Wiki⇒\RightarrowWMT). This suggests that we will need setting-specific demonstration to get the optimal translation quality.

Using out-of-setting demonstrations can benefit translation.

However, we can still gain from using out-of-setting demonstrations as demonstrated by the positive gains in Table 4 and 5, where we find that transfer in target-shared and reversed settings is relatively easier, and that transfer across distant domains can be successful particularly when in-setting example pool is of low quality. This is also supported by the transfer to document-level translation, where both BLEU and document-specific evaluation get improved as shown in Table 6. Results in Table 19 show that the transfer is unstable and could deliver negative results, i.e. worse than zero-shot prompting, partially resonating with previous findings Lin et al. (2021). We leave the question of how to select prompt examples in transfer learning setups to future.

Discussion

Although prompting enables translation with decent performance, it still suffers from many (well-known) problems. Here, we briefly explain the problems we observed from the model’s outputs.

Prompting sometimes rejects translating the input. Instead, it emits either empty or off-target outputs, i.e. translating in a wrong target language. This occurs frequently when translating into Chinese, where the model often translates into traditional Chinese with messy codes, causing unstable performance. Besides overly relying on a language model, prompting tends to under-translate the input, copy source phrases, produce code-switched output, mistranslate entities (e.g. dates) and generate hallucination, as illustrated in Table 7.

We also observe a phenomenon specific to prompting: prompt trap where prompting behaves unpredictable when its input is mixed with prompt template phrases. In the second case in Table 7, the model copies the template phrases, rather than translating them into Chinese. This means that translating prompt itself (not just the input) becomes non-trivial, and that users may attack prompting-based translation systems by manipulating the input format.

We find that the translation quality between German and Chinese is very poor (see Table 13). We argue that the cross-lingual ability of GLM-130B mainly centers around English (although GLM-130B was pretrained on Chinese as well), and thus explore pivoting translation instead. Table 8 shows that pivoting through English greatly improves non-English translation. It’s still unclear whether the current LLM pretraining recipe could achieve promising non-English-centric cross-lingual ability. We might need to consider adding parallel data into the LLM pretraining or finetuning.

Related Work

The capability of prompting heavily depends on its surface representation, where small modifications to the prompt could cause high variance in its performance. This inspires researchers to develop advanced prompting strategies to get the most from LLMs. Gao et al. (2021) proposed to generate prompt templates automatically using T5 Xue et al. (2021) rather than adopting manual templates. Liu et al. (2022) reported selecting prompt examples close to the test input via a kkNN-based retriever, Sorensen et al. (2022) resorted to an information-theoretic approach based on mutual information, while Zhang et al. (2022b) formulated example selection as a sequential decision problem and solved it by reinforcement learning. For reasoning tasks, Wei et al. (2022c) developed chain-of-thought (CoT) prompting letting the model output the intermediate reasoning steps, which inspires researchers to further explore CoT selection Fu et al. (2022) and decomposition Zhou et al. (2022). In contrast to the studies just mentioned, which focus on NLP tasks other than MT, we explore prompting strategies exclusively for translation.

Prompting uses instructions to guide LLMs, which is closely related to neural MT with special prefixes. In multilingual NMT, a target language tag is often appended to the source input to indicate the translation direction Johnson et al. (2017); Arivazhagan et al. (2019); Zhang et al. (2020). Special attribute tags can also be used to control properties of the model output, such as politeness Sennrich et al. (2016a), diversity Shu et al. (2019), and quality Caswell et al. (2019). Besides, retrieved phrases and sentences can be augmented to the input to improve translation quality Zhang et al. (2018); Gu et al. (2018). With the popularity of prompting LLMs, researchers see value in incorporating prompts into neural MT Li et al. (2022); Tan et al. (2021); Garcia and Firat (2022). Still, these methods rely on pretraining or finetuning the model rather than prompting frozen LLMs.

Very recently, concurrent to our work, Vilar et al. (2022) examined the capability of prompting PaLM for translation and discovered that prompting with high-quality examples even chosen randomly performs on par with or better than the one using input-relevant examples. By contrast, Agrawal et al. (2022) explored strategies to select input-specific examples, and observed that input-relevant examples based on n-gram overlap significantly improves the capability of prompts. Our study resonates with both their findings and also explains their conflict: while the quality and input-based semantic similarity correlate with prompting performance significantly, the correlation strength is unfortunately not strong enough so using them as indicators to select examples may produce mixed results. Note that apart from example selection, we also studied using monolingual data and transfer learning for MT prompting, which, to the best of our knowledge, have never been explored before.

Conclusion and Future Work

In this paper, we presented a systematic study on prompting for MT, exploring topics ranging from prompting strategy, the use of unlabelled monolingual data, to transfer learning. We found that prompt template and demonstration example selection both have substantial impact on translation. Some prompt example features correlate significantly with prompting performance; treating them as criteria for example selection benefits translation to some extent but not consistently as the correlations are not strong enough.

Prompting for MT requires retaining the source-target mapping signals in the demonstration. Directly applying monolingual data for prompting sounds interesting but doesn’t work. Constructing pseudo parallel prompt examples by back-/forward-translation via zero-shot prompting is a simple yet effective solution. Regarding transfer learning, we saw positive results when applying a (sentence-level) demonstration to other domains, other language pairs or document-level translation. Unfortunately, the optimality of the demonstration doesn’t generalize across settings and the transfer performance is also unstable. We argue that MT provides a set of unique challenges and call for more efforts on evaluating prompting LLMs for MT.

Prompting also faces a number of other issues, like off-target generation and prompt traps, which we plan to address in the future. We are also interested in examining whether our findings can generalize to other LLMs, like GPT-3, OPT and PaLM. We would also like to explore further how to improve the cross-lingual ability in LLM.

Limitations

Our study heavily depends on the INT-4 quantized GLM-130B, which, unlike GPT and PaLM, was pretrained with both bidirectional and unidirectional training objectives. The quantization might weaken the model’s capability and deteriorate some unknown aspects. It’s unclear how our findings generalize to other pretrained LLMs. In addition, we mainly work on three languages due to resource constraints, and in experiments, results vary greatly across language pairs. Increasing the coverage of experimental languages would make the results more reliable.

Acknowledgments

This work was funded by UK Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guarantee [grant number 10039436 – UTTER]. The computations described in this research were performed using the Baskerville Tier 2 HPC service (https://www.baskerville.ac.uk/). Baskerville was funded by the EPSRC and UKRI through the World Class Labs scheme (EP/T022221/1) and the Digital Research Infras-tructure programme (EP/W032244/1) and is operated by Advanced Research Computing at the University of Birmingham.

References

Appendix A Appendix