Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine

Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, Shuming Shi, Zhaopeng Tu

Introduction

ChatGPThttps://chat.openai.com is an intelligent chatting machine developed by OpenAI upon the InstructGPT Ouyang et al. (2022), which is trained to follow an instruction in a prompt and provide a detailed response. According to the official statement, ChatGPT is able to answer followup questions, admit its mistakes, challenge incorrect premises, and reject inappropriate requests due to the dialogue format. It integrates various abilities of natural language processing, including question answering, storytelling, logic reasoning, code debugging, machine translation, and so on. We are particularly interested in how ChatGPT performs for machine translation tasks, especially the gap between ChatGPT and commercial translation products (e.g., Google Translate, DeepL Translate).

In this report, we provide a preliminary study of ChatGPT on machine translation, which to our best knowledge is also the first one since the release of ChatGPT. Specifically, we focus on three aspects:

Translation Prompt: ChatGPT is essentially a large language model, which needs prompts as guidance to trigger its translation ability. The style of prompts may affect the quality of translation outputs. For example, how to mention the source or target language information matters in multilingual machine translation models, which is usually solved by attaching language tokens Johnson et al. (2017); Fan et al. (2021).

Multilingual Translation: ChatGPT is a single model handling various NLP tasks and covering different languages, which can be considered a unified multilingual machine translation model. Thus, we are curious about how ChatGPT performs on different language pairs considering both the resource difference (e.g., high vs. low) and language family (e.g., European vs. Asian).

Translation Robustness: ChatGPT is developed upon GPT3, which was trained on large-scale datasets that cover various domains. Therefore, we wonder if it can perform robustly well on domain-specific or even noisy sentences.

To trigger the translation ability of ChatGPT, we ask ChatGPT itself for advice and obtain three candidate translation prompts. By evaluating on the Chinese⇒\RightarrowEnglish translation task, we find that the candidate prompts generally work well and show minor performance differences. Nevertheless, we adopt the best-performing prompt for the rest parts of the study. By evaluating the translation among four selected languages on the Flores-101 test sets, we find that ChatGPT performs competitively with commercial translation products (e.g., Google Translate) on high-resource European languages but lags behind significantly on low-resource or distant languages. As for the translation robustness, results on three robustness sets suggest that ChatGPT does not perform as well as the commercial systems on biomedical abstracts or Reddit comments but exhibits good results on spoken language.

Further, we have a discussion on how to improve ChatGPT for machine translation. On one hand, we explore an interesting strategy named pivot prompting for distant languages, which asks ChatGPT to translate the source sentence into a high-resource pivot language before into the target language, improving the translation performance noticeably. On the other hand, with an improved engine GPT-4 OpenAI (2023)https://openai.com/research/gpt-4 launched on March 15, 2023, we re-evaluate the translation ability of ChatGPT and observe a significant boost of performance. The translation performance of ChatGPT becomes comparable to commercial translation products, even for distant languages. Extensive analysis on Google Translate and ChatGPT suggests that ChatGPT with GPT-3.5 tend to generate more hallucinations and more mis-translation errors while that with GPT-4 makes the least errors. In other words, ChatGPT has already become a good translator with GPT-4 as the engine!

ChatGPT for Machine Translation

We provide a brief introduction of the evaluation setting, which mainly includes the compared baselines and test data.

We compare ChatGPT with three commercial translation products, namely, Google Translatehttps://translate.google.com, DeepL Translatehttps://www.deepl.com/translator, and Tencent TranSmarthttps://transmart.qq.com/zh-CN/index. So far, the three commercial systems support translation in 133, 29, and 16 languages, respectively. By default, the results in this report come from the ChatGPT version on 2022.12.16. For new results, we will mark the updated version information correspondingly.

Data.

For multilingual translation, we evaluate the above translation systems on the Flores-101 Goyal et al. (2021)https://github.com/facebookresearch/flores test sets, which consists of 1012 sentences translated into 101 languages. To test the translation robustness, we adopt the test set of WMT19 Biomedical Translation Task (Bawden et al., 2019, i.e., Bio) and the set2 and set3 of WMT20 Robustness Task (Specia et al., 2020, i.e., Rob2 and Rob3). We obtain the first two test sets through SacreBLEU and the third pre-processed by Wang et al. (2021)https://github.com/hsing-wang/WMT2020_BioMedical/tree/master/Bio-18-19-testset. Table 1 lists the information of these test sets. Since this empirical study was conducted upon the very early release of ChatGPT, we were only able to access it through webpage, which cannot respond to large batches. As a result, obtaining the translation results from ChatGPT is time-consuming. Therefore, we randomly sample 50 sentences from each set for evaluation.

Metric.

We adopt the mostly used BLEU score Papineni et al. (2002) as our primary metric and also report ChrF++ Popović (2017) and TER Snover et al. (2006) in some cases. These three metrics are all supported by SacreBLEU Post (2018)https://github.com/mjpost/sacrebleu.

2 Translation Prompts

To design the prompts for triggering the machine translation ability of ChatGPT, we seek inspiration from ChatGPT by asking it for advice. Specifically, we ask ChatGPT with the following prompt:

Provide ten concise prompts or templates that can make you translate.

and obtain the results as shown in Figure 1. The generated prompts look reasonable but share similar formats. Thus, we summarize them into three candidate prompts as shown in Table 2, where [SRC] and [TGT] represent the source and target languages of translation. Note that we add an extra command into Tp2 to ask ChatGPT not to generate double quotes around the translation, which often occurs with the original format. Nevertheless, it is still unstable such that sentences in a batch (in multiple lines) are translated into a single line occasionally.

We compare the three different candidate prompts on the Chinese-to-English (Zh⇒\RightarrowEn) translation task with the test set from Flores-101. Table 3 shows the results of ChatGPT and three commercial systems. While ChatGPT provides reasonably good translations, it still lags behind the baselines by at least 5.0 BLEU points. Concerning the three candidate prompts, Tp3 performs the best in terms of all the three metrics. Thus, we use Tp3 throughout this report by default.

3 Multilingual Translation

We select four languages to evaluate the capability of ChatGPT in multilingual translation, including German (De), English (En), Romanian (Ro), and Chinese (Zh), which are commonly adopted in both research (Wang et al., 2022a; Jiao et al., 2021, 2022b) and competitions (Bojar et al., 2016; Farhad et al., 2021). The first three languages come from the same family with Latin scripts while the last is from another family with Chinese scripts Fan et al. (2021). We test the translation performance between any two languages, which involves 12 directions in total. For clarity and comparison, we report the BLEU scores and the improvement or drop of performance (i.e., +/-) relative to Google Translate. Table 4 presents the results.

We consider the resource difference of languages in the same family. In machine translation, German⇔\LeftrightarrowEnglish translation is usually regarded as a high-resource task supported by over ten million sentence pairs Farhad et al. (2021) while Romanian⇔\LeftrightarrowEnglish translation is supported by much less data Bojar et al. (2016). This resource difference can also be indicated by the data statisticshttps://github.com/openai/gpt-3/tree/master/dataset_statistics of GPT-3 Brown et al. (2020), although we do not know the data information of ChatGPT. As shown in Table 4, ChatGPT performs competitively with Google Translate and DeepL Translate for both German⇒\RightarrowEnglish and English⇒\RightarrowGerman translations. However, it lags behind them significantly on Romanian⇒\RightarrowEnglish and English⇒\RightarrowRomanian. Specifically, ChatGPT obtains a BLEU score on English⇒\RightarrowRomanian that is 46.4% lower than Google Translate and the value is 10.3% on Romanian⇒\RightarrowEnglish. We speculate that the huge resource difference of monolingual data between English and Romanian limits the language modeling capability of Romanian, which partially explains the poor performance on English⇒\RightarrowRomanian. On the contrary, Romanian⇒\RightarrowEnglish can benefit from the strong language modeling capability of English such that the resource gap of parallel data can be somewhat compensated.

Language Family.

We also take the impact of language families into account. In machine translation, translating between different language families is often considered harder than that within the same language family, due to the different cultures and writing scripts. By comparing German⇔\LeftrightarrowEnglish with Chinese⇔\LeftrightarrowEnglish or German⇔\LeftrightarrowChinese translation, we find that the gap between ChatGPT and the commercial systems becomes larger. We attribute to the better knowledge transfer within the same family (i.e., from English to German) than between different families (e.g., from English to Chinese). For language pairs that are both low-resource and from different families (e.g., Romanian⇔\LeftrightarrowChinese), the performance gap can be further enlarged Wang et al. (2022b). Since ChatGPT handles different tasks in one model, low-resource translation tasks not only compete with high-resource translation tasks Jiao et al. (2022a), but also with other NLP tasks for the model capacity, which explains their poor performance.

4 Translation Robustness

We further evaluate the translation robustness of ChatGPT on the WMT19 Bio and WMT20 Rob2 and Rob3 test sets, which introduce the impact of domain bias and potentially noisy data. For example, WMT19 Bio test set is composed of Medline abstracts, which require domain-specific knowledge to handle the terminologies. WMT20 Rob2 are comments from the social media website reddit.com that could contain various errors, including spelling/typographical errors, word omission/insertion/repetition, grammatical errors, spoken languages, Internet slang, and so on Michel and Neubig (2018).

Table 5 lists the BLEU scores. Obviously, ChatGPT does not perform as well as Google Translate or DeepL Translate on the WMT19 Bio and WMT2 Rob2 test sets. The reason may be that commercial translation systems like Google Translate often need to continuously improve their ability to translate domain-specific (e.g. biomedical) or noisy sentences, since they are real-world applications that require better generalization performance over out-of-distribution data. However, these may not be done in ChatGPT.

An interesting finding is that ChatGPT outperforms Google Translate and DeepL Translate significantly on WMT20 Rob3 test set that contains a crowdsourced speech recognition corpus. It suggests that ChatGPT, which is essentially an artificial intelligent chatting machine, is capable of generating more natural spoken languages than these commercial translation systems. We provide some examples in Table 6.

Improving ChatGPT for MT

As presented above, ChatGPT can match the performance of commercial translation systems on high-resource language pairs, but still struggles on low-resource ones, especially those distant languages. Then, one question arises:

The first way to improve ChatGPT for MT is to exploit the potential of ChatGPT in other tasks to assist the target task. Here, we explore an interesting strategy named Pivot Prompting to improve the translation quality between distant languages. Rather than the direct translation between source and target languages, we ask ChatGPT to translate the source sentence into a high-resource pivot language (i.e., English by default) first and then into the target language. Accordingly, we adjust the Tp3 prompt as below:

Please provide the [PIV] translation first and then the [TGT] translation for these sentences one by one:

where [PIV] denotes the pivot language. As a large language model, ChatGPT will naturally condition on both the prompt and the translation result in the pivot language to generate the translation into the target language. Figure 2 shows an example when using pivot prompting.

There are several advantages of pivot prompting:

Knowledge Transfer: While parallel data between two distant languages is often scarce Fan et al. (2021); Wang et al. (2022b), the parallel data between them and the pivot language can be relatively considerable, which is expected to learn better translation ability for source-pivot and pivot-target directions than that for the source-target direction. Thus, pivot prompting will potentially transfer the knowledge of the high-resource pivot language to the low-resource target languages Zoph et al. (2016); Aji et al. (2020); Li et al. (2022); He et al. (2022).

Convenience: Essentially, pivot prompting is similar to the pivot translation technique in previous studies Cheng et al. (2016) but is more convenient for ChatGPT. For the commonly adopted multilingual sequence-to-sequence translation models Fan et al. (2021), pivot translation requires two steps: (1) Input the source sentence and translate it into the pivot language; (2) Input the translation results in pivot language and translate it into the target language. In contrast, ChatGPT can identify both the [PIV] and [TGT] languages and translate the source sentence into the two languages sequentially (see Figure 2), which requires only one step operation.

Table 7 presents our results in BLUE score and length ratio of translation results over references. We obtain the translation results by using Tp3 (i.e., Direct) and pivot prompting (i.e., Pivot) through English (i.e., source-to-English-to-target), respectively. As seen, the latest update for ChatGPT seems to harm the translation quality for German⇒\RightarrowChinese and Romanian⇒\RightarrowChinese translations, compared with the previous version we used (i.e., Directnew vs. Direct). Nevertheless, pivot prompting can significantly improve the translation performance by nearly 3.9 and 6.6 BLEU points for German⇒\RightarrowChinese and Romanian⇒\RightarrowChinese translations, respectively, which demonstrates its effectiveness. By inspecting the translation results, we find that direct translation with Tp3 will under-translate some tokens in source sentences, which can be noticeably fixed by pivot prompting. This can be reflected by the length ratio results. Note that, while pivot prompting is convenient for ChatGPT, how to further accelerate the inference process is still an important research question as we need to generate longer sentences.

2 GPT-4 as the Engine

Another way to improve ChatGPT for MT is to improve its engine. Unsurprisingly, OpenAI released GPT-4 OpenAI (2023) on March 15, 2023, which exhibits all-around stronger capabilities than the GPT-3.5 model behind ChatGPT. Therefore, we re-evaluate the performance for four translation directions. As shown in Table 8, GPT-4 boosts the performance over ChatGPT significantly on all the four directions, bringing the BLEU scores to the level of top commercial translation systems. Note that these results only come from zero-shot settings. With modern techniques like in-context learning with demonstrations Brown et al. (2020); Agrawal et al. (2022), the translation performance could be further improved. In other words, GPT-4 has already become a good translator!

Analysis

Here we conduct some analyses on the translation outputs for a deeper understanding in ChatGPT. By default, we analyze the outputs of Google, ChatGPT, and GPT-4 on Zh⇒\RightarrowEn translation for all the 50 test examples.

We follow previous studies Jiao et al. (2021); Wang et al. (2022a) to analyze the translation outputs using automatic tools, i.e., compare-mthttps://github.com/neulab/compare-mt, at both word level and sentence level.

Essentially, ChatGPT is a large language model that has been trained on a variety of corpora, covering different domains. It could be beneficial to the translation of low-frequency words in the test sets. Specifically, we divide the target words into three categories based on their frequency and calculate the accuracy of word prediction. Table 9 shows the F-measure results. Unexpectedly, ChatGPT turns out to perform the worst on low-frequency words (i.e., <2<2), which we attribute to the immature translation ability of ChatGPT. What’s interesting is that GPT-4 mainly addresses this shortcoming for ChatGPT with little improvement to higher-frequency words.

Sentence Length.

ChatGPT is also trained for various text generation tasks, which usually do not require strict length constraints of generated sentences as machine translation. Therefore, we are curious about how sensitive the translation performance is to the sentence length. We divide the target sentences into three categories based on the sentence length, of which the average value is 23.2 tokens. Table 10 shows the results. As seen, ChatGPT performs the worst on short sentences (i.e., <15<15), with 18.8 BLEU points lower than Google Translate. One observation is that when translating terminologies, e.g., 美国公共广播公司, ChatGPT tends to output the full names (i.e., American Public Broadcasting System) while Google Translate and the reference use the abbreviations (i.e., PBS). As a result, the precision of word prediction will be reduced noticeably, so will BLEU score Papineni et al. (2002), especially for short sentences. GPT-4 can predict the abbreviations properly sometimes, which gives a better translation performance.

2 Human Analysis

In addition to the automatic analysis, we also inspect the translation outputs manually. We ask three annotators to identify the errors in the translation outputs Wang et al. (2022a), including under-translation (i.e., Und-Trans), over-translation (i.e., Ove-Trans), and mis-translation (i.e., Mis-Trans). Based on the translation errors, the annotators rank the translation outputs of Google, ChatGPT and GPT-4 accordingly, with 1 as the best system and 3 as the worst. For translation outputs that are really hard to distinguish, we allow the same ranking (e.g., 1-1-1, 1-1-2 or 1-2-2). To eliminate subjective bias, we do not present the system information of each translation output to the annotators, and the three translation outputs for each test example are also shuffled randomly.

Table 11 shows the results of translation errors. Generally, ChatGPT makes more over-translation errors and mis-translation errors than Google Translate, but slightly less under-translation errors. It suggests that ChatGPT is more likely to generate hallucinations. In contrast, GPT-4 makes the least errors across the three error classes, which demonstrates the best translation performance. This is also confirmed by the ranking results in Table 12, such that GPT-4 is ranked the best (i.e., 1) for 32 times out of 50 test examples, followed by Google Translate and ChatGPT. However, the BLEU score of GPT-4 is still lower than that of Google Translate (i.e., 28.50 vs. 31.66 in Table 8), which indicates that GPT-4 may generate more diverse translations with different lexical choices from the references.

3 Case Study

We present four test examples in Table 13 for an intuitive understanding. The first example shows the hallucination of ChatGPT at the first few tokens and the inaccurate translation of 过量降水. The second example shows that both ChatGPT and GPT-4 translate 广泛耐药结核病 into the full name while the reference and Google Translate do not. The third example shows that GPT-4 can also translate the terminology 美国公共广播公司 into the abbreviation. The last example suggests that GPT-4 is able to translate the terminology 狼孩 more properly based on the context while Google Translate and ChatGPT fail to.

Conclusion

This work presents a preliminary study of ChatGPT for machine translation. We find that ChatGPT performs competitively with commercial translation products (e.g., Google Translate) on high-resource European languages but lags behind significantly on low-resource or distant languages. It also exhibits good results on spoken language while still performs worse than commercial systems on biomedical abstracts or Reddit comments. We further explore an interesting strategy named pivot prompting that can improve the translation performance of distant languages noticeably. With the launch of the GPT-4 engine, the translation performance of ChatGPT is significantly boosted, becoming comparable to commercial translation products, even for distant languages. Extensive human analysis suggests that, ChatGPT has already become a good translator with GPT-4 as the Engine.

Limitations

As a preliminary study, this work is far from complete with various aspects to make it more reliable:

Comprehensiveness: Currently, we randomly select 50 samples from each test set for evaluation due to the response delay of ChatGPT, which is not comprehensive due to the data coverage. Besides, we found that the results of the same query may vary across multiple trials, bringing randomness to the evaluation results. For more reliable results, it is best to repeat the translation multiple times for each test set and report the average result.

Translation Abilities: We only focus on multilingual translation and translation robustness in this report. However, there are some other translation abilities that can be further evaluated, e.g., constrained machine translation and document-level machine translation.

References