A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models
Chenyang Lyu, Zefeng Du, Jitao Xu, Yitao Duan, Minghao Wu, Teresa Lynn, Alham Fikri Aji, Derek F. Wong, Siyou Liu, Longyue Wang
Introduction
Machine Translation (MT), especially Neural Machine Translation (NMT, Bahdanau et al., 2015; Vaswani et al., 2017; Castilho et al., 2017; Stahlberg, 2020; He et al., 2022b; Kocmi et al., 2022) is a fundamental task in natural language processing (NLP) that aims to automatically translate text from one language to another. Despite decades of research, MT still faces many challenges, such as dealing with idiomatic expressions, low-resource translation, handling rare words, and maintaining coherence and fluency in the translation He et al. (2022a). Recently, the emergence of Large Language Models (LLMs), such as GPT-3 and ChatGPT (Brown et al., 2020; Chen et al., 2021; Ouyang et al., 2022; Wei et al., 2022), has significantly advanced the state-of-the-art in MT. The zero-shot MT performance of LLMs is even on par with strong fully supervised MT systems while LLMs can be used in various scenarios beyond MT (Wei et al., 2022; Jiao et al., 2023b; Wang et al., 2023).
However, MT using LLMs also poses new challenges and opportunities that require new directions and methodologies. In this paper, we brainstorm several interesting directions for MT using LLMs, including stylized MT, interactive MT, and Translation Memory (TM) based MT, as well as a potential new evaluation paradigm of translation quality using LLMs. Stylized MT (Sennrich et al., 2016; Niu and Carpuat, 2020) aims to preserve the stylistic features of the source text in the translation output, such as the tone, register, formality or genre. Interactive MT (Knowles and Koehn, 2016; Santy et al., 2019) aims to facilitate the collaboration and feedback between human users and MT systems, such as through chatbots or question-answering systems. TM-based MT (Bulte and Tezcan, 2019; Xu et al., 2020) tends to make use of similar translations retrieved from the TM to improve the MT performance. The new evaluation paradigm using LLMs aims to leverage the power of LLMs for more accurate and efficient evaluation of MT systems from various aspects instead of only evaluating the similarity between system outputs and references.
In addition to the new directions and methodologies, we also discuss the privacy concerns in MT using LLMs and propose basic privacy-preserving methods to mitigate the risks. Privacy in LLM-based MT is becoming increasingly important, as LLMs may inadvertently reveal sensitive information in the source text or the translation output.
To preliminarily investigate the feasibility of the interesting directions mentioned above, we present several examples using ChatGPT for MT under various scenarios, demonstrating the feasibility of the directions. Our results demonstrate the potential of the potential new directions and methodologies for enhancing the quality and diversity of MT output, as well as the importance and challenges of privacy in MT using LLMs. We conclude by highlighting the opportunities and challenges for future research in MT using LLMs and suggesting potential directions for further exploration.
Stylized MT
Stylized MT refers to the ability to generate translations that match a specific style or genre (Toshevska and Gievska, 2022; Wang et al., 2022), such as formal or informal language (Sennrich et al., 2016), poetry or prose, or different dialects or registers. This can be achieved by training MT systems on multi-parallel data that contain translations in different styles or genres, or by using style transfer techniques (Yang et al., 2018; Jin et al., 2022) that can transform a given translation into a desired style. Stylized MT has many potential applications, such as in marketing, literature, or cultural preservation.
However, stylized MT is difficult to achieve before the presence of LLMs as there lacks of such parallel corpora for stylized MT to fit various styles while the zero-shot ability of LLMs make such tasks feasible. We can directly prompt LLMs to translate text with specific style expressed by natural language or we can firstly let LLMs translate the original text and then stylize the translation output. We present an example of translating an introduction for Olympic game from Wikipedia from English to Chinese while following poetic style in Figure 2. This example shows that ChatGPT can handle translation with poetic style, which can be hardly achieved by conventional MT systems.
Nevertheless, stylized MT also poses several challenges. One challenge is how to define and measure different styles or genres in a systematic and scalable way. Another challenge is how to evaluate the quality of stylized MT, as traditional evaluation metrics may not be sufficient to capture the diversity of stylistic variations. Overcoming these challenges requires interdisciplinary collaboration between linguists, literary scholars, and computer scientists.
Interactive MT
Interactive MT Santy et al. (2019); Jiao et al. (2023a) allows users to actively participate in the translation process, either by correcting or refining automatic translations or by providing feedback on the translation quality. This can be achieved by integrating MT systems based on LLMs with interactive user interfaces, such as chatbots or online forums, that allow users to engage with the translation process in real time to provide feedback and more specific requirements. Interactive MT can help to improve the accuracy and fluency of translations, especially in cases where the source language is ambiguous or the domain knowledge is limited.
However, interactive MT also raises several challenges. One challenge is how to design user interfaces that are intuitive and user-friendly, yet also informative and flexible. Another challenge is how to incorporate user feedback into the translation process in a principled and effective way. Overcoming these challenges requires insights from human-computer interaction, NLP, and user experience design.
Translation Memory-based MT
TM has been used for decades to help translators in basic Computer-Aided Translation systems. The general process of using TM in MT is, for a sentence to be translate, to first search for similar translations in TM using, for instance, fuzzy matching techniques, then revised or edit the retrieved similar translation in order to obtain a high quality translation. TM-based MT has already been integrated into conventional NMT systems (Bulte and Tezcan, 2019; Xu et al., 2020; Cai et al., 2021). The use of retrieved similar sentence pairs (Pham et al., 2020) seems to be a natural fit to few-shot prompting techniques when performing MT using LLMs. LLMs has emerged the In-Context Learning (ICL) ability that they can learn specific tasks through task examples given in the prompt.
However, existing works so far have mostly used randomly selected translation examples as prompts and suggest that using semantically similar examples does not significantly further improve the translation performance (Vilar et al., 2022; Zhu et al., 2023). Most of these works used sentence level embedding built by an external model to retrieve similar examples via embedding similarity search. On the contrary, other studies using fuzzy match to retrieve similar translations have shown significant improvements (Moslem et al., 2023). Therefore, the conclusion about the effectiveness of using similar translations in MT using LLMs still remains unclear. Since TMs can provide useful domain and style information that can help LLMs to generate translations that better meets the translation requirement, it is a promising direction to further study how to better integrate TMs into LLMs for MT. Figure 4 illustrates an example of prompting LLM with TMs.
Previous studies on conventional TM-based MT has also shown that conventional Transformer-based NMT system already shows the ability to make use of new TMs that has never been seen by the model during training to largely improve domain-specific translation during inference (Xu et al., 2020, 2022). This indicates that conventional NMT systems learns to understand the relationship between a given source sentence and a similar translation and to select useful information from the given similar translation, rather than simply remember sentences seen during training. This ability is, to some extent, similar to the ICL ability of LLMs. However, to the best of our knowledge, there does not exist research works focusing on finding the relationships between these two abilities.
New Evaluation Paradigm for MT using LLM
Evaluating the quality of MT using LLMs is a challenging task, as existing evaluation metrics may not be sufficient to capture the full range of translation quality. In addition, existing open-access test sets may suffer from the data contamination problem as they are possibly used during the training process of LLMs. Evaluating on these test sets cannot correctly reflect the MT performance of LLMs. A new evaluation paradigm for MT using LLMs should take into account the unique characteristics of LLM-based MT, such as the ability to generate fluent but inaccurate translations or the sensitivity to domain-specific knowledge. Possible approaches to a new evaluation paradigm include using specifically-designed human evaluations Graham et al. (2020); Ji et al. (2022) for such systems, or even directly employ LLMs to evaluate the translation outputfrom LLMs (Kocmi and Federmann, 2023) - although studies show that LLMs would prefer the translation output from LLMs instead of other systems Liu et al. (2023). Besides, we can also use extrinsic evaluation - we use the translation output in other tasks Lyu et al. (2021) and measure the corresponding performance instead directly assessing the translation quality. An example of using ChatGPT to evaluate the translation output for a tweet from Elon Musk is shown in Figure 5.
However, developing a new evaluation paradigm also poses several challenges. One challenge is how to balance the trade-off between evaluation efficiency and evaluation quality, as human evaluations can be time-consuming and expensive and LLM-based evaluation can be biased. Another challenge is how to ensure the reliability and validity of the evaluation results, as different evaluators may have different subjective judgments or biases. Overcoming these challenges requires rigorous experimental design, statistical analysis, and transparency in reporting.
Privacy in MT using LLM
As LLMs become more powerful and widely used in MT, there are growing concerns about privacy and security Xie et al. (2023). In particular, LLMs may inadvertently reveal sensitive information in the source text or the translation output, such as personally identifiable information, confidential business data, or political opinions. Privacy in MT using LLMs aims to mitigate these risks by developing privacy-preserving methods that can protect the confidentiality and integrity of the translation process.
One basic approach to privacy in MT using LLMs is to anonymize sensitive information in the textual input and then pass it to LLMs and get the output, which is then de-anonymized. An example of such issue using LLMs is shown in Figure 6. This is similar to methods integrating terminologies or user dictionary into conventional NMT systems (Crego et al., 2016).
However, privacy in MT using LLMs also poses several challenges. One challenge is how to balance the trade-off between privacy and accuracy, as privacy-preserving methods may introduce additional noise or distortion to the translation output. Another challenge is how to ensure the interoperability and compatibility of privacy-preserving methods across different languages, models, and platforms. Overcoming these challenges requires collaboration between experts in cryptography, privacy, and MT, as well as adherence to ethical and legal standards.
Future Directions
Personalized MT Mirkin and Meunier (2015); Rabinovich et al. (2017) - With the advancements in LLM-based MT, the focus can be shifted towards personalized MT. This approach can enable the provision of customized translations that are tailored to each user’s preferences and needs. It can include translations that are adapted to the user’s language proficiency, domain-specific terminology, or cultural references. One possible approach to provide personalized MT is to prompt LLMs with user-specific preferences or metadata, such as user search histories or social media posts - in other words, incorporating more context when translating text Wang et al. (2017). The zero-shot ability of LLMs makes such tasks feasible, which is difficult to achieve in previous MT systems because such data is usually unavailable. However, personalized MT raises several challenges. One challenge is how to collect and store user-specific data in a privacy-preserving manner. Another challenge is how to measure the effectiveness of personalized MT, as traditional evaluation metrics may not capture the nuances of user preferences and needs. Overcoming these challenges requires careful consideration of ethical, legal, and technical issues.
Multi-modal MT Yao and Wan (2020); Sulubacak et al. (2020) - Another promising direction is multi-modal MT, which involves integrating visual, audio, or other non-textual information into the translation process. This approach can enhance the quality and accuracy of translations in various settings, such as image or video captioning, speech recognition, and sign language translation. LLMs, such as GPT-4 OpenAI (2023), can be employed to develop models that can learn from multi-modal data and generate translations that accurately convey the meaning of the input. However, multi-modal MT poses several challenges, such as data heterogeneity, unbalanced datasets, and domain specificity. Overcoming these challenges would require developing novel algorithms that can learn from multi-modal data and generalize well across different modalities and domains.
Conclusion
In this paper, we discussed several interesting and promising research directions for MT under the scenario of using LLMs. We demonstrated case examples for stylized MT, interactive MT, TM-based MT, and a new evaluation paradigm for MT using LLMs, as well as preserving user-privacy in LLM-based MT. We also pointed out further directions like personalized MT and multi-modal translation. We hope this work can inspire further researches in this area of using LLMs for MT.