Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting
Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, Furu Wei
Introduction
Large language models (LLMs) demonstrate impressive multilingual capability in a wide range of natural language processing tasks, including language generation, knowledge utilization, and complex reasoning (Zhao et al., 2023). Their performance in downstream tasks has been shown to reach or even surpass human-level performance (Brown et al., 2020; Chowdhery et al., 2022; Scao et al., 2022). The capabilities of LLMs stem from the extensive volume of training data they leveraged (Kaplan et al., 2020). The training data for current models is primarily dominated by the English language corpus, but it also encompasses data from other languages, as described in GPT-3 (Brown et al., 2020), PaLM (Chowdhery et al., 2022), and BLOOM (Scao et al., 2022), etc.
There are over 7,000 languages worldwide, with the vast majority being low-resource or extremely low-resource languages (Forkel et al., 2022). Despite the latest GPT-4 model (OpenAI, 2023) demonstrating some generalization capabilities in multilingual tasks as evaluated on the MMLU benchmark (Hendrycks et al., 2021), it is still the case that LLMs do not have equal capability to handle all languages, leading to imbalanced capability across different languages. Furthermore, several evaluation results (Bang et al., 2023; Jiao et al., 2023; Hendy et al., 2023; Zhu et al., 2023) indicate that large models struggle with understanding and generating non-English languages, particularly in low-resource or extremely low-resource languages. Therefore, to democratize language intelligence and minimize performance gaps in different language, it is essential and meaningful to stimulate and enhance the multilingual capability of models in non-English and low-resource languages.
Intuitively, LLMs can improve multilingual capability by augmenting data (Lin et al., 2022) or fine-tuning models (Chen et al., 2021, 2022), but both are computationally expensive. Alternatively, in-context learning with prompts can also boost performance (Brown et al., 2020; Ahuja et al., 2023; Wei et al., 2022c) but is limited to monolingual tasks (Sanh et al., 2022).
This work explores a universal in-context learning approach to enhance the multilingual capability of LLMs. We introduce a simple yet effective method, called cross-lingual-thought prompting (XLT), to enable models to handle various natural language processing tasks across different target languages. Our method employs a generic and language-independent prompt, which eliminates the need to update model parameters. Depending on the task input type, cross-lingual-thought prompting guides the large language model to assume the role of an expert in a specific language for a particular task. Given its predefined meta information, XLT directs LLMs to respond logically through a process involving problem understanding, cross-lingual thinking, task analysis, task execution, and output formatting. During this process, our method is designed to stimulate models’ cross-lingual and logical reasoning skills, enabling them to respond to input requests regardless of the language. For enhanced performance, few-shot learning can also be employed with our method by providing an LLM-generated response output as a demonstration using cross-lingual-thought prompting zero-shot learning.
We conduct a comprehensive evaluation to verify the effectiveness of XLT across seven representative multilingual benchmarks of natural language reasoning, understanding, and generation tasks. Each benchmark includes multilingual data covering both high-resource and low-resource languages. The experimental results demonstrate that our method can significantly improve the performance of all benchmarks across languages under both zero-shot and few-shot learning settings. Notably, XLT achieves an average gain of over 10 points on the MGSM and MKQA benchmarks. Furthermore, we observe that our prompting method significantly reduces the gap between the average performance and the best performance of each task in different languages, indicating its potential to democratize language intelligence.
Cross-Lingual-Thought Prompting
Although LLMs are capable of accepting any input and generating responses, users typically structure their requests in the form of prompts to elicit the desired output. The design of these prompts is crucial for achieving optimal performance on downstream tasks, as LLMs are sensitive to the format of the prompts chosen (Zhao et al., 2021). Through a process called instruction tuning (Wei et al., 2022a), models can develop the ability to follow natural language instructions (Wei et al., 2022b), which can reduce their sensitivity to prompt engineering (Wei et al., 2022a). In accordance with the guidelines of the OpenAI cookbookhttps://github.com/openai/openai-cookbook, we propose a cross-lingual thought prompting template, denoted as the XLT template. This generic template allows LLMs to respond to requests with cross-lingual thought and supports a wide range of multilingual tasks.
Figure 2 displays the XLT template, with the colored sections representing placeholders. Figure 1 showcases an example of instantiated prompt for the Chinese request. The following section will explain the details of constructing XLT.
The XLT template is designed to emulate the process humans employ when handling multilingual tasks. Our template is written in English, as English is the dominant language during LLM pre-training, and existing research indicates that English prompting is more effective for multilingual tasks (Shi et al., 2023). In contrast to the vanilla prompt that only includes a task description, our XLT template aims to elicit multilingual capability through cross-lingual thoughts. This template comprises six logical instructions in sequence. To complete the template, only seven placeholders need to be filled in based on intrinsic knowledge of the task and the request, as depicted in igure 2.
. First, the model receives a role definition that helps establish the model’s behavior. This concept is akin to the system role of ChatGPThttps://platform.openai.com/docs/guides/chat/introduction. To achieve this, we simply need to fulfill the task name with a known category (such as commonsense reasoning or paraphrase identification), along with the language of the task in the task language field.
. Second, we explicitly append the request as the task input. The request is basically structured in terms of the task type so as to make sure the model can comprehend it. For example, in the natural language inference task, the two sentence inputs are specified with “premise” and “hypothesis”, respectively.
. We encourage the model to engage in cross-lingual thought by rephrasing the requested content in English, which is the dominant language used as a pivot language by Shi et al. (2023) and Ahuja et al. (2023). Rephrasing the requested content enclosed in the input tag helps the model better understand the request in its native language and knowledge. Our observations suggest that using keywords such as "retell" or "repeat" while rephrasing the content may result in better performance in practice.
. After rephrasing the task input, we need to complete the task in task goal. This step is comparable to the task description used in conventional prompting methods. In practice, we can get the task information from the literature or seek assistance from ChatGPT to generate effective prompts for solving the task (Jiao et al., 2023).
. We then ask the model to follow the instructions and complete the task step by step. Since LLMs exhibit a strong ability to maintain a chain-of-thought (Wei et al., 2022c), we carefully design instructions to guide the model, with the hope that it will respond to our instructions in a step-by-step manner and utilize the intermediate outputs to aid in solving the task.
. Finally, we should regularize the output format of the model to obtain the exact answer. LLMs are utilized in a zero- or few-shot manner, and they tend to generate texts that may not conform to the format of the target answer. Fortunately, LLMs possess a strong ability to follow instructions, and we can define the output format in terms of output type and output constraint. The output type can be a number, index, or text, while the output constraint is optional and determined based on the task requirements. Output constraint may include length limitations, language specifications, and other relevant factors.
2 XLT for Few-shot Learning
The above construction of XLT can be directly fed to LLMs to yield outputs, which is performed in the zero-shot learning setting. In addition, we also explore incorporating demonstrations into XLT to enable few-shot learning. Different from previous work that just appends model outputs to the corresponding request (Shi et al., 2023) or utilizes a verbalizer to format the output, our method constructs the demonstrations with better formatted model outputs from a step-by-step processing-based XLT. As illustrated in Figure 3, we first sample a few examples from the development set and incorporate the requested parts into XLT. The zero-shot learning is performed over LLM to collect responses that are further aligned with those of the samples. Only response-aligned requests are assembled with the corresponding model responses to form final demonstrations for few-shot learning. In this way, the demonstrations are constructed with rich logical knowledge via XLT, which will cater to the XLT-based generation of new requests. In practice, we can also correct or design the demonstrations for better alignment with the instruction logic.
Experiments
To comprehensively verify the effectiveness of our method on language-independent generality, we evaluate our XLT template on different LLMs covering various natural language processing tasks in multiple languages.
We conduct evaluations on seven typical benchmarks related to reasoning, understanding, and generation tasks that can represent different capabilities of LLMs, encompassing both high-resource and low-resource languages. These benchmarks cover 27 different languages, including English (en), German (de), Russian (ru), French (fr), Chinese Simplified (zh), Spanish (es), Japanese (ja), Italian (it), Vietnamese (vi), Turkish (tr), Indonesian (id), Swahili (sw), Arabic (ar), Korean (ko), Greek (el), Thai (th), Bulgarian (bg), Hindi (hi), Estonian (et), Bengali (bn), Tamil (ta), Galician (gl), Urdu (ur), Telugu (te), Javanese (jv), Haitian Creole (ht), and Southern Quechua (qu). In terms of the language distribution statistics in the Common Crawl Monthly Archiveshttps://commoncrawl.github.io/cc-crawl-statistics/plots/languages and the language performance of LLMs (Shi et al., 2023; Ahuja et al., 2023), we have arranged them in the order of language frequency from high-resource to low-resource. In particular, the frequency of some underrepresented languages is even less than 0.1% (e.g., bn, ta, gl, ur, te, jv, ht, qu).
Arithmetic Reasoning. The MGSM (Shi et al., 2023) benchmark contains grade school mathematical problems and asks the model to calculate the correct answer. It covers 11 languages, and we utilize the accuracy score for evaluation.
Commonsense Reasoning. The XCOPA (Ponti et al., 2020) benchmark contains one premise and two choices. It asks the model to choose which one is the result or cause of the premise. It covers 11 languages from 11 diverse families, and we utilize the accuracy score for evaluation.
Natural Language Inference. The XNLI (Conneau et al., 2018) benchmark contains one premise and one hypothesis and requires the model to determine whether the hypothesis is entailed, contradicted, or neutral conditioned on the premise. It covers 15 languages, and we utilize the accuracy score for evaluation.
Paraphrase Identification. The PAWS-X (Yang et al., 2019) benchmark contains two sentences and requires the model to judge whether they paraphrase each other or not. It covers 7 languages, and we utilize the accuracy score for evaluation.
Question Answering. The MKQA (Longpre et al., 2021) benchmark contains an open-domain question and asks the model to predict a short answer. Since it has unanswerable questions or long questions that do not have precise answers, we remove these questions during evaluation. It covers 25 languages, and we choose a subset of 10 languages, including de, en, es, fr, ja, ru, th, tr, vi, and zh. We utilize the token overlap F1 score for evaluation.
Summarization. The XL-Sum* (Hasan et al., 2021) (250 test samples randomly sampled from XL-Sum per language) benchmark contains a long news article and wants the model to summarize it into a short text. It covers 44 languages, and we choose a subset of 6 languages, including en, es, fr, tr, vi, and zh. We utilize the ROUGE-1 score (Lin, 2004) for evaluation.
Machine Translation. The FLORES* (Costa-jussà et al., 2022) (200 test samples randomly sampled from FLORES-200 per language) benchmark contains parallel text from Wikimedia projects for 204 languages, yielding over 40,000 translation directions. We choose a subset of 12 directions, including high resource to high resource translation (i.e., zh ru and devi), high resource to low resource translation (i.e., zh th and zhjv), and low resource to low resource translation (i.e., thgl and jvth). We utilize the SacreBLEU score (Papineni et al., 2002; Post, 2018) for evaluation.
Among these benchmarks, MGSM, XCOPA, XNLI, PAWS-X, and MKQA are parallel, i.e., the instances are semantics-equivalent across each language. For all benchmarks, we report the results on the test sets using all instances (Table 5), except for XL-Sum and FLORES-200, where we only sample 250 and 200 examples respectively to show the trend of generation performance. In the few-shot setting, we randomly choose examples from the development set if they have, otherwise, we translate the English training set into corresponding languages to construct several examples.
1.2 Baselines
are the vanilla in our experiments that were proposed and suggested in previous work. After determining the prompt, we format each monolingual instance using the English basic prompt. This setting is similar to the monolingual prompting in MEGA (Ahuja et al., 2023). The basic prompts used for the evaluation of each benchmark are listed in Table 5. Note that, we dismiss the baseline using native-language, since MEGA (Ahuja et al., 2023) reveals monolingual prompting is superior to cross-lingual prompting.
(CoT) prompting invokes LLMs to generate a series of intermediate results to solve reasoning tasks Wei et al. (2022c), which is still effective under multilingual scenarios Shi et al. (2023). In experiments, we append the instruction “Let’s think step-by-step and tell me the answer in the end” after the input to prompt LLMs.
leverages the robust capabilities of LLMs in English to tackle multilingual tasks, as suggested by both Shi et al. (2023) and Ahuja et al. (2023). This approach translates instances from other languages into English beforehand. In practice, we utilize the Google Translate API to translate examples into English and apply the basic prompt to format them. Note that, we do not apply this method to generation tasks since they require the output in respective language rather English.
utilizes the proposed template consisting of multiple instructions introduced in Section 2. The instantiated XLT templates for each benchmark are listed in Table 6.
In few-shot learning scenarios, for basic prompt, we use the same template as an additional input to the model. For XLT, we provide the exemplars with XLT template inputs and anticipate desirable step-by-step outputs as outlined in Figure 3. In the subsequent evaluation, we apply the 5-shot setting, except for the XL-Sum* experiments, which use the 3-shot setting due to input length constraints.
1.3 LLMs
We mainly evaluate two LLMs from the GPT-3.5 series models:
text-davinci-003https://platform.openai.com/docs/models/gpt-3-5 is trained using instruction tuning and reinforcement learning from human feedback (Ouyang et al., 2022). It can perform a wide range of natural language tasks with satisfactory results.
gpt-3.5-turbo4 is optimized for chat based on text-davinci-003 and suitable for traditional NLP tasks. It is the most capable GPT-3.5 model.
To verify the compatibility of our XLT template, we further incorporate LLaMA-2-Chat Touvron et al. (2023) (Llama-2-70b-chat-hf) as our base models. It is an open-source model that has been trained through supervised fine-tuning and reinforcement learning from human feedback on the base LLaMA 2 model. In addition, we also refer to the existing results from other LLMs, such as code-davinci-0024, when the evaluation is comparable. During inference, we employ greedy search (i.e., temperature=0) to generate the LLM responses. We find LLMs have excellent instruction-following abilities to respond to our instructions in the given format. Therefore, we just extract the part after “Answer format:” as labels.
2 Experimental Results
We comprehensively evaluate XLT’s performance over seven tasks. The average score of text-davinci-003 is summarized in Figure LABEL:fig:radar(a) and Table 1, and more details are listed in Appendix A. As for the CoT prompting, it can enhance reasoning tasks while becomes less effective on understanding and generation tasks. In terms of the Translate-En prompting, it can boost the performance in the zero-shot settings while may not work well in the few-shot settings. Overall, compared to the three baseline methods, XLT achieves significant improvements over two LLMs for all tasks on both zero-shot and few-shot settings regardless of the language difference, except for a slight drop on the PAWS-X benchmark in the zero-shot setting. It is noted that XLT achieves remarkable gains of nearly 20 points on average in the MGSM benchmark for the arithmetic reasoning task and around 10 points on average in the MKQA benchmark for the open-domain question answering task. The experiments demonstrates the effectiveness of XLT for empowering LLM with multilingual capability.
As for the compatibility test, we list the results of LLaMA-2-Chat on the MGSM benchmark in Table 7. It is notable that LLaMA 2 can also benefit from our cross-lingual-thought, which further demonstrates the generality of our XLT template. However, the gains of LLaMA-2-Chat is not as good as GPT-based models. Our analysis reveals this gap can primarily be attributed to LLaMA 2’s poorer multi-step instruction-following ability.
Furthermore, we try to assess the democratization degree of tasks between languages by defining a “democratization score”, which calculates the average percentage of performance attained by different languages relative to the best performance among all languages. Given the evaluation scores of corresponding to language on a task, the democratization score is formulated as:
Table 2 presents the degree of democratization for tasks across languages under both zero-shot learning and few-shot learning, and we further summarize it in Figure LABEL:fig:radar(b) by averaging all scores per task regardless of the setting and model differences. We can observe that XLT leads to higher democratization scores in general, particularly for XCOPA, and MKQA. As for MGSM, XNLI, and PAWS-X, our XLT can improve performance in multiple languages, where the overall performance of the baseline is consistently lower but the gap between languages is smaller as shown in Tables 7, 9, and 10. In conclusion, our method can reduce the performance gap between languages and improve the language democratization of LLMs.
3 Further Analysis
In this section, we further investigate the factors that affect the performance of XLT and how they affect various multilingual benchmarks.
For the XLT variants, we mainly conduct experiments to compare the following strategies:
Ablating the instructions. Since our XLT consists of six logical instructions, we disable the Role Assigning, Cross-lingual Thinking, and CoT Task Solving instructions separately to analyze the contribution per instruction.
Reordering the instructions. Considering the logicality of our instructions, we further change the order of the instructions in XLT to explore whether LLMs will handle tasks differently and lead to different results.
Changing the content word. As prompts are usually sensitive to the word choice, we verify the robustness of XLT when alternating the rephrasing keyword with “retell”, “repeat”, and “translate” in the cross-lingual thinking instruction.
The outcomes are presented in Table 3, indicating that XLT surpasses almost all the variants, thereby validating the effectiveness and reasonableness of our proposed XLT method.
The results from the “Instruction Ablation" row indicate that: (1) Cross-lingual Thinking yields more significant gains compared to other instructions. This suggests that the LLM’s ability of cross-lingual thinking is activated, allowing it to utilize its knowledge in English to solve tasks effectively; (2) Removing Role Assigning from XLT impedes the model’s understanding of the ultimate goal for diverse multilingual tasks, highlighting the task transferability of XLT; and (3) the better performance of XLT can also be attributed to CoT Task Solving, which requires the model to respond to complex instructions in a step-by-step manner.
The performance drop is evident when the order of our designed logical instructions is switched. When designing XLT, we have taken into account the process by which humans solve multilingual problems, and this experiment further confirms the optimum order of our XLT template. Placing the Role Assigning instruction later may confuse the model initially. Additionally, conducting Cross-lingual Thinking before Task Analyzing is crucial since we rely on the English task-solving abilities of LLMs to handle multilingual tasks.
We can find that different words indeed affect the performance of XLT, but it is less sensitive to the other variants. Through experimentation, we have determined that “repeat" yields better results for text summarization and machine translation, while “retell" is more suitable for the remaining five tasks. Our aim is to provide XLT with a more unified template, while still allowing users to fine-tune specific keywords for optimal performance in their tasks.
3.2 Effectiveness of XLT Few-shot Learning
As mentioned in Section 2.2, the construction of demonstrations for XLT few-shot learning differs from the previous method. We have compared XLT and basic prompt. Here, we focus on the construction of the demonstration input-output pairs and compare various demonstrations that may be used to perform XLT few-shot learning. The illustrations can be found in Figure 4.
Basic prompt input + Basic prompt output: This is the normal demonstration format used in most of the previous work.
Basic prompt input + XLT output: This ablation is to separate the effect of input and output formats in the demonstration.
XLT input + XLT output: This is the method that we used in this work.
Observing the experimental results presented in Table 4, we can conclude that: (1) Our XLT few-shot learning outperforms all other variants, thus confirming its effectiveness. (2) The use of normal demonstrations for XLT few-shot learning leads to a decrease in performance. (3) Merely incorporating XLT as a demonstration input without its output does not result in any improvements. (4) Consistency in the demonstration for few-shot learning is crucial, implying that the demonstration input-output format should align better with its zero-shot learning input-output format.
Related Work
Despite the impressive capabilities of LLMs, it is crucial to determine their impact on natural language processing tasks. Liang et al. (2022) conduct a comprehensive evaluation of LLMs from various perspectives, such as accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency. Bang et al. (2023) extensively evaluate the ChatGPT model on multiple natural language processing tasks and find that the model performs well in high-resource languages but exhibits certain limitations in low-resource and non-Latin script languages. Additionally, studies by Jiao et al. (2023) and Hendy et al. (2023) compare different GPT models with supervised models for machine translation tasks and find that GPT models have competitive translation abilities in high-resource languages but perform less effectively in low-resource languages. It is worth noting that achieving multilingual generative AI capability necessitates cross-lingual knowledge to further improve the model’s performance. In this context, Ahuja et al. (2023) evaluate the multilingual task understanding ability of GPT models and attempt to enhance their task processing abilities in other languages using English knowledge. Our work also focuses on evaluating the multilingual capabilities of LLMs, including reasoning, understanding, and generative capabilities. Our evaluations indicate that LLMs exhibit differences in high-resource and low-resource abilities, which necessitates additional efforts to enhance their multilingual capability.
2 Multilingual Task Processing
Multilingual knowledge has been shown to be exploitable and transferable between languages to improve model performance (Devlin et al., 2019; Conneau et al., 2020; Raffel et al., 2020; Ouyang et al., 2021; Chi et al., 2021). While much research has been devoted to multilingual understanding tasks, multilingual generation tasks are more challenging, particularly when the target language is low-resource or non-English (Ma et al., 2021; Liu et al., 2020). Two methods can enable models to support multilingual task processing: one is training a supervised model that covers multiple languages for multilingual processing (Costa-jussà et al., 2022), and the other is training a pre-trained model and using fine-tuning to transfer knowledge among languages to achieve multilingual capability (Chen et al., 2021, 2022). However, the emergence of LLMs has made it possible to directly process multilingual tasks via in-context learning (Brown et al., 2020; Ahuja et al., 2023). These LLMs, with hundreds of billions or even trillions of parameters, require a significant amount of computation resources for training, making traditional fine-tuning methods less feasible. To improve the generative ability of LLMs, researchers explore in-context learning methods that do not require updating model parameters, such as few-shot prompting (Vilar et al., 2022), automatic prompt learning (Shin et al., 2020), task-instruction prompting (Ye et al., 2023), chain-of-thought prompting (Wei et al., 2022c), etc. Our work builds upon these methods and proposes an optimized, generic, and language-independent prompt to enhance the multilingual capability of LLMs.
Conclusion
This work investigates the language processing capabilities of large language models in multilingual settings and expects to develop a universal framework for handling diverse multilingual tasks. To accomplish this goal, we propose a generic prompt, referred to as XLT, to enhance the multilingual capability and reduce the performance gaps among languages in tasks related to language understanding, reasoning, and generation in non-English and low-resource languages. Although our method is generally applicable across tasks and languages, we discovered that prompting design factors such as instruction logic and word choice have explicit impacts on its effectiveness. Cross-language thinking in XLT is particularly effective. Finally, we hope this work can inspire further research to prioritize the development of generic prompting. By doing so, large language models can encompass a wider range of modalities and languages.
Acknowledgements
Tianyi Tang and Xin Zhao are supported by National Natural Science Foundation of China under Grant No. 62222215, Beijing Natural Science Foundation under Grant No. 4222027 and L233008.
Limitations
Due to limitations imposed by the evaluation benchmarks and OpenAI API cost, we conducted tests on 27 languages, which merely scratch the surface of the vast array of languages in the world. Besides, our XLT template is based on English. It deserves to explore whether the template written in task language can lead to better performance and how to better construct the instruction in each language. Furthermore, we only verify the effectiveness of our method on two GPT-based models (i.e., text-davinci-003 and gpt-3.5-turbo) and LLaMA-2-Chat. It is worthwhile to investigate the generality of our template on more models, such as BLOOM and PaLM.
References
Appendix A Additional Experiments
Table 7 presents the results of the MGSM benchmark. XLT significantly improves the arithmetic reasoning capabilities of both models, particularly for gpt-3.5-turbo in the zero-shot setting. We hypothesize that gpt-3.5-turbo may have undergone supervised fine-tuning (Ouyang et al., 2022) with arithmetic reasoning samples in the chain-of-thought format, which enables XLT to activate its arithmetic reasoning ability directly. For both low-resource languages (e.g., sw, th, bn, and te) and high-resource languages, XLT can further enhance the performance. Even under the few-shot setting, XLT can still significantly improve the reasoning performance of both models and reduce the performance gap for all languages. Notably, for some high-resource languages, such as de, ru, fr, and es, the performance is comparable to English.
The XCOPA benchmark results are presented in Table 8. Our XLT approach significantly enhances the performance of both models in both settings, as compared to basic prompting. In the zero-shot setting, XLT demonstrates significant improvements for relatively low-resource languages (e.g., sw, th, et, ta, and ht), but it underperforms the baseline for some high-resource languages such as zh and it. In the few-shot setting, XLT brings enhancements for both high- and low-resource languages. Our findings suggest that XLT is more effective for low-resource languages, particularly for gpt-3.5-turbo on sw, th, ta, and ht, where it yields improvements of over 10 accuracy points.
A.2 Results on Understanding Tasks
Table 9 presents the results of the XNLI benchmark. In the zero-shot setting, our XLT significantly outperforms the basic prompt in all languages. Additionally, when using few-shot setups on high- and low-resource languages, both text-davinci-003 and gpt-3.5-turbo show significant improvements compared to the basic prompt. Specifically, for low-resource languages such as th, bg, hi, and ur, XLT achieves an average improvement of 9.4 accuracy scores for text-davinci-003 and 5.3 accuracy scores for gpt-3.5-turbo. This demonstrates that XLT is effective for both models, but text-davinci-003 has better natural language inference capabilities.
Table 10 displays the comparisons on the PAWS-X task, where XLT outperforms basic prompt in all languages, particularly for low-resource languages under the few-shot setting. We observe a slight performance drop on average in zero-shot learning compared to gpt-3.5-turbo for some high-resource languages (e.g., en, de, and fr). Based on our analysis of intermediate outputs, we infer that the drop in performance may be due to cross-lingual thinking that alters the original meaning of the two sentences, leading to difficulties in judgment. Additionally, a comparable pattern is evident in a previous study (Ahuja et al., 2023), where non-Latin script languages (ja, zh, and ko) exhibit significantly poorer performance than English or German in the few-shot setting. Nevertheless, by demonstrating the construction of XLT, we can guide the model on how to think across different languages and effectively address the aforementioned issues.
A.3 Results on Generation Tasks
The MKQA benchmark outcomes are listed in Table 11. Across all languages in the zero-shot and few-shot settings, the XLT template shows a significant improvement over the basic prompt. It is worth noting that text-davinci-003 performs worse than gpt-3.5-turbo in this task, and we speculate that the latter is optimized for open question answering, which is common in daily chat. Additionally, our findings indicate that XLT can notably enhance the performance of under-resourced languages. XLT brings over 10 points of improvement for these languages. (e.g., zh, ja, vi, and tr) This aligns with previous benchmarking studies and is particularly noteworthy in this evaluation. We suspect that high-resource and low-resource languages share the same cross-lingual thinking as English to greatly leverage the LLM’s ability to solve English open-domain QA.
The results of the XL-Sum* benchmark are presented in Table 12. It can be observed that XLT outperforms the basic prompt in both zero- and few-shot settings across all languages. Additionally, the LLM model exhibits a significant improvement in generating summaries under the few-shot setting compared to the zero-shot setting. This suggests that providing fewer examples can effectively guide the model in summarizing multilingual texts. Furthermore, the few-shot results revealed an interesting finding that text-davinci-003 performed better when gpt-3.5-turbo and text-davinci-003 use basic prompt. However, once XLT is enabled, gpt-3.5-turbo outperforms text-davinci-003, highlighting the effectiveness of our approach.
Machine translation is a special generation task where the source and target are two different languages. The experiment in this part is to verify how XLT boosts machine translation tasks. Since English has been specified as the pivot language in the cross-lingual thinking in XLT, we exclude English-centric tasks to avoid language redundancy and focus on 12 non-English translation directions in the FLORES* benchmark, which includes both high-resource and low-resource languages. As shown in Table 13, XLT achieves impressive zero-shot results for all languages compared with basic prompt. For example, it significantly improves translation quality in Chinese-to-X or X-to-Chinese. The result emphasizes that XLT will potentially transfer the knowledge of a high-resource pivot language like English to the target language. While the benefit of XLT may not be as obvious for high-to-high translations, it becomes more significant for high-to-low, low-to-high, and low-to-low translations. For instance, XLT improves the translation performance of gpt-3.5-turbo by nearly 4.0, 2.8, and 3.3 BLEU points for thgl, jvzh, and zhth translations, respectively, demonstrating its effectiveness regardless of whether the source language is high-resource or low-resource. Noticing that Hendy et al. (2023) have shown that few-shot configurations do not yield significant improvements over the zero-shot setup for translation tasks, we do not evaluate the few-shot paradigm on FLORES* in this work and leave it for future exploration.