A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, Pascale Fung

Introduction

ChatGPT is a successor of the large language model (LLM) InstructGPT Ouyang et al. (2022) with a dialog interface that is fine-tuned using the Reinforcement Learning with Human Feedback (RLHF) Christiano et al. (2017) approach. ChatGPT has gathered 100 million monthly active users in such a short period of time Hu (2023) and is being used by businesses and consumers alike for a myriad of mostly textual tasks. One reason for its unprecedented popularity is that ChatGPT, through its scale and via RLHF, has shown impressive abilities in many areas of NLP as well as emergent abilities. Another reason is that its dialog interface allows users to interact with the underlying LLM more effectively and efficiently via interactive chats that are akin to multi-turn prompting.

However, despite its powerful abilities, anecdotal reports on ChatGPT consistently showed remaining challenges - for example, it fails in some elementary mathematical Gilson et al. (2022); Goldberg (2023); Frieder et al. (2023); Choi et al. (2023); Davis (2023b) and commonsense reasoning tasks Guo et al. (2023); Davis (2023b); it hallucinates with human-like fluency and eloquence on things that are not based on truth Shen et al. (2023); Thorp (2023); Smith (2023); and as a general-purpose language model trained from everything on the web, its language coverage is questionable Lu et al. (2022); Jiao et al. (2023). Consequently, it is not clear what people can or cannot use ChatGPT for despite its popularity.

Since OpenAI never published any benchmarking results on ChatGPT at the time, seeing this need, in February 2023, we proposed a comprehensive framework for quantitatively evaluating interactive LLMs such as ChatGPT through standard public test sets on major NLP tasks such as question answering, reasoning, summarization, machine translation, sentiment analysis, language identification, task-oriented dialogue, and misinformation detection. We evaluate its multilingual performance as well as vision-language multimodal abilities. With additional experiments, we also quantitatively evaluated its primary limitations in reasoning and hallucination. In addition, we conducted experiments to test its multi-turn interactivity as a means for better prompt engineering. We aimed to provide insights to users of ChatGPT on the strengths mentioned above and limitations, as well as how they can improve outcomes with interactivity. To the best of our knowledge, this is the first published benchmark of ChatGPT from a third party. More recently, the GPT-4 technical report OpenAI (2023) published a number of human task benchmarks.

The true scope of all emergent capabilities of generative models, including ChatGPT, is still unclear. Thus, any benchmarking exercise cannot be 100% “comprehensive” in the scientific sense. We aim to show not just researchers but also users what ChatGPT can and cannot do by presenting interpretable benchmarking results in a zero-shot setting without access to APIs so that the general audience can replicate our evaluation with the test sets we have provided in a zero-shot setting. This version of ChatGPT is 15 December 2022.

The following are the major insights we have gained from the evaluations:

Multitask, Multimodal, and Multilingual: For 9/13 NLP datasets, ChatGPT outperforms previous LLMs with zero-shot learning. It even outperforms fully fine-tuned task-specific LMs on 4 different tasks. In other cases, ChatGPT is on par or slightly lower than fully fine-tuned for specific NLP tasks; ChatGPT fails to generalize to low-resource and extremely low-resource languages (e.g., Marathi, Sundanese, and Buginese). There is an overall performance degradation in low-resource languages, especially in non-Latin scripts in the case of translation; its weakness lies in generation rather than understanding part of the translation process; ChatGPT enables a code intermediate medium to bridge vision and language, even though the multi-modality ability is still elementary compared to vision-language models.

Reasoning: We tested 10 different reasoning categories with 634 samples in total. Based on our experiments, ChatGPT shows more weakness in inductive reasoning than in deductive or abductive reasoning. ChatGPT also lacks spatial and mathematical reasoning while showing better temporal reasoning. Further, we found that ChatGPT is relatively better at commonsense reasoning than non-textual semantic reasoning. Finally, while ChatGPT shows acceptable performance in causal and analogical reasoning, it is bad at multi-hop reasoning capability, similar to other LLMs’ weakness (Ott et al., 2023).

Hallucination: Similar to other LLMs (Radford et al., 2019; Muennighoff et al., 2022; Workshop et al., 2022), ChatGPT suffers from the hallucination problem. It generates more extrinsic hallucinations – factual statements that cannot be verified from the source.

Interactivity: One of the primary differentiating factors of ChatGPT from its predecessors is its multi-turn dialog interactivity. This enables ChatGPT to perform multiple tasks within a dialog session. There is also significant performance improvement (8% ROUGE-1 on summarization and 2% ChrF++ on low-resource machine translation) via multi-turn interactivity in various standard NLP tasks. This process is akin to prompt engineering with feedback from the system.

Multitask, Multilingual, and Multimodal Evaluations of ChatGPT

ChatGPT has become very well-known in such a short period of time to general public users, not just those who are in AI, machine learning, and NLP communities who might be more familiar with LLMs. One of the main reasons is that, in addition to media reports, innumerable use cases of ChatGPT are shared by both non-academic and academic users online (Marr, 2022; Gordon, 2023; Shankland, 2023). There have been debates and panels on whether ChatGPT is approaching Artificial General Intelligence, as it seems to be able to carry out a multitude of tasks without specific fine-tuning (Desk, 2023b; Johnson, 2023; Kingson, 2023). On the other hand, there has also been as much sharing of its failures in simple tasks (Gilson et al., 2022; Choi et al., 2023; Shen et al., 2023).

Instead of relying on anecdotal examples, we first evaluate ChatGPT’s performance in various standard NLP tasks in a zero-shot manner to obtain a basic/better understanding of its multi-task ability. We compile results from the existing literature on ChatGPT and compare them with the state-of-the-art fully-fine-tuned and zero-shot models across multiple tasks. We evaluate ChatGPT performances on 21 datasets covering 8 tasks, i.e., summarization, machine translation, sentiment analysis, question answering, task-oriented dialogue, open-domain knowledge-grounded dialogue, and misinformation detection tasks. We sample testing cases from existing standard test sets for each task with a sample size ranging from 30 to 200 samples.

The result of the multitask evaluation is shown in Table 1. ChatGPT is shown to achieve remarkable zero-shot performances on multiple tasks, surpassing previous state-of-the-art zero-shot models on 9 out of 13 evaluation datasets with reported zero-shot LLMs’ performances. In most tasks, especially task-oriented and knowledge-grounded dialogue tasks, task-specific fully-fine-tuned models outperform ChatGPT. Compared to the latter, ChatGPT yields lower performance in most tasks while still surpassing the performance on 4 datasets.

Furthermore, from the evaluation results, we also observe several limitations of ChatGPT: 1) limited language understanding and generation capabilities on low-resource languages, 2) lacking reasoning ability as shown from the results in QA, and 3) performing task-oriented and knowledge-grounded dialogue tasks. More detailed experimental setup and analysis for each task are shared in Appendix §C. We also provide the complete list of all the datasets used in our evaluation in Appendix J.

Given that ChatGPT has the ability to generate conversation-like responses, we test it on conventional dialogue tasks: 1) knowledge-grounded open-domain dialogue and 2) Task-oriented dialogue. Task setups are explained in Section C.6.

To quantitatively measure ChatGPT’s performance on knowledge-grounded dialogue, we utilize 50 samples from the test set of OpenDialKG (Moon et al., 2019), which contains open-ended dialogues grounded on a knowledge path. According to human judgment, the responses from ChatGPT are of high quality with fluent response generation and incorporating the provided knowledge in the response. However, the automatic evaluation results are relatively low compared with fine-tuned GPT2. We postulate this is because ChatGPT responses are longer than the golden answers and include content from its parametrized knowledge injected during pre-training.

We investigate and discuss how ChatGPT’s emergent abilities and interactivity could potentially be leveraged for ToD as well in two setups. Firstly, A) modular approach: testing dialogue state tracking (DST) and response generation using oracle actions. DST is mediocre while ChatGPT successfully leverages all information provided while answering the questions with a 71.1% inform rate and 5.65 BLEU score. Next, B) Unified approach: a direct approach to simulate the ToD interaction while leveraging information in a structured database. We observed the limitations of ChatGPT: 1) ChatGPT cannot keep the belief state across multiple turns within the interaction, 2) ChatGPT’s response tends to be wrong if the query introduces a basic level of reasoning 3) ChatGPT tends to generate hallucinated information beyond the given knowledge, which is not desirable for ToD. We provide details and examples of the modular and unified approaches in Appendix C.6.2.

2 Evaluating Multilinguality of ChatGPT

Training data size affects language understanding and generation ability of LMs (Raffel et al., 2022; Cahyawijaya et al., 2021; Rae et al., 2021; Workshop et al., 2022; Chowdhery et al., 2022; Hoffmann et al., 2022). As an LLM, the same premise also applies to ChatGPT, but the question is to what extent. We investigate this question through a series of experiments by analyzing 1) the language understanding capability through sentiment analysis (SA) and language identification (LID) tasks, and 2) the language generation capability through machine translation using English as the pivot language. Based on the size proportion in CommonCrawl (i.e., the primary source of language pre-training data used in various LLMs) , we group languages into 4 language resource categories, i.e., high-resource language (HRL) (¡≥\geq1%), medium-resource language (MRL) (≥\geq0.01%), low-resource language (LRL) (≥\geq0.0001%), and extremely low-resource language (X-LRL) (¡0.0001%). The statistics of the languages are shown in Table 9 and other details are described in Appendix D.

We investigate the language understanding ability of ChatGPT on 4 languages from different language categories in NusaX (Winata et al., 2022), i.e. English, Indonesian, Javanese, and Buginese, through sentiment analysis and language identification tasks. ChatGPT fails to generalize to extremely low-resource languages. As shown in Table 2, there is a clear correlation between ChatGPT performance with the language resource category. This result aligns with the findings from prior works (Chowdhery et al., 2022; Workshop et al., 2022; Muennighoff et al., 2022), where LLMs, including ChatGPT, yield a lower performance for lower resource languages. Interestingly, the performance gap between English, Indonesian, and Javanese is considered marginal compared to the performance gap with Buginese. This suggests that ChatGPT has a limitation in generalizing toward extremely low-resource languages. Furthermore, we also find that ChatGPT can understand low-resource languages, such as Javanese, without having the knowledge to identify the language itself. Moreover, ChatGPT displays better human-preferred responses when it has no knowledge about the language. For instance, as illustrated in 8, ChatGPT lets the user know that its prediction is uncertain when it does not completely understand the language and also provides broader information regarding the language.

2.2 Language Generation

We assess the multilingual language generation ability of ChatGPT through machine translation. We experiment with 6 languages: French, Chinese, Indonesian, Korean, Javanese, and Sundanese from the FLORES-200 dataset (Team et al., 2022; Goyal et al., 2022a). For each language, we sample 30 English-XXX parallel sentences and perform two directions of translation using English as the pivot language. The correctness of the translation results is manually validated by a native speaker of the corresponding language.

Based on our evaluation results (Table 3), similar to other LLMs (Workshop et al., 2022; Muennighoff et al., 2022), ChatGPT produces better English translation quality from high-resource languages, such as French and Chinese. While for low-resource languages, such as Javanese and Sundanese, ChatGPT tends to generate several mistranslated words/phrases and sometimes even hallucinate some objects. Moreover, we also observe that sometimes ChatGPT translates the English sentence into a different but related language other than the requested target language (see §5). This fact suggests that the generalization of LLMs, including ChatGPT, to low-resource languages, remains an open challenge. Moreover, we also find that ChatGPT can handle Latin script languages better than non-Latin script languages, especially in generating sentences using those scripts.

3 Evaluating Multimodality of ChatGPT

Since ChatGPT is a purely text-prompted language model, it is unlikely to explore its multimodal capabilities with visual inputs like contemporary vision-language works (Rombach et al., 2022; Ramesh et al., 2021; Yu et al., 2021a; Radford et al., 2021; Dai et al., 2022; Lovenia et al., 2022; Dai et al., 2023a). However, thanks to its code understanding and generation abilities, programming codes can serve as the intermediate medium to bridge vision and language (Rasheed, 2020; Shiryaev, 2022). Given textual prompts, ChatGPT can generate code representations of visual images using the SVG (Scalable Vector Graphics) format or APIs (e.g., HTML Canvas element, Python Turtle graphics). For example, as shown in Figure 1, ChatGPT can generate a well-formed and suitable intermediate representation in code format to synthesize images given the dialogue context and user prompts.

In this way, even though the generated images are symbolic and their quality is not comparable to the ones generated by modern text-to-image models (Ramesh et al., 2021; Rombach et al., 2022), it is worth exploring due to three reasons. Firstly, it helps us investigate the visual understanding and reasoning abilities of ChatGPT, which can be seen as an emergent skill after the very large-scale pre-training on text and code data. Furthermore, representing images with code is a more explainable way to understand the model’s behaviors and rationales in text-to-image generation. Third, it is a natural way to evaluate ChatGPT’s ability on multi-turn interaction by asking for post-editing and corrections of the generated images.

To systematically evaluate the image generation ability of ChatGPT through code generation, we designed a national flag drawing task. This task tests how ChatGPT’s textually described knowledge (language) converts into the drawing (vision) through the SVG (code), using multi-turn conversations. The task contains three steps. Firstly, we ask ChatGPT to illustrate the appearance of the flag. Next, based on the description, we ask ChatGPT to generate the SVG code of that flag. Finally, if the generated image contains errors, we iteratively ask ChatGPT to fix them. There are four types of errors: 1) layout, 2) color, 3) missing components, 4) shape/size. We uniformly collect 50 national flags from different continents and conduct the flag-drawing task on ChatGPT. The prompts and full results are shown in Appendix E. The generated flag images are evaluated by the aforementioned four error types as criteria. We further assess the image quality with five grades, A ∼\sim E, which indicate zero to four (or above) errors. An overview of the result evaluation is provided in Table 4.

We share our major two findings from the task: 1) ChatGPT is capable of drawing, yet better with a self-generated textual description. As demonstrated in Table 4 and Appendix E, by following the task formulation, ChatGPT can generate plausible national flags using the SVG format. To better understand the behavior of ChatGPT, we perform an ablation study by removing the description generation step. As illustrated by Figure 1, the performance drops dramatically without first prompting the textual flag description, which is generated by ChatGPT itself. Explicitly describing the appearance of the flag and then drawing disentangles the image generation process, which can be considered as a chain-of-thought reasoning. 2) ChatGPT is an elementary illustrator. Among the four error types, the majority lies in the shape/size error, which happens 68% of the time. For the other three error types (layout, color, missing components), they appear 34%, 20%, and 18% of the time, respectively. For instance, ChatGPT cannot generate the exact shape of the maple leaf in the Canadian flag while it gets the layout and color correctly (Figure 3). This is a natural defect of text-only language models as they never see actual visual data and textual data is usually conceptual.

Reasoning Evaluations of ChatGPT

Reasoning is one of the most actively discussed and debated abilities of LLMs as scaling the model parameter size also increases the implicit knowledge in LLMs Wei et al. (2022a); Wang et al. (2022); Huang and Chang (2022). Mahowald et al. eloquently argues that “language ability does not equal to thinking” or “reasoning” in LLMs, and that LLMs have poor reasoning skills despite possessing human-level language skills.

In the NLP literature, evaluating a model’s reasoning often means evaluating its various skills in arithmetic, commonsense, and symbolic reasoning in different NLP tasks that require such skills Talmor et al. (2020); Zelikman et al. (2022); Wei et al. (2022b). However, the reasoning itself is a much broader concept thus it is hard to conclude whether a model can “reason” or not based on those aforementioned, and current works on reasoning are scattered. This is in line with the anecdotal experience of users with ChatGPT – some of the examples demonstrate surprisingly good “reasoning” abilities compared to previously introduced LLMs but at the same time ChatGPT fails in very simple reasoning problems Desk (2023a); Venuto (2023); Qiao et al. (2022); Cookup.ai (2022); Labs (2022).

Thus, we investigate the reasoning ability of ChatGPT in a more fine-grained manner, which includes deductive, inductive, abductive, analogical, causal, multi-hop, mathematical, temporal, and spatial reasoning, via question-answering tasks. We categorize available QA tasks into each category by avoiding overlap (i.e., choosing testsets that require mainly one specific category of reasoning). Composed results and corresponding datasets for each category are shown in Table 5. For evaluation, we manually check the accuracy of the answer as well as verify the rationales and explanations generated by ChatGPT. A detailed explanation of task setup is explained in Appendix F.

Inductive, deductive, and abductive reasoning are common forms of logical reasoning, a process of deriving a conclusion or judgment based on given evidence or past experience and observations (Rogers et al., 2022; Wason and Johnson-Laird, 1972; Huang and Chang, 2022). We first investigate basic reasoning skills with bAbI tasks (Weston et al., 2016b), 30 examples each from task 15 (inductive) and task 16 (deductive). One major investigation is that ChatGPT is a lazy reasoner that suffers more from induction. Interestingly, when ChatGPT was asked to answer a question given premises without any prompt engineering, it performed poorly in induction (0 out of 30) while it achieved much better performance in deduction (19 out of 30). However, when ChatGPT is explicitly asked for reasonable inference inductive reasoning increases to 20 out of 30. Yet, it is still not as good as in deduction. When we repeat the analysis on advanced tasks, specifically on CLUTRR Sinha et al. (2019) for induction and EntailmentBank for deduction Dalvi et al. (2021), the same conclusion holds based on our experiment.

It is often investigated in public sharing about ChatGPT errors cases that it lacks the reasoning ability that requires non-text semantic understanding such as mathematical, temporal, and spatial reasoning. Not surprisingly, it could only score 23.33% (7/30) for the MATH dataset (Saxton et al., 2019), which tests mathematical reasoning. Overall, ChatGPT correctly answers 86.67% of the time (26/30), suggesting that it has a decent temporal reasoning ability. ChatGPT falls short of the spatial reasoning tasks, with success rates of 43.33% for StepGame and 43.75% for SpartQA. We investigate the errors that it often fails to understand clock direction (e.g., “W is at K’s 3 o’clock”) and diagonal spatial relations.

It is understanding and reasoning about everyday concepts and knowledge that most people are familiar with, to make judgments and predictions about new situations Storks et al. (2019). Recent works show that LLMs perform impressively well on commonsense reasoning benchmarks Qiao et al. (2022); Huang and Chang (2022); Bhargava and Ng (2022). Based on our evaluation with CommonsenseQA Talmor et al. (2018), PiQA Bisk et al. (2020) and Pep-3k Wang et al. (2018b), ChatGPT shows surprisingly good commonsense reasoning capability, perhaps due to its large parametric memory.

Factuality and Hallucination

LLMs are known to be susceptible to generating nonfactual, untruthful information, which is referred to as hallucination (Lee et al., 2022; Ji et al., 2022a, b; Su et al., 2022; Dai et al., 2023b; Xu et al., 2023). Many anecdotal witnesses show ChatGPT also seems to suffer from the same problem as other LLMs. To evaluate this aspect of ChatGPT, we first explore existing fact-checking and QA test sets and also illustrate the challenge of hallucination in ChatGPT by sharing hallucination examples.

We evaluate ChatGPT with test sets that consist of scientific and social claims related to COVID-19 (Lee et al., 2021). ChatGPT is able to detect misinformation 92% (46/50) and 73.33% (22/30, excluding verification-refusing cases) accuracy on covid-scientific and covid-social respectively. In comparison to its previously reported performance, ChatGPT’s performance on covid-scientific is impressive. Interestingly, for more societal-related claims, ChatGPT often refuses to make verification. However, it cannot avoid the criticism that parameterized knowledge is obtained by better memorization as it still shows worse performance in questions designed to cause imitative falsehoods. We test on 66 test samples from TruthfulQA (Lin et al., 2022), which tests the extent of LLMs to mimic human falsehood and 35.38% of the time ChatGPT fails to answer truthfully.

From various tasks, we often find extrinsic hallucinations, including both untruthful and factual ones, across various tasks such as Machine Translation and question answering, which causes degradation in performance. The intrinsic hallucinations are barely found as discussed in tasks about summarization and knowledge-grounded open-domain dialogue. We share examples of these hallucination types detected from different task explorations in Table 19.

Evaluating Interactivity in ChatGPT

ChatGPT has a built-in interactive ability thanks to conversational data fine-tuning and RLHF. We further delve into the benefit of exploiting this interactive ability of ChatGPT in three NLP tasks, summarization, machine translation, and multimodal generation. Our experiments demonstrate the potential of employing multi-turn interaction to refine the quality of the generated responses and improve the task performance of ChatGPT.

Summarization models aim to extract essential information from documents and to generate short, concise, and readable text (Yu et al., 2021b; Su et al., 2021). In real-world applications, people may want to improve the summary based on the previously generated summary. We ran experiments with 50 documents from SAMSum (Gliwa et al., 2019) and conducted a two-turn iterative prompt approach. ChatGPT usually generates an overly long summary. By adding a follow-up prompt after the first summary, ‘‘Please make the summary shorter’’, ChatGPT could provide a much shorter summary than the first response. Experimental results show that with the second length control prompt, the refined summaries achieve 7.99 and 1.64 gains on ROUGE-1 and ROUGE-2 respectively.

One of the capabilities of ChatGPT is to perform text translation from one language to another. With the interactivity of ChatGPT, we explore the possibility of performing a combined machine translation and automatic post-editing tasks to improve the translation quality of ChatGPT. For the experiment, we adapt the dataset used in §2.2.2. As shown in Figure 2, despite the translation and post-editing being done using a single ChatGPT model, the multi-turn approach method helps to improve the correctness of the translation by making partial corrections or even full corrections in some cases. We provide experimental setup details and examples of the post-editing in Appendix K.

The multi-turn interaction ability of ChatGPT enables the refinement of text-to-image generation. It is one of the most natural ways for humans to create artwork or product designs by requesting an AI tool iteratively. Through interaction with ChatGPT over multiple turns, a process of creating an interesting painting can be achieved (Figure 7).

To quantitatively study how this ability impacts image generation, we conduct at most three rounds of post-editing for the flag-drawing task. As shown in Figure 4, in the first round of generation, ChatGPT rarely generates errorless SVG images except for some simple flags (e.g., Nigerian and German). We observe that 34% and 36% of samples experience improvement (i.e., fewer errors) from turn 1 to 2 and from turn 2 to 3, respectively. We also tested with the InstructGPT, which has the same backbone model as ChatGPT but lacks conversation ability. InstructGPT cannot achieve salient improvements by directly putting the intermediate results in the input context (Section H.3).

Evaluation of GPT-4

As a successor of ChatGPT, GPT-4 was introduced in March 2023 with its technical report OpenAI (2023). Since the GPT-4 technical report focuses on professional and academic benchmarks (e.g., SAT), we evaluated GPT-4 on other LM abilities with our framework.We evaluated GPT-4 ‘gpt-4’ version API on 27 October 2023. We provide the details of the GPT-4 evaluation in Appendix I. Our findings from the GPT-4 evaluation are described as follows:

In terms of multitasking ability, GPT-4 is as good as ChatGPT with on-par performances in most of the common NLP tasks we tested.

GPT-4 shows better performance at language identification of extremely low-resourced languages (e.g., Buginese) and also better at machine translation, which aligns with the results reported on their technical report.

On commonsense reasoning, GPT-4 displays a very high performance which is close to ChatGPT. This finding aligns with the result reported on OpenAI (2023).

While for other reasoning tasks, GPT-4 generally yields higher performance than ChatGPT, especially on inductive, mathematical, multi-hop, temporal, and spatial reasoning skills.

Overall, our findings on GPT-4 showcase its superior ability against ChatGPT in various aspects, while there is still room for improvement in extremely low-resource languages and reasoning skills, especially on complex reasoning tasks.

Conclusion and Discussion

ChatGPT outperforms SOTA LLMs in a zero-shot manner on various tasks and even surpasses fine-tuned models on some tasks. However, there are still some failure cases (§2.1) and it produces responses with altered nuance and meaning. Therefore, dealing with these special cases is a complex but important task. In terms of multilinguality, ChatGPT achieves strong performance in many high-resource and medium-resource languages. Nevertheless, ChatGPT still lacks the ability to understand and generate sentences in low-resource languages (§2.2), which is also supported by Lai et al. (2023). Additionally, ChatGPT lacks the generation ability of non-Latin script languages (§2.2.2), despite the languages being high-resource. These raise the concern of language diversity and inclusivity in ChatGPT Joshi et al. (2020); Aji et al. (2022). Regarding multimodality, our flag drawing experiments showed the potential of ChatGPT’s multimodal ability. It would be an interesting research direction to further explore ChatGPT’s multimodal ability to answer “can textual models like ChatGPT switch to a multimodal backbone?”

The impressive performance of ChatGPT has sparked interest in expanding its usage beyond traditional NLP tasks into more complex domains requiring sophisticated reasoning such as problem-solving, decision-making, and planning. Our evaluation of its reasoning abilities shows that they are not reliable. Specifically, our findings indicate that ChatGPT exhibits a tendency to be a lazy reasoner and that its capabilities are inconsistent across various reasoning abilities; To support the further expansion of its use cases, it is necessary to prioritize the development of systems with robust complex reasoning capabilities, which should also be facilitated by the creation of more comprehensive benchmarks for assessing these abilities, such as works by Laskar et al. (2023b); Qin et al. (2023); Davis (2023a), particularly when multiple abilities are required to complete the tasks.

ChatGPT, like other LLMs, still makes things up (Ji et al., 2022a). To ensure factuality, it is possible to build LLMs with an interface to an external knowledge source, like Blenderbot 3.0 Shuster et al. (2022) and LaMDa Thoppilan et al. (2022). Meanwhile, there are many forms of hallucinations that are not necessarily counterfactual but still undesirable. The RLHF process of ChatGPT can ensure human feedback to mitigate undesirable responses. However, researchers need to work on coming up with more automatic and scalable methods to detect and mitigate hallucinations and other undesirable artifacts.

Compared with the previous LLMs, the interactive ability of ChatGPT has made a leap according to both qualitative and quantitative measures. Through interactivity, ChatGPT can recite from its own description, which is a very important ability. A similar exploration of this ability in LLMs has also been explored in other research works Sun et al. (2022); Wang et al. (2023). However, sometimes ChatGPT retains the wrong answer even after receiving multiple rounds of prompts from the user. Improving the ability of ChatGPT to handle multiple rounds of user feedback is also an important challenge.

Limitation

The experiments are done with the UI of ChatGPT provided by OpenAI (15 December 2019 version), before the ChatGPT API was released, thus, the number of samples for evaluation is limited (30-200). However, tasks of evaluation should not be affected much because most of the recent updates/releases of ChatGPT are related to safety concerns. Moreover, It is possible to augment our benchmarks with other technical benchmarks for research purposes, especially now that the ChatGPT APIs are available. There has been recent automatic or human-in-the-loop evaluations such as Laskar et al. (2023a) Nevertheless, many of the benchmarks are not necessarily interpretable to laypeople for general purposes, such as named entity recognition and etc. Our paper provides an easier-to-follow guideline.

Due to the page limit, many parts of the experimental setup details are added to the Appendix while the overall structure of evaluation and major insights stay in the main content. This may cause the reader inconvenience to follow the experiments. However, we publicly release the codebase that can help the community replicate the exact same evaluation either on ChatGPT or other LLMs easily.

Ethics Statement

Previous works have discussed the ethical implications or concerns associated with ChatGPT (and other LLMs) Jabotinsky and Sarel (2022); Susnjak (2022); Blanco-Gonzalez et al. (2022); Aydın and Karaarslan (2022); Jeblick et al. (2022). Agreeing with the previous literature, the responsible design and usage of LLMs including ChatGPT is an important and pressing challenge today. There are common issues with these models, such as fairness, toxicity, demographic bias, and safety, which need to be addressed. In the case of ChatGPT, OpenAI constructs safety layers and uses RLHF and potentially other means to filter out undesirable system responses. However, this is still not perfect and requires future research to further improve the robustness of the safety layer. This process is resource-intensive and opaque to the public. We hope to see a more open discussion and sharing of responsible design of LLMs from various organizations including OpenAI in the future.

This paper conducts an evaluation of ChatGPT for academic purposes only. We comply with the terms and conditions of ChatGPT stated in https://openai.com/policies/terms-of-use. Moreover, we comply with all the licenses of all the data (i.e., test sets/benchmarks) that are used in this evaluation.

Acknowledgments

This work has been partially funded by MRP/055/18 of the Innovation Technology Commission, Hong Kong SAR Government; the Hong Kong Fellowship Scheme by the Hong Kong Research Grants Council, and PF20-43679 Hong Kong PhD Fellowship Scheme, Hong Kong Research Grants Council.

References

Appendix

The appendix consists the following content:

D: Details for Multilinguality Evaluation

K: Examples from Machine Translation and Post-Editing

Appendix A Background and Related Work

Compared to existing LLMs, ChatGPT has unique characteristics. First, it has the ability to interact with users in a conversation-like manner, while retaining its accumulated knowledge and generalization ability gained from pre-training. This is achieved by pre-training ChatGPT on a large-scale conversational-style dataset, that is constructed by transforming a large-scale instruction-tuning corpus used for building InstructGPT into a conversational format, then fine-tuning the model based on a reward model to further improve the generation quality and align the generation with human

Second, ChatGPT is trained with a better human-aligned objective function via Reinforcement Learning from Human Feedback (RLHF) (Christiano et al., 2017). Conventional natural language generation models, including dialogue models, are trained with maximum likelihood estimation (MLE) and might not be aligned with human preferences. For instance, for dialogue systems, humanness, engagement, and groundedness are some examples of essential criteria for success. Such discrepancy between training objectives and evaluation metrics becomes a bottleneck to performance improvement. By using RLHF, ChatGPT aligns more closely with human preferences in generating text than by using MLE.

As ChatGPT has become available to public users through an easily accessible UI, there have been many discussions from a wide range of communities, not just from AI or NLP, but also from other disciplines. A line of discussion is the specific emergent ability and strength of ChatGPT in more technical perspectives. Guo et al. (2023) conducts linguistic analyses of ChatGPT’s writing against human experts and found that ChatGPT responses are strictly focused on the given question, more formal, objective, and less emotional. Nov et al. (2023) also studies ChatGPT’s generated medical advice if it passes the Turing test. Frieder et al. (2023) show that “significantly below those of an average mathematics graduate student.” There are many investigations of ChatGPT’s understanding and potential applications in different fields such as law Choi et al. (2023), medical domain (Blanco-Gonzalez et al., 2022; Jeblick et al., 2022) and finance Birch (2022); Dowling and Lucey (2023). Jeblick et al. (2022) conduct a case study of the application of ChatGPT on simplified radiology reports. Another important line of discussion is the ethical concerns over the use of ChatGPT. The most active discussion is over the use of academic writing and exam integrity (Jabotinsky and Sarel, 2022; Susnjak, 2022). OpenAI also discusses the misuse of LM for disinformation and remedies. https://openai.com/blog/forecasting-misuse/ Zhuo et al. study AI ethics of ChatGPT in criteria of bias, reliability, robustness, and toxicity.

A.2 LLM benchmark and evaluation

With the advancement of LLMs’ generalization ability, there have been efforts to understand their capabilities, limitations, and risks. Recently, several benchmarks with a collection of a large number of NLP datasets, such as BIG-Bench (Srivastava et al., 2022) and AI LM Harness (Gao et al., 2021), have been introduced. Moreover, HELM (Liang et al., 2022) is proposed to conduct a holistic evaluation of LLMs that considers scenarios and metrics with a top-down approach. In this work, we instead focus on specific limitations and unique findings of ChatGPT that had not been discussed with previous LLMs.

There are also other works that discuss LLMs’ emergent abilities through thorough surveys or case studies. Mahowald et al. (2023) thoroughly studies LLMs capabilities by distinguishing formal and functional linguistic competence with reference to cognitive science, psychology, and NLP to clarify the discourse surrounding LLMs’ potential. Other works focus on more specific abilities such as mathematical skills (Davis, 2023b), reasoning (Webb et al., 2022a; Qiao et al., 2022). Also, there have been overviews of existing LLMs (Gozalo-Brizuela and Garrido-Merchan, 2023; Wolfe, 2023)

A.3 ChatGPT Evaluation

To the best of our knowledge, this benchmarking exercise is the first of its kind. Since the introduction of ChatGPT with its advancement, there has been a huge amount of assessments of ChatGPT to understand its limits. Mao et al. (2023) provides a survey of recent assessments of ChatGPT in broad categories of 1) Language and Reasoning Ability, 2) Scientific Knowledge, and 3) Ethical Considerations. Laskar et al. provide extensive automatic or human-in-the-loop evaluations on 140 tasks. Qin et al. mainly evaluated the reasoning abilities of ChatGPT while Zhuo et al.; Ray focus on other important aspects such as ethics, robustness, reliability, limitations, and future scope of ChatGPT. Kocoń et al. examined whether the high quality of the LLM can indicate a tool’s usefulness to society by evaluating ChatGPT’s capabilities on 25 diverse analytical NLP tasks, most of them subjective even to humans. After the introduction of ChatGPT, GPT-4 has been introduced by OpenAI. However, OpenAI is not disclosing any internal benchmarking of ChatGPT. Even in their GPT-4 technical report OpenAI (2023), they have shown the performance of GPT4 in terms of human-level exams. So, it is important that there are 3rd party evaluations of generative models.

Appendix B General Experimental Details

The experiments were done with the UI (15 December 2019 version) of ChatGPT provided by OpenAI, before the ChatGPT API was released. The number of samples for evaluation is 30-200. We’ve prioritized sample diversity, hand-picking tasks that encapsulate the abroad spectrum of scenarios a language model is likely to encounter, thus creating a representative snapshot of potential real-world applications. All experiments are single-run.

Appendix C Multitask Evaluation of ChatGPT

We test on 100 samples from two common summarization datasets: half from SAMSum (Gliwa et al., 2019), a dialogue summarization dataset, and another half from CNN/DM (Hermann et al., 2015; Nallapati et al., 2016), news summarization datasets. The large version of Bart Lewis et al. (2020b) model fine-tuned on both datasets is conducted for comparison. Moreover, OpenAI’s text-davinci-002 is used as the previous SOTA zero-shot model. We calculate ROUGE-1 scores for evaluating the generated summary. According to the evaluation, ChatGPT achieves a similar zero-shot performance with text-davinci-002, which is expected since they evolved from the same GPT3 pre-trained checkpoint. However, the fine-tuned Bart still outperforms zero-shot ChatGPT by a large margin.

C.2 Machine Translation

We evaluate the machine translation ability of ChatGPT on both high-resource and low-resource languages using the ChrF++ metric Popović (2015). Specifically, we incorporate 8 high-resource languages, i.e., French (fra), Spanish (spa), Chinese (zho), Arabic (ara), Japanese (jpn), Indonesian (ind), Korean (kor), and Vietnamese (vie), and 4 low-resource languages, i.e., Javanese (jav), Sundanese (sun), Marathi (mar), and Buginese (bug) for our evaluation. For a fairer comparison in our multitask experiment, we strictly follow the definition of high-resource and low-resource languages from NLLB Team et al. (2022). For each language pair, we sample 30 Eng↔\leftrightarrowXXX parallel sentences from the FLORES-200 dataset (Team et al., 2022; Goyal et al., 2022a). The result of our experiment suggests that ChatGPT can well perform XXX→\rightarrowEng translation, but it still lacks the ability to perform Eng→\rightarrowXXX translation.

C.3 Sentiment Analysis

Sentiment analysis has been widely explored for both high-resource and low-resource languages Wang et al. (2018a); Wilie et al. (2020); Ilmania et al. (2018).

We explore the sentiment analysis ability of ChatGPT through 4 languages with diverse amounts of resources in NusaX (Winata et al., 2022): English (eng), Indonesian (ind), Javanese (jav), and Buginese (bug). For each language, we sample 50 sentences from the corresponding dataset for our experiment and measure the macro F1 score as the evaluation metric. We compare the results with two baselines, i.e., supervised state-of-the-art performance from Winata et al. (2022) and zero-shot multilingual LLM from Cahyawijaya et al. (2022). ChatGPT outperforms the previous state-of-the-art zero-shot model by a large margin except for the Buginese, where it performs on par. This shows that ChatGPT still has a limited understanding of extremely low-resource languages.

C.4 Question Answering

Since Question Answering (QA) is a broad topic, we classify QA datasets into different categories based on the knowledge/reasoning type required to do the task, e.g commonsense reasoning, spatial reasoning, temporal reasoning, etc., to have a clearer analysis on ChatGPT’s abilities. For each category, we select several datasets, and for each dataset, we sample 30 instances and test ChatGPT on the subset. Based on our experiment results, ChatGPT outperforms the existing zero-shot and some of the fine-tuned state-of-the-art performance on question answering. Furthermore, ChatGPT achieves near-perfect scores on three tasks, i.e., bAbI task 15, EntailmentBank, and Pep-3k.

C.5 Misinformation Detection

We test ChatGPT’s ability to detect misinformation with the test sets that consist of scientific and social claims related to COVID-19 (Lee et al., 2021) with 100 samples. We take half from scientific (covid-scientific) and another half from social (covid-social) sets. We evaluate the accuracy of the veracity by manually checking the generated text. ChatGPT could detect misinformation 92% (46/50) and 73.33% (22/30, excluding verification-refusing cases) accuracy on covid-scientific and covid-social respectively.

C.6 ChatGPT on Dialogue Tasks

Open-domain dialogue systems interact with humans with generated responses automatically and aim to provide users with an engaging experience. To boost informativeness, these systems leverage external knowledge, including structured knowledge such as knowledge graphs Zhao et al. (2020); Ji et al. (2022b) and unstructured knowledge such as free text Xu et al. (2022, 2023).

Prompt used for experiment: ‘‘Can we try dialogue generation? I will give you turns, and you can generate the next turn, but only one.\n \n You can also consider the knowledge of XXX for your reference in the dialogue.’’

C.6.2 Task-Oriented Dialogue Experimental Setups

We investigate ChatGPT’s ability for both dialogue state tracking and response generation in 50 dialogue turn samples taken from MultiWOZ2.2 (Zang et al., 2020). In detail, we ask the model to provide the belief state as domain-intent: [slot1, value1], … in the prompt following previous zero-shot Lin et al. (2021) and few-shot Madotto et al. (2021) approaches, and provide an exhaustive list of domain-intent-slot-value for the given dialogue. For the response generation, we provide only the oracle dialogue actions (e.g. ’Hotel-Inform’:[’area’, ’centre’]), and ask ChatGPT to generate a TOD response given the dialogue history. We assess DST with joint goal accuracy (JGA), the ratio of dialogue turns where the predicted dialogue state is exactly the ground truth, and response generation with BLEU and inform rate(%)

We explore ChatGPT’s ability to simulate a TOD interaction in an end-to-end manner by providing nothing more than a structured database and giving the instruction: ‘‘Use the following knowledge base to complete the task of recommending a restaurant as a task-oriented dialogue system’’.

Result Analysis: We could investigate whether ChatGPT is able to complete basic retrieval queries and respond to users’ requests such as “Give me some restaurants that serve Italian food” or “I would prefer cheap options please”. However, there are several limitations that we could investigate as follow.

Long-term Multi-turn Dependency: ChatGPT cannot keep the belief state across multiple turns within the interaction. For instance, asking for Italian food will overwrite the previous turn’s belief state by asking for restaurants with a rating of 3 or higher. However, if the user explicitly asks to recall the earlier preferences, ChatGPT is able to correct the retrieved information and incorporate the previous belief state. This is interesting as it shows that the information previously given in multi-turn is still usable, but needs to be called explicitly.

Basic Reasoning Failure: ChatGPT’s response tends to be wrong if the query introduces a basic level of reasoning such as when it is asked for “recommendation for restaurants with European food” (ChatGPT has to filter the types of cuisine which are based on countries) or “recommendation for restaurants with a rating of 3 or higher” (ChatGPT needs to understand rating 3, 4 and 5). Even with a basic knowledge base, ChatGPT fails to answer correctly 66% of the time.

Extrinsic Hallucination: ChatGPT tends to generate hallucinated information beyond the given knowledge. This is especially harmful in TOD as ChatGPT will sometimes hallucinate some prices for hotel booking, or availability for restaurants.

We provide the example for the modular and unified approaches for Task-Oriented Dialogue in Table 6 and Table 7, respectively.

Appendix D ChatGPT on Multilinguality

We present the statistics of language under study in Table 9. In the following section, we provide the insights that we find during our experiment in exploring multilingual capability of ChatGPT.

As shown in Table 10, ChatGPT correctly classifies the languages for English and Indonesian 100% of the time. While for the language identification for Javenese and Buginese, ChatGPT either misclassifies the samples as other languages or is unable to determine the language. Nevertheless, ChatGPT performance on the sentiment analysis in Javanese is only slightly lower compared to English and Indonesian which suggests that ChatGPT can understand the semantic meaning of sentences in low-resource languages without having the knowledge to identify the language itself. This limitation of language identification in LMs aligns with the result from BIG-bench Srivastava et al. (2022).

As shown in Table 8, ChatGPT lets the user know that its prediction is uncertain when it does not completely understand the language and also provides broader information regarding the language, such as location and tribe of which the predicted language is spoken. This fact provides evidence regarding the benefit of using the RLHF approach compared to other training approaches for aligning LLMs with human preferences.

Despite being high-resource and medium-resource languages, the translation from English to Chinese and Korean is much inferior to the other languages with Latin scripts, i.e., French or Indonesian. Similarly, prior works focusing on transliteration Chau and Smith (2021); Muller et al. (2021) have shown the effectiveness of utilizing Latin scripts over other scripts, e.g., Cyrillic, Georgian, Arabic, etc, especially for low-resource languages. Interestingly, this problem of using non-Latin scripts is less severe for translation from Chinese and Korean to English, which suggests that ChatGPT can better neutralize the effect of non-Latin scripts as source languages Wan (2022), but it still lacks the ability to generate non-Latin script languages.

Appendix E Multimodality: Flag Drawing Task

We uniformly collect 50 national flags from different continents and conduct the flag-drawing task on ChatGPT. The flag-drawing task contains three steps:

Ask ChatGPT to illustrate the appearance of the flag using the prompt “Describe how the flag looks like”.

Based on the description, ask ChatGPT to generate the SVG code of that flag by prompting “Generate a code snippet to represent that flag in SVG format”.

If the generated image contains errors, we iteratively ask ChatGPT to fix them.

There are four types of evaluation criteria: 1) layout 2) color 3) missing components 4) shape/size. In each round of fixing, we ask ChatGPT to revise only one type of error with the prompt “. Revise the image”. We terminate the conversation once the generated flag becomes perfect or we have already passed two rounds of fixing.

The generated flag images are evaluated by the aforementioned four error types as criteria. We further assess the image quality with five grades, A ∼\sim E, which indicate zero to four (or above) errors. We assign grades to each round so that we can assess the number of improvements and degradation through conversational interactions (post-editing). The full results are shown in Figure 4.

Appendix F Details for Reasoning Evaluations

Table 11 shows the categories of reasoning that are evaluated in this paper as well as corresponding datasets. The following section introduces each of the categories and detailed experimental setup and/or analysis.

Inductive and deductive are categorized by “a degree to which the premise supports the conclusion” based on logic and philosophy (Qiao et al., 2022; Rogers et al., 2022; Hawthorne, 2021). Inductive reasoning is based on “observations or evidence” while deductive is based on “truth of the premises” (i.e., necessarily true inference) (Douven, 2017). Another way to categorize is based on the “direction of reasoning” – deductive is from premise to conclusion while abductive is from conclusion to the most probable premise that supports the conclusion (Walton, 2014).

Inductive and deductive reasoning are common forms of logical reasoning that are categorized by “a degree to which the premise supports the conclusion” based on logic and philosophy Qiao et al. (2022); Rogers et al. (2022); Hawthorne (2021). Deductive reasoning involves processes of driving specific conclusions based on more general premises. On the contrary inductive reasoning involves specific observation of patterns, processing them on increasingly abstract cycles of hypothetico-deductive reasoning to draw a more general conclusion (Lawson, 2005). Comparing the two types of reasoning, deduction requires less “guessing” from the perspective of ChatGPT, as induction requires figuring out rules (Rogers et al., 2022). The former can be viewed as top-down while the latter is bottom-up.

Deductive reasoning involves processes of driving specific conclusions based on more general premises. On the contrary, inductive reasoning involves specific observation of patterns, processing them on increasingly abstract cycles of hypothetico-deductive reasoning to draw a more general conclusion (Lawson, 2005). Comparing the two types of reasoning, deduction requires less “guessing” from the perspective of ChatGPT, as induction requires figuring out rules (Rogers et al., 2022). The former can be viewed as top-down while the latter is bottom-up.

We explore ChatGPT’s ability of inductive and deductive reasoning in two different levels: 1) basic and 2) advanced. Basic-level tasks are the prerequisites to probe reasoning. While solving these tasks does not necessarily indicate full reasoning capability, if ChatGPT fails on any of these tasks, then there are likely real-world tasks that it will fail on too if they require similar reasoning mechanisms. Consequently, the advanced-level tasks are there to probe those capabilities in real-world tasks where the noises are present, and solving them requires a more systematic generalization. Additionally, we choose tasks that do not require or are dependent on external knowledge and the solution could be only derived by premises to focus on dissecting the capability of each reasoning mechanism.

ChatGPT answers “It is not specified what ¡attribute¿ ¡entity¿ is.” for most of the time when it was asked a question requiring inductive reasoning. However, when ChatGPT is explicitly asked for reasonable inference with a prompt “Based on the given facts, do a reasonable inference on this question using inductive reasoning:”, its ability for inductive reasoning increases. Yet, it is still not as good as in deduction as the same prompt engineering also helps increases its ability for deductive reasoning.

We could derive similar insight as ChatGPT only correctly answered for half of the time while it could make inferences deductively well for 90% of the time. CLUTRR Sinha et al. (2019) requires induction on extracting relations between entities, and in the ChatGPT responses, it often asks for more information to make inferences. An interesting finding along with CLUTRR was that ChatGPT can’t differentiate son and grandson but can differentiate daughter and granddaughter when it induces the logical rules governing kinship relationships. We show all performances in Table 12 and some of the prompting samples in Table 13. We follow Qiao et al. (2022) categorization on the deductive and inductive reasoning datasets, but we only use the QA part of EntailmentBank, that the authors took from ARC dataset Clark et al. (2018), as we aim to test for reasoning capability. Regarding EntailmentBank, it might trigger the universe-related knowledge out of ChatGPT, which could help the model to derive the correct answer, although the test set is designed to test deductive reasoning skills. One of the future explorations would be with checking the rationale of ChatGPT as a follow-up question.

F.1.2 Abductive Reasoning

Abductive reasoning is the inference to the most plausible explanation given observations. For instance, “if Jenny finds her house in a mess when she returns from work, and remembers that she left a window open, she can hypothesize that a thief broke into her house and caused the mess” An example provided by Bhagavatula et al. (2020).. We test ChatGPT’s language-based abductive reasoning ability with 30 samples from α\alphaNLI dataset (Bhagavatula et al., 2020), which requires the model to select the most plausible explanation given the conclusion. Based on our test, it could achieve 86.7% (26 out of 30) accuracy.

F.2 Non-textual Semantic Reasoning

Mathematical capabilities or numerical reasoning has been frequently mentioned to be lacking for LLMs, not only ChatGPT (Frieder et al., 2023). Frieder et al. test ChatGPT’s capability with publicly available datasets as well as the human-curated dataset, which consists of 728 prompts. The shared findings for ChatGPT’s mathematical capabilities include 1) ChatGPT often understands the question but fails to provide correct solutions; 2) it shows inconsistent poor performance on graduate-level advanced mathematics; 3) it has a great ability to search for mathematical objects. Refer to detailed findings in the original paper. We also test separately on MATH dataset. Not surprisingly, it could only score 23.33% (7/30) for the MATH dataset (Saxton et al., 2019), which tests mathematical reasoning.

Temporal reasoning is mentioned a few times in the literature but is less common than others. It tests the understanding of the time duration of and the relation between events. For this category, we conduct experiments on the dataset TimeDial (Qin et al., 2021), which solely requires temporal reasoning. We follow the format of the task in the BIG-bench benchmark (Srivastava et al., 2022), which is multiple-choice (single correct answer), Overall, ChatGPT correctly answers 86.67% of the time (26/30), suggesting that it has a decent temporal reasoning ability. Also, compared to Chinchilla and Gopher which have the accuracy of 68.8% and 50.9% respectively, ChatGPT shows a promising improvement for LLMs in that aspect.

Spatial reasoning is using an understanding of spatial relations among different objects and spaces. For spatial reasoning, we utilize two existing datasets: SpartQA (Mirzaee et al., 2021) and StepGame (Shi et al., 2022a), which compose of story-question pairs about k relations of k+1 (where k is up to 10) entities written in natural language. ChatGPT is asked to answer spatial relations between two entities based on the provided descriptions of different entities. ChatGPT falls short of the spatial reasoning tasks, as shown in Table 15, with overall success rates of 43.33% for StepGame and 43.75% for SpartQA. ChatGPT could only score 25% on SpartQA (hard), which covers multiple spatial reasoning sub-types, and 23.33% for stepGame (Hard) with k=9. ChatGPT could not provide any spatial relations but instead generated “It is not specified in the given description”. Even with the fine-tuned models, as the number of relations (k) increases in context description, performance drops (Shi et al., 2022a).

To understand spatial reasoning ability at a more elementary level, we test with less complicated examples from StepGame which we refer to as StepGame (Basic). It does not involve multi-hop reasoning but purely spatial relation between two entities. (e.g, “C is sitting at the top position to Y. What is the relation of the agent Y to the agent C?”). We test for basic spatial relations with 8 labels from StepGame {left, right, above, below, lower-left, lower-right, upper-left, upper-right}. When we test on StepGame (Basic), ChatGPT scores higher (63.33%).

We investigate the errors that it often fails to understand clock direction (e.g., “W is at K’s 3 o’clock”) and diagonal spatial relations. We further analyze the results by breaking down the test examples of StepGame (Basic) into two comparisons: i) types of directions (basic cardinal vs. diagonal) and ii) ways of spatial description for cardinal directions (basic cardinalThose of which spatial relations are described with explicit vocabulary. vs. clock-position cardinal). We take 20 more samples for each category (basic cardinal, diagonal, clock-position cardinal) and tested them as illustrated in Table 14.

ChatGPT poorly infers with clock-position description. Although it is a simple cardinal direction, ChatGPT could only correctly answer for 5 samples (25%), which is clearly poorer performance in comparison to performance with the basic cardinal description (17 correct answers).

ChatGPT is worse at the diagonal position. It correctly answers around half of the time (55%), which is worse than basic cardinal points (85%). Even with analysis from StepGame (Hard), among the correct 7 answers, there is only one diagonal direction that ChatGPT gets correctly while the others are all cardinal points. For those answers that require diagonal points, ChatGPT only could infer cardinal points for some examples.

F.3 Commonsense Reasoning

To evaluate ChatGPT’s capability on commonsense reasoning, we first test it on two widely used benchmark datasets CommonsenseQA Talmor et al. (2018) and PiQA Bisk et al. (2020). CommonsenseQA focuses on general commonsense question answering such as “Where is a business restaurant likely to be located?”, and PiQA is about physical commonsense reasoning: given a sentence such as “When boiling butter, when it’s ready, you can ”, the goal is to fill in the blank with one of two answer options, “Pour it onto a plate” and “Pour it onto a jar”. We use the validation split for both of the datasets since there are no labels provided on the test set that we retrieve. We also further probe ChatGPT by evaluating a more challenging commonsense reasoning dataset in a more comprehensive way. We use Pep-3k Wang et al. (2018b), which requires the model to recognize plausible but possibly novel events, such as “man swallow paintball”. Each instance in the Pep-3k is an s-v-o predicate, and the task is to judge if the predicate is plausible or not. But instead of evaluating ChatGPT’s performance only based on the binary judgment, we also check if the answer contains relevant rationales (explanations) that lead to its judgment.

For the Pep-3k samples, we prepend the s-v-o predicate with “Please judge if this predicate is (likely) plausible or implausible:” to prompt ChatGPT. We show the results in Table 16. As we see, ChatGPT performs quite well on the three datasets in terms of answer accuracy, which matches our anticipation. Furthermore, as we also check the rationales in ChatGPT’s answer when evaluating Pep-3k samples, we can see that ChatGPT does quite well not only in terms of answer accuracy but also in generating reasonable reasoning procedures to support its answer. We show a concrete example in Table 17. As we can see, ChatGPT’s answer explains well what kinds of materials are usually cut through with knives (i.e., food, paper, or wood). Then, it reasons why rocks cannot be chopped with a knife by explaining ‘rocks are much harder than these materials.’ While our findings are based on 30 samples from each dataset, we see the potential in ChatGPT’s commonsense reasoning capability, and further large-scale investigation is worth exploring.

F.4 Causal, Multi-Hop, and Analogical Reasoning

Causal reasoning is the process of identifying the relationship between causes/actions and effects/changes (i.e., causality) (Thomason, 2018; Huang and Chang, 2022). We test ChatGPT on 30 samples of human-annotated explainable CAusal REasoning dataset (E-CARE) (Du et al., 2022) and it could score 24 samples correctly (80%). Note that our evaluation is mainly based on whether the model can make a judgment on correct causes or effects instead of its generated explanation of why the causation exists.

To be able to reason over a larger context, a system has to perform multi-hop reasoning over more than one piece of information to arrive at the answer Mavi et al. (2022). We test ChatGPT’s multi-hop reasoning capability on 30 samples of HotpotQA dataset Yang et al. (2018) and we find that ChatGPT has difficulty performing with such capability, only answering 8 samples correctly, although the questions posed are only 2-hops. It is worth noting that ChatGPT oftentimes generates the answer in a short passage of explanations, thus we evaluate manually each of the ChatGPT responses to check its accuracy. This aligns with the findings that LLMs are also limited in several ways, and fail to produce accurate predictions due to their inability to accomplish complex reasoning, such as solving tasks that require multi-hop reasoning Ott et al. (2023).

Analogical reasoning is a way of thinking that relies upon an analogy, comparing two or more objects or systems of objects (Bartha, 2013) to drive a conclusion. We test with 30 samples from Webb et al. (2022b) and evaluate based on human evaluation, to see if the generated answer match with/contain the gold answer. ChatGPT could correctly answer all 30 examples, which may reveal that ChatGPT has a good capability in analogical reasoning skills.

Appendix G Details for Hallucination Evaluations

There exist two categories of hallucination Ji et al. (2022a). Intrinsic hallucinations that refers to the LLM generation that contradicts the source/input content. Extrinsic hallucinations that refers to the LLM generations that cannot be verified from the source/input content (i.e., output that can neither be supported nor contradicted by the source). In Table 19, we share examples of these hallucination types detected from different task explorations. With the setting of tasks we test, we often find extrinsic hallucinations, including both untruthful and factual ones, across various tasks such as Machine Translation, Question answering.

The intrinsic hallucinations are barely found. For instance, in the abstractive summarization task, in which neural models usually suffer from intrinsic hallucination, ChatGPT’s generated summarisation did not include any intrinsic hallucination examples based on our experiments. It rather shows a factual extrinsic hallucination, for instance, ChatGPT could correctly paraphrase “Britain and five other countries” from source input into “P5+1 (US, UK, France, China, Russia, and Germany),” which is assessed to be factual. We could also observe an interesting intrinsic hallucination for our proposed multi-modal task, the flag drawing task. ChatGPT is first asked to generate a description of how the flags look before it is asked to generate code for the flag. Although it generates the correct description as “The flag of Mexico consists of three vertical bands […]”, the final drawing (SVG code) consists of horizontal bands.

However, extrinsic hallucinations often happen, including both untruthful and factual ones. In the QA task, we often find extrinsic hallucination to be non-factual which harms the final performance. For instance, in the question of asking for the relationship among entities, although step kindship is never mentioned in the question, ChatGPT answers the question with step kinship, as illustrated in Table 19. We could also observe that ChatGPT’s weakness with extrinsic hallucination also degrades machine translation. When it is asked to translate the text “Like some other experts, he is skeptical about whether diabetes can be cured, noting that these findings have no relevance to people who already have Type 1 diabetes.” into Korean, it contains a piece of information that was not found in the source, “저주파 치료” (transcutaneous electrical nerve stimulation) in the translated text.

Appendix H Details for Interactivity Evaluation

Figure 5 shows an example of how multi-turn interaction helps to control the length of the summary.

Given an input dialogue as the context, we first input the prompt ‘‘Summarize the above dialogue’’ to the ChatGPT.

To refine the summary, we simply input another prompt – ‘‘Please make the summary shorter’’ after the first response.

Evaluation: We calculate the ROUGE scores (ROUGE-1, ROUGE-2, and ROUGE-L) of the first and second summaries and compare between turns.

H.2 Interactivity on Machine Translation

We explore the capability on translation from English to the target language. For the experiment, we adapt the dataset used in §2.2.2 which samples 30 parallel sentences from 6 language pairs in NusaX Winata et al. (2022), Chinese, French, Indonesian , Korean, Javanese, and Sundanese.

Query model to translate to the target language using ‘‘What is [TARGET_LANGUAGE] translation of the following sentence?\n\n[INPUT_SENTENCE]’’

Query for the post-editing using the following prompt template: ‘‘Could you perform a post-editing to ensure the meaning is equivalent to ‘‘[INPUT_SENTENCE]"?’’

Evaluation: The post-editing results are manually validated by a native speaker in the corresponding language to validate: 1) whether the post-edited sentence is better than the translation one, and 2) whether the post-edited sentence is the correct translation of the given English sentence.

Based on the evaluation, performing automatic post-editing through interactive LLMs, such as ChatGPT, yields consistently better translation results compared to a single-turn machine translation, which is especially useful for translation in low-resource languages. We provide per-language examples of the machine-translated and post-edited sentences in Appendix K.

H.2.2 Experiment 2: Automatic post-editing

To further strengthen our hypothesis, we conduct an additional experiment on the automatic post-editing (APE) shared task dataset on WMT 2022 Bhattacharyya et al. (2022), which focuses on English→\rightarrowMarathi post-editing task. Marathi (mar) is also a low-resource language with 0.02% data size on CommonCrawl. We sample 50 samples from the corresponding dataset.

Evaluation: 1) human-targeted translation error rate (HTER)HTER is the official evaluation metric used in the APE 2022 shared task., SacreBLEU Post (2018) and METEOR Banerjee and Lavie (2005) between the Marathi generated sentence compared to the human post-edited sentence, 2) HTER, SacreBLEU, METEOR, and semantic similarity score, i.e., BERTScore Zhang* et al. (2020), between the English back-translated sentence and original English sentence.the back translation process is done via Google Translate (https://translate.google.com/).

As shown on Table 20, the single-turn translation without post-editing produces a slightly better evaluation score on the Marathi language, but the multi-turn with post-editing consistently yields better evaluation performance on the back-translated English text on all metrics. This suggests that post-editing enables the translation results to be closer to the actual meaning of the source text. Nevertheless, the translation to the Marathi language is much worse compared to the baseline MT provided from the APE 2022 shared task Bhattacharyya et al. (2022) which further supports the limitations of ChatGPT on generating sentences in low-resource and non-Latin script languages.

H.3 Interactivity on Multimodal Generation

We show an example of a multi-turn flag drawing of InstructGPT, which has the same backbone model as ChatGPT but lacks conversation ability, in Figure 6. Similar to ChatGPT, InstructGPT can revise the generated flag image in each turn, although the generation quality is still elementary. Figure 7 shows the process of creating an interesting painting by prompting ChatGPT with varied requirements through multiple turns.

Appendix I Results for Evaluation of GPT-4

I.2 Results on Mulilinguality

I.3 Results on Reasoning

Appendix J List of Evaluation Datasets

We provide a detailed list of all the datasets used in our experiment on LABEL:tab:datasets-complete.

Appendix K Examples from Machine Translation and Post-Editing