GPTEval: A Survey on Assessments of ChatGPT and GPT-4

Rui Mao, Guanyi Chen, Xulang Zhang, Frank Guerin, Erik Cambria

Introduction

ChatGPT (OpenAI, 2023b) has generated significant scholarly interest across various disciplines due to its impressive dialogue-based task-processing capabilities. This has enabled users to explore and evaluate its performance across a wide range of tasks and disciplines, thereby sparking considerable enthusiasm in the field of Artificial Intelligence (AI). While many researchers have concentrated on evaluating ChatGPT and GPT-4 (OpenAI, 2023a) within their specific domains of expertise, a comprehensive review encompassing the assessments in multiple tasks and disciplines can offer a holistic understanding of the strengths and limitations of these GPT models (see Appendix A).

We focus on ChatGPT and GPT-4, because many evaluation reports have come to similar conclusions – ChatGPT and GPT-4 are the state-of-the-art (SOTA) large language models (LLMs) by now. The scope of our survey encompasses quantitative evaluations carried out on ChatGPT or GPT-4, specifically focusing on their language proficiency, scientific knowledge, and ethical considerations.

Our main findings are summarized as follows: a) ChatGPT and GPT-4 are strong in language understanding and generation, adeptly engaging in user interactions through dialogues, enabling them to tackle diverse NLP tasks and provide explanatory outputs. However, their current status falls short of being a comprehensive AI, as their performance lags behind expert models in numerous domains involving domain-specific knowledge.

b) ChatGPT performs satisfactorily in general science knowledge and can answer science questions that desire open responses. However, it can also make mistakes, especially for questions that require multi-step reasoning. The exceptional language proficiency poses challenges for users in assessing the accuracy of factual information, giving rise to a range of ethical concerns.

c) Existing evaluation methods may be unreliable. The current evaluation methods heavily depend on prompt engineering and benchmark datasets. Varying prompts can yield disparate evaluation results. Additionally, the comparison of expert systems often relies on (in-domain) datasets that were utilized for training those systems. It remains uncertain whether the examined data, such as public datasets and scientific knowledge, have been inadvertently exposed during the training of ChatGPT and GPT-4. These factors may contribute to an unfair comparison between LLMs and their respective baselines.

The contributions of this work are threefold: (1) We conduct a comprehensive survey of recent assessments focusing on the language proficiency and scientific knowledge of ChatGPT and GPT-4. (2) We compare their assessment results across various tasks and disciplines to highlight the strengths and weaknesses of the GPT models. (3) We critically analyze the existing assessment methods employed, offering recommendations for future evaluation studies and delivering our ethical considerations associated with the GPT models.

Language and Reasoning Ability

Cabrera and Neubig (2023) compared several chatbots with a novel LLM evaluation toolkit, Zeno Build. The performance was evaluated by Critique metricshttps://docs.inspiredco.ai/critique/ such as ChrF (a character- and word n-grams similarity-based metric, Popović, 2015), BERTScore (a BERT embedding similarity-based metric, Zhang et al., 2019), and UniEval Coherence (a coherence probability-based metric, Zhong et al., 2022). They found that ChatGPT surpassed all the baselines, e.g., GPT-2 (Radford et al., 2019), LLaMa (Touvron et al., 2023), Alpaca (Taori et al., 2023), Vicuna (Chiang et al., 2023), MPT-Chat (MosaicML, 2023), and Cohere Commandhttps://docs.cohere.com/docs/command-beta in the three evaluation metrics. However, Cabrera and Neubig (2023) also highlighted that ChatGPT exhibited vulnerabilities that were noticeable in various aspects, such as occurrence of hallucinations, inadequate exploration for additional information, and repetition of content. Although ChatGPT exceeded other LLMs, Bang et al. (2023) observed that SOTA models still outperformed ChatGPT on task-oriented dialogue, and open-domain knowledge-grounded dialogue, based on automatic evaluation metrics.

Generation

was almost only evaluated on text-to-text generation tasks (but not data-to-text). For machine translation (MT), an early assessment suggested SOTA MT systems could defeat ChatGPT by a large margin (Bang et al., 2023). Later on, by testing on more language pairs, more datasets, and better prompts, Hendy et al. (2023) found ChatGPT yielded competitive performance for high-resource languages, but still had limited capabilities for low-recourse languages. This was further validated by Jiao et al. (2023), where the authors additionally concluded that ChatGPT was not good at translating between distant languages and translating sentences that contain domain biases or noise, compared to commercial MT systems (e.g., DeepL and Google Translate), but the gap could be narrowed for GPT-4. However, through a large-scale human evaluation and error analysis by expert translators, Karpinska and Iyyer (2023) had very different conclusions. They found that, when doing paragraph-level translation, ChatGPT’s translations were overwhelmingly preferred compared to those from Google Translate. ChatGPT could largely reduce errors, including mistranslation, grammatical errors, inconsistency errors, and more. For summarization, by testing on multiple summarization datasets in multiple languages, Qin et al. (2023); Bang et al. (2023); Lai et al. (2023); Zhang et al. (2023a) observed that ChatGPT largely underperformed SOTA systems in doing either abstractive or extractive summarization. Interestingly, Zhang et al. (2023a) found that its performance, especially the faithfulness, improved if we asked ChatGPT to first extract salient sentences and then generate the summarization, based on the extracted sentences, although it still lost to SOTA. However, similar to what happened when assessing ChatGPT’s translation ability, in a small human evaluation (i.e., 50 summaries rated by 2 annotators), annotators could not distinguish summaries generated by ChatGPT from those by human beings (Soni and Wade, 2023). This conclusion should be further confirmed by larger populations and rigorous experiments.

Although ChatGPT is often considered to perform well in generation tasks that need creativity, there have not been many related assessments. Jentzsch and Kersting (2023) asked ChatGPT to produce jokes and suggested that while ChatGPT generated jokes, it struggled to produce “new” jokes. Approximately 90% of the generated jokes were repetitions of the same 25 jokes. Chu and Liu (2023) conducted reader experiments and found that ChatGPT wrote more engaging and persuasive short stories than its human counterparts. However, the conclusion was the opposite if the aim was long stories. They suggested that ChatGPT might focus on retaining information from the instruction when writing a long story, which limited its flexibility.

Affective Computing.

Amin et al. (2023) compared ChatGPT to naive machine learning models on Personality Detection, Sentiment Analysis, and Suicide Detection tasks. The baselines were trained with task-specific datasets using RoBERTa embeddings, Word2Vec, and Bag-of-Words (BoW) vector representations. ChatGPT yielded better results on Sentiment Analysis, while lacking significant improvements, compared to a basic BoW model. The baselines outperformed ChatGPT on the other two examined tasks. Qin et al. (2023); Bang et al. (2023) found that fine-tuned models exceeded ChatGPT on English sentiment analysis tasks. Kocoń et al. (2023) also found that ChatGPT was weaker in emotion detection than other SOTA models.

Information Retrieval.

Wei et al. (2023) examined ChatGPT on relation extraction, named entity recognition (NER), and event extraction tasks, showing that the performance of the basic version of ChatGPT was much weaker than supervised methods, while largely exceeding 50 shot or fewer shot baselines. Then, they introduced a multi-turn question answering (QA) framework, where the modified querying process did not help ChatGPT to exceed the supervised baseline. Qin et al. (2023) observed that ChatGPT yielded much weaker performance than fine-tuning-based models on CoNLL03 NER dataset (Sang and De Meulder, 2003). Sun et al. (2023) evaluated ChatGPT and GPT-4 on passage re-ranking tasks, showing that GPT-4 outperformed SOTA supervised baselines across all three datasets, including a multilingual dataset with 10 languages. ChatGPT slightly fell behind SOTA methods, while largely exceeding BM25. Bubeck et al. (2023) found that GPT-4 (77.4% accuracy) significantly outperformed the SOTA model (40.8%, Payne, 2020) on a personally identifiable information detection task.

GPT as Human Annotator.

Wang et al. (2023c) compared ChatGPT with SOTA natural language generation evaluation metrics, e.g., BERTScore, BARTScore (Yuan et al., 2021), ROUGE (Lin, 2004), and more. ChatGPT slightly outperformed the strongest BARTScore-based setup on the SummEval dataset (Fabbri et al., 2021) in coherence, relevance, consistency, and fluency dimensions, while was surpassed by BARTScore on the NewsRoom dataset (Grusky et al., 2018). ROUGE-1 yielded the highest correlation scores on the sample- and dataset-level evaluations on the RealSumm dataset (Bhandari et al., 2020). The above three datasets were used for evaluating text summarization, although presenting inconsistent results. ChatGPT achieved the largest improvements on the OpenMEVA-ROC (Guan et al., 2021) story generation dataset. ChatGPT and baselines were comparable on the BAGEL dataset (Mairesse et al., 2010) for informativeness, naturalness, and quality evaluations. Liu et al. (2023c) found that GPT-4 could exceed other metrics on the dialogue generation dataset from Mehri and Eskenazi (2020). Kocmi and Federmann (2023) evaluated ChatGPT-generated machine translation accuracy scores with its former versions and other SOTA scoring systems, finding that ChatGPT was less accurate than those baselines. Gilardi et al. (2023) observed that ChatGPT achieved higher intercoder agreements than MTurk crowd-workers and educated annotators, and maintained the highest annotation accuracy in tweet frame and stance annotation tasks. The relevance annotation accuracy of ChatGPT was comparable to humans, while topic annotations were not as useful as MTurk.

2 Multilingualism

SuperCLUE benchmarks (Xu et al., 2023) compared ChatGPT-like foundation models in regard to their basic language ability, professional ability, and Chinese-featured ability. By 18th May 2023, they reported that GPT-4 and ChatGPT achieved the second (76.67) and third best (66.18) results after humans (96.50), exceeding other LLMs. GPT-4 and ChatGPT had better basic language ability than the other two metrics in Chinese. Both models achieved human-like accuracy on role-playing, chit-chatting, and coding. However, their ability to understand Chinese poetry, literature, classical Chinese, and couplets was far inferior to that of humans. Huang et al. (2023) also ranked GPT-4 and ChatGPT as top-2 on a multi-discipline (52 subjects) Chinese evaluation.

Multilingual NLP Tasks

were examined by Lai et al. (2023), e.g., multilingual part-of-speech (PoS) tagging, NER, relation classification, NLI, QA, commonsense reasoning, and summarization. The researchers analyzed 36 languages and discovered that task-specific fine-tuned models outperformed ChatGPT in the majority of examined tasks, except PoS tagging. ChatGPT exhibited superior performance in English tasks compared to tasks in other languages; for low- and extremely low-resource languages, ChatGPT performed significantly worse than baselines. Noticeably, despite the use of non-English languages in the target tasks, ChatGPT improved its performance with English prompts. Wei et al. (2023) found that direct usage of ChatGPT could not yield satisfying results in Chinese information extraction. Wang et al. (2023b) tested ChatGPT and GPT-4 on English-to-Chinese and English-to-German summarization, showing that although ChatGPT and GPT-4 exceeded other LLM baselines on a zero-shot setup, they fell behind a fine-tuned mBART-50 (Tang et al., 2021) on most of the examined datasets. Bang et al. (2023) argued that ChatGPT generally yielded weak performance on low-resource languages in language understanding and generation , while achieving higher proficiency in comprehending non-Latin scripts compared to its proficiency in generating them.

3 Reasoning

was tested by Bang et al. (2023). They found 56 out of 60 answers correct (with appropriate prompts) for deductive reasoning, i.e. applying general rules to specific situations or cases. This was stronger than other types of reasoning. 26 out of 30 were scored for abductive reasoning, i.e. forming plausible explanations or hypotheses, based on limited evidence or incomplete information. 33 out of 60 were scored for inductive reasoning, i.e., drawing generalized conclusions from examples or specific observations.

Commonsense Reasoning.

Bang et al. (2023) tested ChatGPT via three commonsense datasets, showing that 80 out of 90 of ChatGPT’s predictions were correct. ChatGPT was able to give good explanations of the reasoning steps to support its answer. However, Qin et al. (2023); Laskar et al. (2023) showed that the commonsense reasoning accuracy of ChatGPT was lower than fine-tuned baselines by a large margin. Davis (2023) found significant flaws in common benchmarks for common sense, including the CommonsenseQA dataset used by Bang et al. (2023), which he explicitly addressed. Davis (2023) listed several examples of commonsense and particularly physical reasoning failures that had been found shortly after the release of ChatGPT, and pointed to others. However, there does not exist a thorough assessment of the GPT models’ commonsense reasoning ability. More generally, Davis (2023) pointed out that “many important aspects of commonsense reasoning and commonsense knowledge are not tested in any existing benchmark”. Bubeck et al. (2023) probed a small number of their own real-world physical reasoning tasks with GPT-4, finding that it had good knowledge. They concluded that the model was able to learn an understanding of the real-world environment through the medium of text.

Causal Reasoning.

Bang et al. (2023) found that 24 out of 30 causes or effects could be correctly identified. Gao et al. (2023) systematically evaluated event causality identification, causal discovery, and causal explanation generation. Compared to SOTA models, ChatGPT and GPT-4 yielded lower scores in causality identification. They outperformed baseline models on the causal discovery, although the compared models, e.g., BERT- (Devlin et al., 2019) and RoBERTa-base Liu et al. (2019) were relatively weak. The generation of causal explanations yielded inconsistent findings in terms of AVG-BLEU and ROUGE-l metrics, while the human evaluation affirmed that both GPT models attained a level of accuracy comparable to that of human performance. Kıcıman et al. (2023) examined ChatGPT and GPT-4 on the causal discovery, counterfactual reasoning, and actual causality inferring, finding that they outperformed other LLMs and SOTA models largely on the first two tasks.

Psychological Reasoning

is the ability of humans to reason about other’s unobservable mental states (a.k.a Theory of Mind (ToM)). Kosinski (2023) and Moghaddam and Honey (2023) designed sets of False-Belief questions and quantified results suggested that both ChatGPT and GPT-4 had ToM ability, but that was still inferior to a human’s. However, Marcus and Davis (2023) pointed out flaws in the Kosinski (2023) study because the test material was in the training data. Holterman and van Deemter (2023) further tested GPT-3 and GPT-4 on more ToM tasks summarized in Kahneman (2000). They acknowledged the potential problem of the test material being in the training data, so they substituted various nouns in the scenario. However, this is unlikely to be adequate to stop a neural model from generalising from those examples. Borji (2023) found that chatGPT failed on a variant of a classic ‘Sally-Anne Test’ (used to test children). Bubeck et al. (2023) found that chatGPT answered correctly on a similar variant Sally-Anne, and they further tested on a range of more advanced ToM scenarios, with probing questions, e.g. to infer the counterfactual impact of actions on mental states. They found that GPT-4 had superior abilities and suggested that GPT-4 had a very advanced level of ToM.

Task-Oriented Reasoning.

Qin et al. (2023) evaluated dialogue, logical reasoning, complex yes/no QA, symbolic reasoning (last letter concatenation and coin flip), date understanding-, and tracking shuffled objects-oriented logical reasoning. However, ChatGPT underperformed fine-tuned baselines on most of the tasks, excluding the logical reasoning tasks. To ascertain whether ChatGPT relies on profound comprehension of truth and logic in their reasoning or merely exploits shallow memorized patterns, Wang et al. (2023a) proposed a dialectical evaluation task, finding that despite displaying high confidence, ChatGPT demonstrated an inability to hold its belief in the truth in a wide range of reasoning tasks, e.g., mathematics, first-order logic, commonsense, and generic reasoning.

Natural Language Inference (NLI)

aims to examine if a statement can be inferred, contradicted, or neutral, compared to another statement. Liu et al. (2023b) compared ChatGPT and GPT-4 with RoBERTa. However, both models encountered difficulties when dealing with novel and out-of-distribution data. They yielded relatively modest performance on NLI that needed logical reasoning. Qin et al. (2023) also proved that the NLI ability of ChatGPT was lower than that of supervised models. Ambiguity is one of the difficulties of NLI. For example, whether “John and Anna are not a couple” contradicts “John and Anna are married” depends on whether “married” means “both married” or “married to each other”. Given an ambiguous NLI premise, Liu et al. (2023a) asked ChatGPT and GPT-4 to either generate disambiguations of a premise with respect to the hypothesis or recognize disambiguation (i.e., deciding whether the disambiguation is an interpretation of an ambiguous premise). Their human evaluation showed that for the first task, GPT-4 achieved correctness at 32%, while, for the second task it was at the level of random guessing, suggesting that resolving tricky ambiguity remained challenging for ChatGPT.

Scientific Knowledge

Frieder et al. (2023) proposed a new mathematical benchmark dataset at the graduate level to test ChatGPT’s mathematical reasoning, related to textbook exercises, Olympiad problems, proof completion, algebra and probability theory problem solving, and theorem-proof and definition understanding. They found that ChatGPT only achieved a passing grade (50% of points) on 6 out of 17 testing sets. ChatGPT performed badly on mathematical problem solving, e.g., Olympiad problems, textbook exercises, algebra, and probability theory, while presenting comparatively better grades in definition understanding and proof completion. It seems ChatGPT is not good at solving mathematical problems that are weakly related to language memory or generation, because compared to the textual understanding, e.g., theorem-proofs and definitions, the other takes depend on multi-step reasoning. Qin et al. (2023) evaluated ChatGPT in arithmetic reasoning, finding that ChatGPT achieved the highest score on 3 out of 6 datasets, while fine-tuning-based methods exceeded ChatGPT on the rest of the 3 datasets. Bubeck et al. (2023) showed that GPT-4 largely exceeded the Minerva expert model (Lewkowycz et al., 2022), while ChatGPT fell behind Minerva on average.

Computer Science.

Bordt and von Luxburg (2023) used an exam with 10 different exercises. ChatGPT and GPT-4 achieved 20.5 and 24 points out of 40 full points. GPT-4 slightly exceeded the average score (23.9) of 200 students. Bordt and von Luxburg (2023) believed that passing the exam by ChatGPT should not be misconstrued as an indication of its comprehension of computer science, because numerous topics addressed in the exam were extensively documented and readily available online. The coding skills of ChatGPT and GPT-4 largely exceeded the SOTA baselines (Bubeck et al., 2023). GPT-4 even achieved human-level accuracy in LeetCode tests, while ChatGPT yielded lower accuracy, reaching only half of the average human performance. Li et al. (2023b) compared in-context learning (ICL)-based ChatGPT to SOTA models on Text-to-SQL tasks. To further improve the performance of the examined ICL models, Chain-Of-Thought (COT) and extra knowledge evidence sentences were also incorporated. ChatGPT (40.08%) exceeded the strongest baseline Codex (36.47%, Chen et al., 2021), while largely lagging behind humans (92.96%) in execution accuracy.

2 Natural Science

ChatGPT was evaluated as if it is a college student who needs to finish homework, clicker questions, programming exercises, and exams in the first-year Calculus-based Physics (Kortemeyer, 2023). Overall, ChatGPT achieved 53.05% after weighing different testing modules. This score met the minimum requirement for course credit, yet it adversely affected the overall grade-point average, falling below the necessary threshold for graduation. ChatGPT showed outstanding performance in clicker and programming questions, achieving scores higher than 90%. However, its performance in homework and exams was subpar. Additionally, ChatGPT’s mathematical difficulties in the field of physics lowered its overall score.

Chemistry.

Clark (2023) asked ChatGPT to finish two real chemistry exams with closed- and open-response questions. 44% closed-response questions were correctly answered by ChatGPT, although this is lower than the student average score (69%). Conversely, when it came to open-response questions, ChatGPT’s performance was even lower than that of the least successful student.

Medicine.

Gilson et al. (2023); Kung et al. (2023) showed that ChatGPT achieved college student level on the United States Medical Licensing Examination (USMLE). Antaki et al. (2023) found ChatGPT yielded low accuracy on neuro-ophthalmology and high accuracy on general medicine, indicating it had not grasped specialized medical knowledge well. Hirosawa et al. (2023) examined ChatGPT’s diagnosis with 30 clinical vignettes, showing that the rate of top-10 differential suggestions covering the correct diagnosis reached 93.3%, while the accuracy of top-1 suggestions was just 53.3% (physicians were 93.3%). Rao et al. (2023) suggested that ChatGPT achieved an impressive final diagnosis accuracy (76.9%), while its initial diagnosis accuracy was just 60.3%.

3 Social Science

Dahlkemper et al. (2023) aimed to investigate the extent to which students possessed accurate perception regarding the scientific accuracy and the linguistic quality of ChatGPT’s responses. They provided 102 physics students with ChatGPT answers and (masked) expert answers to evaluate the scientific accuracy and linguistic quality of both parties, perceived by the students, based on three physics questions. Despite the fact that all responses generated by ChatGPT in their study were incorrect, imperfect, or misleading, the evaluation results indicated that when confronted with a difficult question, students perceived ChatGPT’s scientific accuracy to be on par with that of the expert solution, attributing this to the higher linguistic quality of ChatGPT. However, in cases where the questions were relatively easy, both the scientific accuracy and linguistic quality of the expert’s answers surpassed that of ChatGPT.

Law.

ChatGPT was tested with four law class exams by Choi et al. (2023), including Constitutional Law, Employee Benefits, Taxation, and Torts. ChatGPT performed better on 12 essay questions than on the 95 multiple-choice questions. Although it passed all four exams, it ranked almost the lowest among law school students in each class. Liu et al. (2023b) examined ChatGPT and GPT-4 in Law School Admission Test (LSAT) and the Chinese Civil Service Examination (CCSE). Compared with RoBERTa, the advantage of GPT-4 was much larger than that of ChatGPT. GPT-4 even exceeded the average level of a human on LSAT. Nevertheless, when compared to the human ceiling, significant disparities persisted in both models.

Economics.

ChatGPT was examined with the Test of Understanding of College Economics (TUCE, Geerling et al., 2023). It demonstrated proficiency by providing accurate responses to 19/30 microeconomics questions and 26/30 macroeconomics questions. Such performance placed it in the upper echelons, ranking within the top 9% and top 1% among 3,255 and 2,789 college students who had successfully finished a full semester of microeconomics and macroeconomics. Xie et al. (2023) tested ChatGPT in stock predictions. ChatGPT showed weaker performance than traditional machine learning algorithms in numerical feature-based predictions. Through error analysis, it was determined that ChatGPT was also weak in comprehending investor sentiment expressed in text.

Ethical Considerations

asks models to be fair in terms of gender, race, language, culture, and more. Seghier (2023); Yong et al. (2023) found that ChatGPT’s responses were much worse in languages other than English (see § 2.2 ). Zhuo et al. (2023) assessed two types of social biases, namely gender bias and race bias. They concluded that, for text generation, although biases still existed, ChatGPT had mitigated them to a large extent compared to its predecessors, and that, for dialogue generation, ChatGPT could generate unbiased responses.

Robustness

requires models to maintain their performance when the inputs are different from the training data. Such inputs could be noisy data, outliers and attacks. For noisy data, Zhuo et al. (2023) tried two NLP datasets while Ye et al. (2023) tried another 12 datasets. They found that although ChatGPT defeated preceding LLMs, it was still far from perfection. Wang et al. (2023d) assessed ChatGPT on outliers, with similar results. Besides generating bad responses, the unsuccessful handling of such inputs may sometimes cause more serious consequences, for example, data leakage or Denial of Service attacks. Peng et al. (2023) assessed whether ChatGPT would suffer from SQL injection if it served as a text-to-SQL interface (Li et al., 2023a), while the answer was positive.

Reliability

requires the generated text to be faithful. Neural text generators hallucinate (Ji et al., 2023) and, therefore, produce unfaithful texts. ChatGPT/GPT-4 is no exception (Bang et al., 2023). Zhuo et al. (2023) assessed ChatGPT on fact-based question-answering datasets and concluded that it was not improved compared to its predecessors. This makes ChatGPT unreliable in tasks where faithfulness is vital. For example, it might make up references when writing scientific articles (Athaluri et al., 2023) and make up legal cases when serving in the legal domain (Deroy et al., 2023).

Toxicity

asks models not to generate harmful, offensive and pornographic content. The GPT models were designed to normally refuse to generate toxic content. Nevertheless, Derner and Batistič (2023) found that through role-playing ChatGPT still produced offensive content. A quantitative study by Zhuo et al. (2023) showed that, by feeding toxic prompting, only 0.5% of the ChatGPT responses were toxic. Nonetheless, this ability was not very robust. In line with Derner and Batistič (2023), they also found that it was susceptible to prompt injections achieved by role-playing.

Discussion

In specific NLP tasks with rich training resources, ChatGPT and GPT-4 may not perform as well as expert models (see Tables 1 and 3). There is also a certain distance from humans (see Tables 2 and 4). However, when it comes to scientific knowledge, extensive multidisciplinary testing has showcased their superiority compared to earlier models (see Table 5). Notably, in computer science and law exams, GPT-4 has achieved a level of accuracy that closely matches or even surpasses the average human performance (see Table 6). However, when we are quoted a study that shows an AI system performing at a similar level to humans, these are typically based on average performance over many cases. This average hides the fact that the AI is typically outperforming on some specialist knowledge that is difficult for humans, and underperforming on some examples that are relatively easy for humans. The same has been shown in computer vision (Russakovsky et al., 2015). Thus, human-level average accuracy achieved by AI is not equivalent to human performance and intelligence.

The GPTs may exhibit instances of hallucinations and false information (Cabrera and Neubig, 2023; Bang et al., 2023). This may be attributed to the fact that the “next word prediction”-based pre-training only taught GPTs “what is right” while neglecting “what is wrong”. To learn to generate a correct last word, e.g., “bird” after the context “if an animal has wings and can fly, it is likely a”, GPTs have to learn logic, commonsense, linguistics, and science. However, without learning “if an animal has wings and can fly, it is a penguin” is wrong, GPTs may struggle to distinguish penguins from other birds by flight ability and yield penguins’ flying hallucinations, because GPTs likely have learned “penguins are birds, having wings” from corpora. In reality, incorrect examples are far more than the negations we can see from corpora. False information may exhibit in ambiguous cases if GPTs do not know the boundary between positive and negative examples. In contrast, humans learn knowledge from both right and wrong applications to sidestep obvious fallacies (NASEM, 2018).

The inference process of GPT models is also dissimilar to humans. Humans are known to have two types of reasoning: “thinking fast” and “thinking slow” Kahneman (2000). However, GPT models seem to only have thinking fast; they do a feedforward propagation and perform relatively fast inference (Bubeck et al., 2023). Humans in contrast can spend longer deliberating for hard questions. This requires a longer process of iterative inference, not necessarily dictated by a fixed number of iterations, but instead iterating until the system achieves a satisfactory state (van Bergen and Kriegeskorte, 2020). Nevertheless, it is clear that GPT models can do a certain type of reasoning. Especially, it has been shown that adding “Let’s think step by step” to a prompt can allow GPT to use its own output as a sort of scratchpad to help it chain together multiple steps to arrive at a solution (Bubeck et al., 2023). In this way GPT is simulating “thinking slow”, however it is limited to “linear” sequences of thought and has severe limitations in tasks that require planning (ibid.); it does not backtrack to try other possible alternatives. Such backtracking would require some sort of short-term memory or workspace to remember what has been tried and what is yet to be tried. These are well-known shortcomings of neural models (Minsky, 1991).

Bubeck et al. (2023) noted the context-dependence of GPT’s mathematical knowledge; “changes in the wording of the question can alter the knowledge that the model displays.” The general phenomenon of sensitivity to input phrasing is well known from other language models also (Mao et al., 2022). Similarly, Lai et al. (2023) found a huge drop in commonsense knowledge when testing in languages other than English. This also explains why assessments by different papers can arrive at contradictory conclusions about its knowledge. It demonstrates that the internal representation in GPT suffers from severe entanglement (Bengio et al., 2013). While more extensive training can cause it to give correct responses in more contexts, it cannot change the fundamental fact that knowledge is not disentangled from the text contexts in which it appears; therefore there will always be the possibility that a change to the text of a question, that preserves the semantics, will elicit a factually incorrect response. Fixing this issue may require a hybrid system with separate facts. A commonsense-based neurosymbolic AI framework, such as the one proposed by Cambria et al. (2022) for sentiment analysis, moreover, can help increase the explainability of the reasoning processes required for decision-making, which is crucial for sensitive applications involving ethics, privacy and health.

2 Evaluation

Concerns about the validity of NLP evaluations in general (Belz et al., 2023) are applicable to this paper also. The following points are worth noting: 1) ChatGPT as a commercial product is updated periodically. On one hand, this means the “flaws” identified in early studies are fixed in later versions and different studies about the same task are not fully comparable. On the other hand, since the repair may be done by including datasets from previous assessments in fine-tuning, the data leakage may make the assessment unfair (Min et al., 2022; OpenAI, 2023a). This has been well-evidenced in Aiyappa et al. (2023) and de Wynter et al. (2023). 2) The design of prompts highly influences the results, leading to biased comparisons. Take MT as an example, by seeking better prompts, Hendy et al. (2023) and Jiao et al. (2023) had very different conclusions compared to Bang et al. (2023). In future studies, it is necessary to ensure transparency in prompt design and facilitate a fair comparison between LLMs and baselines. 3) The factors that matter in previous NLP evaluations are still valid, which may include the choice of evaluation corpora and metrics, the design of human evaluations, the task formulation and so on. For instance, Martínez (2023) argued that OpenAI’s assessment of GPT-4 (OpenAI, 2023a) on the Uniform Bar Exam is misleading because they incorrectly included test-takers who re-took the exam in their comparison. Zhang et al. (2023b) misconcluded that GPT-4 had solved all MIT math and computer science curricula because they improperly used GPT-4 as the judge. 4) The assessments of some key abilities, e.g., creativity and logical reasoning, lack either objective criteria or large-scale benchmarks. Consequently, the evaluation can only capture a portion of the overall capabilities of GPT. People seem to commonly believe that AI will cause mass unemployment (Hatzius et al., 2023). However, in these fields, the evaluation is also very limited. It would be expected to see what kind of leading results the GPT models achieved compared with humans in the field of work that can be replaced by AI.

3 Ethics

Several works have found that human perception about the reliability of ChatGPT’s output can be misled by its seemingly scientific language style (Gudibande et al., 2023; Dahlkemper et al., 2023). Given AI-generated content, it would be necessary to allow users to retrieve source references that were generated by humans to improve response reliability. Caution should be also exercised when employing AI-generated data for training new AI models, as this practice carries the risk of introducing irreparable flaws into the resultant models (Shumailov et al., 2023). Even if the Reinforcement Learning from Human Feedback (RLHF, OpenAI, 2023b) can somewhat mitigate biased, inaccurate, and toxic responses, and enhance human-centric output preferences, RLHF may be misled by human-biased feedback, e.g., system gaming (Leike et al., 2017), positive reward cycles (Ho et al., 2015), and tangled social norms and cultural context (Liu, 2023). There are also concerns regarding the non-transparency of the training set of ChatGPT, which could potentially lead to data leakage, including trade secrets or personal privacy, as user inputs might be utilized for fine-tuning (Min et al., 2022). These limitations appearing in the training and inferring processes may raise additional biases and ethical concerns.

Limitations

Several assessment papers are pre-print, and these papers may contain non-authoritative viewpoints and findings. Despite our efforts to filter out assessments with evident shortcomings, it is important to acknowledge the potential presence of unverified information in these papers. As argued in § 5.2, ChatGPT and GPT-4 are commercial products. The current assessments and the conclusion upon the assessments may be not suitable and reproducible in their later versions.

References

Appendix A ChatGPT and GPT-4 benchmark

We visualize the performance gaps between the surveyed GPT models and baselines on different tasks in Tables 1-6. The visualization depicts varying degrees of discrepancy, with darker colors indicating larger gaps and lighter colors indicating smaller gaps. The degree of discrepancy is measured by g/b−1g/b-1, where gg is the average score of a GPT model on the major metric of an evaluation task; bb is the average score of a baseline on the same major metric. The color red indicates instances where the GPT models outperformed the baselines, while blue indicates cases where the GPT models lagged behind. The teal color represents the gaps between the GPT models and the ground truth, while gray indicates a lack of comparison. Task abbreviations and meanings can be viewed in Table 7. The summary of the surveyed assessment papers can be viewed in Tables 8 and 9.