A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, Jimmy Xiangji Huang
Introduction
In recent years, the introduction of transformer-based Vaswani et al. (2017) language models (LMs), such as BERT Devlin et al. (2018), GPT (Radford et al., 2019), T5 Raffel et al. (2020), etc. have led to significant advancements in NLP Liu et al. (2019); Sanh et al. (2019); Lan et al. (2019); Lewis et al. (2020); Clark et al. (2020). The effectiveness of these models was evaluated by fine-tuning them on benchmark datasets Wang et al. (2018, 2019), achieving state-of-the-art (SOTA) performance across various tasks. Recently, large language models (LLMs) such as GPT-3 Brown et al. (2020) have demonstrated in-context-learning capability without requiring any fine-tuning on task-specific data. The impressive performance of GPT-3 and other LLMs Scao et al. (2022); Tay et al. (2022); Thoppilan et al. (2022); Fedus et al. (2021); Hoffmann et al. (2022); Zeng et al. (2022) in few-shot learning scenarios is a major finding as this helps LLMs to be more efficient, making it possible to use LM-as-a-service Sun et al. (2022) to empower a set of new real-world applications.
Intuitively, in-context learning works by learning through analogies drawn from the given demonstration examples Dong et al. (2023). After a large-scale pre-training with a self-supervision objective, LLMs can identify task-level prior patterns from the given prompt and generate a relevant continuation. Large-scale pretraining also helps them to acquire emergent capabilities like Chain of Thought Wei et al. (2022a). However, training only with self-supervision lacks grounding to real-world concepts and may not align well with its inference-time use cases resulting in unhelpful, hallucinated and sometimes toxic output Ouyang et al. (2022).
Thus, instead of learning meta-tasks in an implicit way from raw texts, recent approaches Wei et al. (2021); Sanh et al. (2021); Muennighoff et al. (2022); Chung et al. (2022); Ouyang et al. (2022) proposed learning tasks in an explicit way with a large scale prompted (supervised) meta-pretraining (a.k.a., instructional tuning) to follow instructions. In addition to that, Ouyang et al. (2022) proposed to use Proximal Policy Optimization (PPO) to fine-tune the LLM policy with human feedback in a reinforcement learning (RL) framework, introducing GPT-3.5 (text-davinci-003)https://beta.openai.com/docs/model-index-for-researchers. ChatGPT is the latest addition in this series that additionally uses dialog-based instructional data in the supervised and RL-based meta-training stages. ChatGPT has shown the ability to solve numerous tasks (e.g., question answering, text summarization, code generation, etc.) as a single model, instigating the question of “Is ChatGPT Turing complete?”.
Despite its impressive capability in performing a wide range of challenging tasks, there remain some major concernsAn ongoing crowd effort that collects its mistakes: here. about using LLMs like ChatGPT to solve real-world problems OpenAI-Blog (2022). Putting aside their high computational cost, which can be prohibitive in many practical scenarios, a primary concern is that they can fail on simple tasks involving reasoning and commonsense Marcus (2022). Second, they can perpetuate biases present in the training data, leading to unfair or prejudiced results. Another concern is their ability to be used for malicious purposes, such as generating fake or misleading text. This can be a problem when it comes to misinformation or propaganda generation that could have real-world negative impacts. While many researchers and practitioners have raised such concerns regarding ChatGPT, a systematic study evaluating ChatGPT’s performance on NLP benchmarks is still missing (as of 20 Jan, 2023, when the paper was submitted to ACL-2023 for reviewing).
In this regard, this paper aims to conduct a comprehensive evaluationTo help facilitate further research, we will release all our prompts with ChatGPT-generated responses here: https://github.com/ntunlp/ChatGPT_Eval. of ChatGPT on benchmark datasets to investigate its effectiveness and limitations in various scenarios, such as language understanding and generation capability, commonsense reasoning, open domain knowledge, the existence of new capabilities, along with studying its potential limitations, such as biases, misinformation generation, ethical concerns and etc. Meanwhile, we discover a unique capability that was not reported and analyzed for any LLMs before. We observe that ChatGPT can answer multiple arbitrary (unrelated) knowledge-based queries from a single input prompt (Sec. 4). We also report several limitations found in existing datasets while evaluating ChatGPT. In short, we conduct an extensive evaluation by analyzing 255K Chatgpt generated responses across 140 benchmark NLP datasets.
Methodology
We use several benchmark datasets and tasks for a zero-shot evaluation of ChatGPT. We categorize our evaluation into two groups: (i) Leaderboard-based Evaluation, and (ii) Task-based Evaluation. Figure 1 shows the list of all tasks that we used for evaluation in this paper. More details about the tasks and the datasets that we evaluate can be found in Appendix C, Table 15.
Evaluation:
Since ChatGPT is a conversational language model that gives human-like responses, for most of the tasks (e.g., usually discriminative classification tasks like sentiment analysis), we require human intervention to validate its responses. While for some other tasks (e.g., generative tasks like summarization or machine translation), we only use the available automatic metrics for evaluation. During the initial phase of our evaluation when the ChatGPT API was not available, a human annotator went to https://chat.openai.com/ and provided the input prompt. Afterward, the ChatGPT-generated responses were manually evaluated by at least two annotators against the gold labels. If there was a disagreement, another annotator chimed in and we considered the majority voting. When the API became available, we used the gpt-3.5-turbo model to generate the responses for different datasets. Below we describe our evaluation procedure for different types of tasks.
For discriminative tasks, after providing an input sample to ChatGPT, the generated response is compared against the gold label. Though most of the responses generated by ChatGPT are evaluated by human annotators, it was challenging to assess all generative responses solely through human annotators in scenarios when the size of the datasets was large. In such cases, we design an evaluation script for the respective dataset to first parse the results and then compare the parsed results with the gold labels. Subsequently, any samples where the script could not parse the result properly were manually reviewed by the human annotators. We denote this evaluation approach as evaluation script + human-in-the-loop (see Appendix D for details).
For generative tasks, such as summarization or machine translation where automatic evaluation metrics like ROUGE Lin (2004) or BLEU Papineni et al. (2002) are available, we solely evaluate the performance of ChatGPT using these automatic metrics instead of any human intervention.
Results and Discussion
We summarize our general observation based on our evaluation of ChatGPT in the following:
As a general purpose instruction following multitask model, ChatGPT performs worse than the SOTA single task fine-tuned models (Table 1).
ChatGPT can often perform on par with an average human in Algorithmic Tasks (Table 2).
For the same input prompt, different versions of ChatGPT may yield significantly different results (see Table 4).
Though the basic reasoning capability of ChatGPT is exceptional with Chain-of-thought (CoT) Wei et al. (2022b) prompting, ChatGPT sometimes faces severe catastrophic forgetting in newly defined reasoning tasks when CoT prompting is not used (Table 4 and Table 26).
ChatGPT can attend to multiple questions in a query and respond accordingly. However, adding many questions may reduce the model’s performance (Section 4).
Though ChatGPT has multilingual capability, its performance in underrepresented languages is very low (Table 8 and Table 24).
Though ChatGPT’s open-domain knowledge capability is extremely high (Table 6), it often suffers in several Commonsense Reasoning tasks (e.g., PIQA, SIQA, HellaSwag, WinoGrande) compared to the competing models, such as, PaLM 540B and LLaMA 65B (Table 10).
For text summarization, the ChatGPT cannot outperform the current SOTA models based on the ROGUE metric (Table 7). However, our annotators prefer ChatGPT’s generated summaries over the SOTA models (Appendix E). This suggests that we may need a new summarization metric to evaluate ChatGPT like instruction-tuned LLMs.
ChatGPT has a very strong Zero-shot mathematical (Table 11) and coding capability in comparison to other LLMs (Table 12).
ChatGPT is found to be more ethical than prior SOTA models (Table 5), while being less biased and more truthful (Table 9).
ChatGPT sometimes considers utilitarian morality and can respond to ethical dilemma-related queries (Section 3.3).
The evaluation of ChatGPT-like LLMs should include human intervention instead of fully automatic evaluation (Figure 2 and Table 16).
2 Performance based on NLP Leaderboards
In this section, we demonstrate the performance of ChatGPT in five NLP leaderboards: (i) SuperGLUE Wang et al. (2019), (ii) Big-Bench Hard Suzgun et al. (2022), (iii) Massive Multitask Language Understanding (MMLU) Hendrycks et al. , (iv) Ethics Benchmark Hendrycks et al. (2021a), and (v) Inverse Scaling Tasks Wei et al. (2022b).
We evaluate ChatGPT on the full SuperGLUE leaderboard, consisting of 10 datasets to measure an NLP model’s natural language understanding capability. We compare its performance with T5-11B Raffel et al. (2020), PaLM-520B Chowdhery et al. (2022) and PaLM 2-L Google (2023) models.
Table 1 shows the evaluation results. We observe that fine-tuned models perform exceptionally better than ChatGPT in most datasets. Meanwhile, in comparison to the 1-shot models, ChatGPT achieves competitive performance in BoolQ, CB, COPA, and WiC datasets while outperforming both models in the RTE dataset. Moreover, it outperforms the zero-shot PaLM-540B model in 5 out of 8 datasets in SuperGLUE. Though none of the models that we compared did evaluation on AX-b and AX-g datasets, we find that ChatGPT achieves 100% parity in gender bias coreference resolution in the (AX-g) dataset and a score 56.7 in terms of the Matthews Correlation Coefficient (MCC) metric in the AX-b dataset. We also find that ChatGPT obtains a very low score in the ReCoRD dataset compared to other models. Similar to GPT-3 Brown et al. (2020), we also observe quite low performance on the WiC dataset using ChatGPT.
Performance in Big-Bench Hard: We compare the performance of ChatGPT on the Big-Bench Hard benchmark with the following models: Codex Chen et al. (2021a), InstructGPT Ouyang et al. (2022); Brown et al. (2020), PaLM-540B Chowdhery et al. (2022) and PaLM-2 Google (2023). We show the overall results in Table 2 and detailed results in Table 26 in the Appendix.
We find based on the average across all tasks that ChatGPT outperforms both InstructGPT and PaLM-540B models when CoT prompts are used, while it fails to outperform these models when no-CoT, i.e., Answer-only (AO) prompts are used. In task-specific comparisons, ChatGPT outperforms both InstructGPT and PaLM-540B in the algorithmic task but fails to outperform in the NLP tasks. While ChatGPT outperforms PaLM-540B in several scenarios, it could not outperform the recently introduced PaLM 2-L model in any tasks. Though CoT prompts significantly improve the performance of ChatGPT in Big Bench Hard, we surprisingly find that even the zero-shot performance of ChatGPT outperforms its performance with few-shot AO prompts. This opens up the question for future evaluation of ChatGPT in this benchmark via tuning the AO prompts.
Performance in MMLU: We compare the performance of ChatGPT in the MMLU benchmark with models of various sizes (from 65B to 540B), as well as the PaLM 2-L Google (2023) model.
The overall evaluation results based on the average across 57 tasks can be found in Table 3. We find that the zero-shot ChatGPT outperforms all 5-shot models that are sized between 65B to 280B. Its performance (average score of 67.0) is also comparable to the 5-shot PaLM model (average score of 69.3). However, the recently released PaLM 2-L model outperforms ChatGPT by a large margin (an absolute difference of 11.3 and 14.2 from the PaLM 2-L and Flan-PaLM 2-L models, respectively). While the 3-shot ChatGPT slightly improves the performance from the zero-shot one (67.0 to 68.9), it still performs much below than the PaLM 2-L based models. While comparing the results of ChatGPT in various categories (Humanities, Social Sciences, and STEM), we find that it performs the best in the Social Science category and worst in the STEM category. We refer readers to Table 25 in the Appendix for a more detailed evaluation result per task.
Performance in Inverse Scaling Tasks: For inverse scaling Wei et al. (2022b), we evaluate the performance of two versions of ChatGPT: (i) the December 15 version in chat.openai.com and (ii) the latest API version gpt-3.5-turbo. We compare the results with the PaLM model Chowdhery et al. (2022) in the standard settings: (a) when CoT prompts are used, and (b) when not used (i.e., direct). Our results are shown in Table 4.
We observe that different versions of ChatGPT lead to different results for both CoT and no-CoT scenarios. We also find that the latest version of ChatGPT may not necessarily lead to better results. Based on the average across all 11 tasks, the December 15 version outperforms the gpt-3.5-turbo version by a score of 3.24 when CoT prompting is used, while the difference is surprisingly much higher (a difference of 24.73) when CoT prompting is not used. Thus, an in-depth evaluation of different versions of ChatGPT is important before being used in the real world. While the older version (e.g., Dec. 15) of ChatGPT outperforms the latest version in most tasks, we find that both versions are generally better than the PaLM-8B and the PaLM-62B models but usually fail to outperform the PaLM-540B model. Moreover, we find that both versions of ChatGPT obtain significantly better results when CoT prompting is used. Meanwhile, we surprisingly observe a very low performance in both versions in ÷ as digit and ÷ as digit instead sub-tasks when CoT prompts are not used. Though the score slightly improves (from 1 to 14) for the gpt-3.5-turbo model in the ÷ as digit task, it obtains a very poor score without CoT prompting in 6 out of 8 sub-tasks of Redefined Math (except Redefine e and ). Very poor performance in these tasks without CoT prompting gives a strong indication that ChatGPT is prone to give incorrect answers via memorizing the original mathematical notation from its pre-training data without properly understanding the new instructions (see Appendix J for some examples).
We find some cases in the Redefined Math task where ChatGPT gives the correct answer but provides incorrect reasoning (see Figure 2(b) for an example). Meanwhile, we observe some cases where ChatGPT gives incorrect answers even though its reasoning is correct (see Figure 2(a) for an example). We also find that the correct answer for the same input type may depend on the reasoning approach that ChatGPT is following (see Figure 3).
Performance in the Ethics Benchmark: We show the performance of the zero-shot ChatGPT model in the Ethics Benchmark in Table 5. For comparisons, we use two fine-tuned SOTA models, ALBERT-xxlarge Lan et al. (2019) and RoBERTa-large Liu et al. (2019), as demonstrated in Hendrycks et al. (2021a). We use both Test and Hard Test versions of this benchmark for evaluation in terms of the following concepts: Justice, Deontology, Virtue, Utilitarianism, and Commonsense. More details on each task are given in Appendix C.
We find based on average across all ethical concepts that ChatGPT outperforms prior SOTA models. Specifically, it significantly outperforms prior models in terms of Justice and Virtue in both Test and Hard Test versions of the dataset. More importantly, in the Hard Test, except Utilitarianism, ChatGPT significantly outperforms prior SOTA models in all other ethical concepts (though in non-Hard Tests, it fails to outperform in some concepts).
3 Performance based on NLP Tasks
Open-Domain QA: We compare the performance of ChatGPT with LLaMA Touvron et al. (2023) and PaLM-540B (both few-shot and zero-shot) Chowdhery et al. (2022) for the open-domain QA task in the following datasets (as demonstrated in Table 6): (i) TriviaQA Joshi et al. (2017), (ii) WebQuestions Berant et al. (2013), and (iii) NQ-Open Kwiatkowski et al. (2019). We find that ChatGPT not only significantly outperforms the zero-shot LLaMA-65B and PaLM-540B models, but also it outperforms the few-shot version of the PaLM-540B model. This gives a strong indication that the pre-training knowledge of ChatGPT is more extensive than LLaMA and PaLM models.
In addition, we conduct a thorough investigation and comprehensive human evaluation of ChatGPT on the EfficentQA dataset Min et al. (2021), which is also an open-domain QA dataset and derived from the NQ-Open dataset. We select EfficientQA in this regard since it is smaller than other open-domain QA datasets we used for evaluation. Based on our extensive analysis, we observe several key insights in the EfficientQA dataset. For instance, many questions in this dataset are time-sensitive, while many answers contain outdated gold answers. Additionally, as ChatGPT was trained in 2021, it fails to answer questions that require knowledge of recent events. Moreover, we find some examples where ChatGPT gives a correct answer but the gold answer in the dataset is outdated. Though we observe an accuracy of 68% by ChatGPT in the EfficientQA dataset, fixing these outdated answers with the correct answers increases the accuracy to 71.1%. We show a few responses of ChatGPT in the EfficientQA dataset demonstrating some of the above findings in Appendix G.
Reading Comprehension: We compare the performance of ChatGPT with the LLaMA 65B model (zero-shot) and the PaLM-540B model (few-shot and zero-shot) for the reading comprehension task as demonstrated in Table 6. We find that in terms of accuracy, ChatGPT outperforms both few-shot and zero-shot PaLM-540B models as well as the LLaMA-65B (zero-shot) model in the RACE dataset (both Middle and Hard versions) Lai et al. (2017). While in the SQuAD 2.0 dataset Rajpurkar et al. (2018), based on the Exact Match (EM) metric, it fails to outperform the PaLM models.
Commonsense Reasoning: For the commonsense reasoning capability evaluation, we also compare ChatGPT with the zero-shot LLaMA-65B model and the PaLM-540B model (few-shot and zero-shot). While we find from Table 10 that ChatGPT outperforms all other models in the SIQA Sap et al. (2019), ARC easy (ARC-e) and ARC challenge (ARC-c) Clark et al. (2018), and OBQA Mihaylov et al. (2018) datasets, it obtains significantly lower scores in the PIQA Bisk et al. (2020), HellaSwag Zellers et al. (2019), and WinoGrande Sakaguchi et al. (2020) datasets.
Mathematical Reasoning: We find from Table 11 that ChatGPT shows strong mathematical performance on all datasets, outperforming all prior models (Minerva-540B Lewkowycz et al. , PaLM-540B Chowdhery et al. (2022), and LLAMA Touvron et al. (2023)) on the MATH dataset Hendrycks et al. (2021b), as well as the GSM-8K Cobbe et al. (2021), and Multilingual Grade School Math (MGSM) Shi et al. (2022) datasets.
Natural Language Inference (NLI): We find from Table 6 that ChatGPT outperforms both few-shot and zero-shot PaLM-540B model Chowdhery et al. (2022) in the Adversarial NLI (ANLI) Nie et al. (2020) benchmark datasets for the NLI task.
Text Summarization: For text summarization, we use the current SOTA models to compare the performance with ChatGPT as results for LLMs like PaLM-540B and LLaMA-65B are not available for the summarization task. We use the following datasets for evaluation: CNN-DM See et al. (2017); Hermann et al. (2015) and XSUM Narayan et al. (2018) for news article summarization, while the DialogSUM Chen et al. (2021b) and SAMSum Gliwa et al. (2019) datasets for dialogue summarization. For these datasets, we evaluate ChatGPT using (i) Restricted Prompting: Writing a summary in not more than X words, and (ii) Unrestricted Prompting: Writing a summary without any word-limit restrictions in the summary.
We show our results in Table 7. We find that except CNN/DM, ChatGPT achieves much better performance when restricted prompts have been used. This could be due to the fact that the average gold summaries in XSUM, SAMSum, and DialogSum datasets are quite smaller and so the restricted prompting helps improve the ROUGE score. However, we find that ChatGPT does not necessarily properly follow the restrictions in words (exceeding the word restriction 73.5% times on average) when it generates its responses (Appendix F for more details). In comparison to the SOTA models, we find that the ROUGE scores of the zero-shot ChatGPT model are much lower than the SOTA results. We further randomly collected 100 samples (50 for XSUM and 50 for CNN/DM) to conduct a human evaluation of the summaries generated by ChatGPT and Ravaut et al. (2022) (see Appendix E for more details). We find that our annotators prefer ChatGPT 78% times in CNN/DM and 92% times in XSUM. This is consistent with the recent findings Liu et al. (2023d); Goyal et al. (2022), where summaries from GPT-3.5 are preferred compared to fine-tuned models in reference-free evaluation.
Machine Translation: We evaluate ChatGPT for the machine translation task in various languages (English (en), French (fr), German (de), Romanian (rn), Kazakh (kk)) under various scenarios. Similar to Chowdhery et al. (2022), for English-centric language pairs, we use the WMT’14 Bojar et al. (2014) for English-French translation in high-resource scenarios, WMT’16 Bojar et al. (2016) English-German in medium-resource while English-Romanian for low-resource scenarios; WMT’19 Barrault et al. (2019) for direct translation between non-English languages: German-French and for extremely low-resource language pairs: English-Kazakh. We find that while translating from English to other languages, ChatGPT outperforms the zero-shot PaLM model. Whereas, the opposite happens when the translation is done from other languages to English. Moreover, for non-English translation (between German and French), we observe that ChatGPT even outperforms the SOTA fine-tuned models. Nonetheless, in other datasets, ChatGPT could not outperform the fine-tuned SOTA models.
Code Generation: We evaluate the coding ability of ChatGPT on the MBPP Austin et al. (2021) and the HumanEval Chen et al. (2021a) datasets. Based on our results shown in Table 12, we find that in terms of the pass@1 metric, ChatGPT outperforms all models in the HumanEval dataset. While ChatGPT obtains a score of 73.8 in the MBPP dataset in terms of pass@1, it outperforms the 3-shot LLaMA in that dataset while also achieving performance comparable to the fine-tuned and 3-shot PaLM-540B models in the same dataset.
Bias and Misinformation: For bias evaluation, we use the WinoBias Zhao et al. (2018) dataset to evaluate the performance on both Type 1 and Type 2 versions of the data for the co-reference resolution task in pro-stereotype and anti-stereotype scenarios. The bias in this dataset is computed via measuring the difference between these two scenarios. For misinformation generation evaluation, we use the TruthfulQA Lin et al. (2022) dataset.
Based on our experimental results in these datasets in Table 9, we find that in the WinoBias dataset, ChatGPT obtains impressive performance on the Type 2 version of the dataset (100% accuracy in pro-stereotype and almost 100% in anti-stereotype scenarios), with a very low difference (0.51%) between these two types. However, in the Type 1 version of the dataset, there is a high bias in ChatGPT response, as the difference between the accuracy of pro-stereotype (96.97%) and anti-stereotype (80.30%) is about 16.67%. Thus, asking ChatGPT to answer based on world knowledge without any syntactic cues in the Type 1 task (contrary to the Type 2 task that can be resolved using syntactic information), leads to more bias. In the TruthfulQA dataset, we find that in terms of truthfulness and informativeness, it obtains a score of 0.78 and 0.70, respectively (in comparison, the LLaMA 65B model Touvron et al. (2023) achieves a score of 0.57 and 0.53, respectively).
Ethical Dilemma: We generate the ChatGPT responses for a set of 25 manually constructed questions that integrate racial, political, social, and religious biases as well as abstract decision problems. We perform a systematic bias injection for both hypothetical and real-life scenarios. Response to each question is generated three times for a rigorous evaluation. While we do not evaluate whether the ChatGPT-generated responses for the given questions are right or wrong, we will release all responses generated by ChatGPT for readers’ discretion (see Appendix H for some ChatGPT-generated responses). By analyzing the responses, we observe that ChatGPT can identify the Trolley Problem. We also observe that most of the time ChatGPT remains neutral and provides expert-like opinions putting arguments for all possible scenarios.
Other Tasks (Sentiment Analysis & NER): In the IMDB dataset Maas et al. (2011) , we obtain 92.3% accuracy for sentiment analysis. For NER (Named Entity Recognition), we use the WNUT 17 Derczynski et al. (2017) dataset to obtain Precision: 18.03, Recall: 56.16, and F1: 27.03.
PolyQuery Synthesis
In this section, we present a unique capability of ChatGPT that we discover in the course of our study. Specifically, it can identify multiple queries (potentially for different objectives) in a single prompt and retrieve responses for all these queries from the latent representation of the model. Retrieving a set of arbitrary information in this way makes it an impressive feature, paving the way to use the ChatGPT API in real-world limited-budget scenarios by solving multiple tasks at once based on a single input prompt. To our best knowledge, no prior work investigated this feature of LLMs. We name this capability as PolyQuery Synthesis.
To do a systematic evaluation, we create a small dataset from the EfficientQA dev split Min et al. (2021) and Web-Questions Berant et al. (2013) test split. For each dataset, we combine 5 different samples into a single sample and create a prompted and non-prompted (non-instructional) input. In total, we use 100 samples from each dataset for evaluation. We also show an example in Figure 4.
We generate responses for 13 different models from OpenAIhttps://beta.openai.com/docs/models/overview; see Table 13 for the result. We observe that ChatGPT shows strong performance on both prompted and non-prompted queries. While davinci-003 and davinci-002 perform reasonably in prompted queries, their performance is much worse in non-prompted queries. We did not observe this in the original davinci model. Based on the performance variations in different models, we suspect that instructional tuning (both supervised and RL) enables this emergent feature in ChatGPT and davinci-{001,002,003} series. An example of responses from all the models can be found in the Appendix in Table 21 and Table 22. We also compare the result with single sample input and observe that PolyQuery Synthesis usually leads to some drop in performance.
Conclusions and Future Work
This paper evaluates the effectiveness and limitations of ChatGPT in standard academic datasets. To our best knowledge, this is the first work that conducts an extensive evaluation of ChatGPT in benchmark NLP datasets. We observe that even though ChatGPT obtains impressive zero-shot performance across various tasks, it is still far from reaching human-level performance in many tasks. Moreover, potential biases and ethical concerns, as well as misinformation generation risks of ChatGPT are discussed. In addition, a unique capability of ChatGPT has been studied. Though there may have numerous other capabilities of ChatGPT that go unnoticed in this paper, future work should nonetheless investigate the capability of ChatGPT on more tasks. We will make all our prompts and ChatGPT-generated responses publicly available.
Limitations
Even though there has been a lot of hype on social media regarding various application areas of ChatGPT, there may have other capabilities of ChatGPT that are not investigated in this paper. Since the instruction-tuning datasets of OpenAI models are unknown (not open-source), some datasets used for evaluation may or may not exist in the instruction-tuning training data of OpenAI. Another limitation of this research is that most of the numerical value of the results may change as OpenAI trains new models with more data and filters. While the experimental results may change over time, this work will still give a concrete direction on what to expect from a general-purpose dialogue model and potential shortcomings.
We also want to add a disclaimer in the result comparison between different models. In this research, we were only able to generate textual responses from the ChatGPT model. That means we did not have access to the log-probability of the model. Thus the model was only evaluated on generative responses. At the time of the research performed, we did not do any log-probability ranking-based evaluation due to the limitations of the ChatGPT API. We also strongly believe that the evaluation of a ChatModel should be generative instead of ranking accuracy. While doing our literature review and collecting results from different LLM papers (i.e., Google (2023); Touvron et al. (2023); OpenAI (2023)) we often did not find details about their evaluation approach, reference evaluation script, or even prompts used for the task. To alleviate this issue, we did rigorous prompt testing on ChatGPT before the evaluation of each task. We tried our best to make sure that ChatGPT responds to answer choices instead of generating open-ended text. While we are quite confident about our evaluation (due to human evaluation), we want to worry that the compared models mentioned in this paper may not always generate suitable targeted words from the answer choices while generating text. However, we included all the potential LLM baselines in this paper because it depicts a reasonable comparison. Since many different institutes are not releasing research details (i.e., checkpoint, model details, evaluation script), we believe that adding these relevant numbers to the table will help see the model in a comparative manner. For chatbot evaluation, we sincerely want to invite the community to adopt the generative evaluation since it depicts a real-life scenario and human-centric interaction with the model.
While this paper evaluates ChatGPT across 140 datasets, there remain many other tasks that are not evaluated in this paper. For instance, tasks in the Biomedical and the Clinical domain Luo et al. (2022); Lee et al. (2020); Alsentzer et al. (2019); Beltagy et al. (2019); Gu et al. (2020); Peng et al. (2019), NER across more datasets Tjong Kim Sang and De Meulder (2003); Malmasi et al. (2022); Fu et al. (2022); Laskar et al. (2022a), Multi-Document and Query-Focused Text Summarization Laskar et al. (2020a); Zhong et al. (2021); Su et al. (2022); Laskar et al. (2022d), Low-Resourced Hedderich et al. (2021) NLP problems, Data-to-Text Generation Kantharaj et al. (2022); Rahman et al. (2023), Entity Linking Wu et al. (2020); Ayoola et al. (2022); Laskar et al. (2022b, c), Answer Re-Ranking Task Garg et al. (2020); Laskar et al. (2020b), etc.
While our study may open up new ideas and thought-provoking arguments on the evaluation of Chat-based models, we want to acknowledge that the breadth of such evaluation is extremely limited at this moment. However, we believe that this evaluation effort will generate new research questions and priorities Red Teaming LLMs.
Ethics Statement
The paper does not leverage any 3rd-party to conduct the human evaluation of the ChatGPT responses and so no additional compensation was needed. All the human evaluations in this paper are conducted by the authors. Since this paper only evaluates the performance of ChatGPT and investigates its effectiveness and limitations, conducting the human evaluation by the authors does not lead to any unwanted biases or ethical concerns. Only the publicly available academic datasets are used that did not require any licensing. Thus, no personally identifiable information has been used while evaluating ChatGPT responses.
Acknowledgements
We would like to thank all the anonymous reviewers for their excellent review comments. This work was supported by the Natural Sciences and Engineering Research Council (NSERC) of Canada and the York Research Chairs (YRC) program. Jimmy Huang (jhuang@yorku.ca) and Shafiq Joty (srjoty@ntu.edu.sg) are the contact authors of this paper.
References
Appendix A Frequently Asked Questions (FAQ)
ChatGPT represents a generational leap in terms of the multi-task capability of machine learning models. It surpasses (or promises to surpass) most of the potential AGI testshttps://en.wikipedia.org/wiki/Artificial_general_intelligence#Tests_for_testing_human-level_AGI defined earlier (though some of them are defined jokingly). The technical details and model weights are kept hidden citing security and competitiveness OpenAI (2023) of the current market. While these reasons are highly debatable in the research community, there is no doubt that such systems will be reproduced in the near future. Evaluation serves as a valuable means to estimate and address various research questions regarding model size, data size, and more. For instance, we refer to this blog posthttps://blog.eleuther.ai/gpt3-model-sizes/ which attempts to estimate the size of the language model based on evaluation results from the API-served model. Moreover, it is important to emphasize that Evaluation of Generative Texts serves as a form of interpretability, empowering researchers and downstream users to understand the capabilities, biases, and tendencies of the models. Evaluating such potential models often leads to the exploration of emergent capabilities, helping researchers bridge the gap between smaller and larger models (often with data augmentation), or, at the very least, gaining insights into what can be expected at different scales. This, in turn, aids in making informed decisions regarding model training and serving specific use cases.
Which version of ChatGPT was used for this paper?
Our initial evaluation was performed manually on the website chat.openai.com. Once the API became available from OpenAI, we utilized the gpt-3.5-turbo API to generate responses for our prompted samples. However, we show the API version for all the evaluated datasets in Table 15.
Why did we conduct a zero-shot evaluation?
Though the consensus from the GPT-3 paper Brown et al. (2020) is to evaluate LLMs in a few-shot manner with in-context evaluation, the basic expectation of the community is always to interact with an LLM in a single-shot question. Since the release of T0++ Sanh et al. (2021) and the FLAN model Wei et al. (2021), we have seen that instruction tuning has enabled LLMs to perform zero-shot evaluation better than non-instruction-tuned models. Presumably, ChatGPT, being a larger instruction-tuned model trained on an extremely large dataset, makes it an appealing test subject to evaluate and understand what to expect from an instruction-tuned model.
In addition, since the Evaluation of Generative Texts of large language models is complex and may require manual evaluation of each sample, some of the prior works often report one-shot results instead of zero-shot to automate the evaluation process by providing a response pattern to the LLM. However, we believe that conducting a zero-shot evaluation would greatly benefit the current research field and provide insights into the model’s real-world performance. While the main purpose of this paper is to conduct a zero-shot evaluation of ChatGPT, some prior research prioritize the performance in terms of few-shot scenarios depending on various tasks. Thus, we also include the few-shot performance of ChatGPT in a few places so that we can have a better comparison.
Why did we evaluate ChatGPT on prompted samples instead of dialogue datasets?
The main training novelty of ChatGPT comes from Proximal Policy Optimization (PPO) based prompted sample fine-tuning while leveraging human in the loop. The training of supervised policy in Ouyang et al. (2022) is similar to the prompted sample training method mentioned in Sanh et al. (2021); Wei et al. (2021). Since the training data is prompted samples of different NLP tasks, we decided to evaluate it in challenging instruction-based prompted datasets collected from various NLP benchmarks. However, we acknowledge that the evaluation of multi-hop dialogue datasets is also important but not covered in this work. We keep it as a future work. For clarity & managing the expectations of the readers, we add benchmark datasets in the title of the paper.
How was the ethical dilemma dataset created? Why do you evaluate ChatGPT on the Trolley problem?
The impressive performance of ChatGPT may potentially lead to applying it in AI agents like autonomous cars, and robots, or in exploratory research. This is called the Agentic behavior of large LLMs. Though trolley problem is a thought experiment, it depicts some fundamental decision problems which can indicate the roots of many derivative biases. Because of this, we decide to evaluate it in the trolley problem.
A set of 25 questions is created by one of our authors inspired by Michael Sandel’s lecture, The Moral Side of Murder Sandel (2019). The questionnaires mainly evaluate moral dilemmas. In addition to that, We tried to explain the importance of the trolley problem in the FAQ section. All of our ethical questions (not restricted to only the trolley problems) and ChatGPT responses are added to the repository folder. Evaluation of the “moral dilemma” is quite a complicated task and may differ in different parts of the world. So we didn’t ask the question “If the answer to the certain ethics question is acceptable or not” rather we commented on patterns (i.e., ChatGPT provides expert-like opinions putting arguments for all possible scenarios) and attached all the responses in Supplementary. We believe that a few systematic thought-provoking questionnaires may introduce many new seeds of ethical evaluation datasets.
We found this unique capability while working on the EfficientQA dataset (an ODQA dataset). To make sure that the emergent capability is not dataset dependent, later we add another additional open-domain QA dataset (Web-Question). We observe that most of the time similar capabilities can be also found in other prompted datasets (e.g., WiC, COPA, etc.). However, their mixing of multiple samples results in a prompted sample that sounds and reads very artificial. Because of this reason, we only evaluate ODQA datasets where both prompted and non-prompted samples sound and read like a natural form of subsequent queries.
Why non-CoT results in many Inverse Scaling tasks are extremely low?
Though ChatGPT achieves good performance on all datasets in the Inverse Scaling benchmark when CoT prompts have been used, it surprisingly performed very poorly in many tasks, especially in Redefine Math sub-tasks when CoT prompts are not used. We hypothesize that ChatGPT is prone to hallucination, and tends to answer based on memorization of the original task learned during its pre-training stage, instead of answering with proper reasoning when no step-by-step instruction to solve a new task is provided. However, a sharp reduction in performance is still an interesting finding and may require more information on the datasets used for training text-davinci-003 and ChatGPT to find the root cause of it.
What is the citation Strategy in tables?
While adding results to various tables, our objective was to provide insight into potential competing models or results that directly signify some strong observations. We acknowledge here that the paper is missing results on several effective smaller models, such as GPT-J (Wang, 2021), GPT-NeoX (Black et al., 2022), T5 (Raffel et al., 2020), T0 Sanh et al. (2021), FLAN-T5 (Chung et al., 2022). We also had to consider page restrictions for the ACL version of the paper. However, feel free to email us with more insightful results for your favorite model, and we will do our best to cite those results in our arXiv version.
Why did we use the Dev Set instead of the Test Set for some datasets?
Many of the datasets that we used for evaluation had a test split for which the gold labels are not publicly available. Meanwhile, as ChatGPT provides generative responses, for most datasets we require human intervention to compare the ChatGPT generated responses against the gold labels. For this reason, for the datasets that do not have a test split containing gold labels publicly available, we report the results on the development split similar to the recent literature Sanh et al. (2021); Chowdhery et al. (2022); Rae et al. (2021); Du et al. (2022); Touvron et al. (2023).
Appendix B Literature Review
The impressive success of pre-trained language models (Radford et al., 2019; Devlin et al., 2018; Liu et al., 2019; Raffel et al., 2020; Brown et al., 2020; Liu et al., 2023c; Zhou et al., 2023) has led to the development of several conversational language models, including, Meena (Adiwardana et al., 2020), LaMDA (Thoppilan et al., 2022), DialoGPT (Zhang et al., 2020), etc. These models are pre-trained on a huge amount of raw data Raffel et al. (2020); Ortiz Suárez et al. (2019); Gao et al. (2020) crawledhttps://commoncrawl.org/ from the web to obtain state-of-the-art performance via task-specific fine-tuning Devlin et al. (2018); Pfeiffer et al. (2021); Li and Liang (2021); Hu et al. (2021); Lester et al. (2021) on various benchmark datasets Wang et al. (2018, 2019); Hu et al. (2020); Liang et al. (2020); Lu et al. (2021).
ChatGPT is also a large conversational language model. It leverages the in-context learning method that works by learning through analogies drawn from the given demonstration examples Dong et al. (2023). After a large-scale pre-training with a self-supervision objective, in-context learning helps LLMs to identify task-level prior patterns, while acquiring emergent capabilities like Chain of Thought Wei et al. (2022a). However, training only with self-supervision lacks grounding in real-world concepts that may lead to hallucination and toxic output generation Ouyang et al. (2022). Thus, instead of learning meta-tasks in an implicit way from raw texts, recently Wei et al. (2021); Sanh et al. (2021); Muennighoff et al. (2022); Chung et al. (2022); Ouyang et al. (2022) proposed learning tasks in an explicit way with a large scale prompted (supervised) meta-pretraining (a.k.a., instructional tuning) to follow instructions. In addition to that, Ouyang et al. (2022) proposed to use Proximal Policy Optimization (PPO) to fine-tune the LLM policy with human feedback in a reinforcement learning (RL) framework, introducing GPT-3.5 text-davinci-003https://platform.openai.com/docs/models. ChatGPT is the latest addition in this series that additionally uses dialog-based instructional data in the supervised and RL-based meta-training stages.
Dialogue Evaluation:
For dialog-based evaluation, Liu et al. (2016) investigated evaluation metrics for dialogue response generation and showed that BLUE-based automatic metric doesn’t correlate well. Lowe et al. (2017) propose an evaluation model ADEM that learns to predict human-like scores to input responses. Using the optimal error rate in determining whether a phrase is human or machine-generated, Hashimoto et al. (2019) provides HUSE, a unified framework that assesses variety and quality. Finally, Adiwardana et al. (2020) introduced a Mini-Turing Benchmark (MTB) which is a collection of 1,477 conversational contexts.
Instruction Datasets:
In recent years, Mishra et al. (2021) constructed a natural instruction dataset via crowdsourcing 61 instructions of 6 task types. Wei et al. (2021) introduce prompting techniques that transform regular tasks into human instructions on 62 text datasets with 620 instructions. Later, Bach et al. (2022)https://github.com/bigscience-workshop/promptsource scales up everything to 176 datasets and 2052 instructions. Both of the benchmarks were proposed for around 12-13 task types. Finally, Wang et al. (2022)https://github.com/allenai/natural-instructions scales up the task type to 76 and proposes around 1616 tasks with 1616 instructions. In contrast to this, Ouyang et al. (2022) annotated 14378 instructions of 10 task types and achieved impressive performance with LLMs via following instructions. To our best knowledge, ChatGPT is also trained based on a similar instruction-based data pipeline but not open-sourced https://openai.com/blog/chatgpt/. Following this, we evaluate ChatGPT on publicly available prompted datasets while creating new datasets when needed.
ChatGPT Evaluation:
Recently few concurrent works have attempted to evaluate ChatGPT on many different tasks based on different benchmarks and tasks. Table 14 shows a brief literature review on the ChatGPT evaluation effort.
Appendix C Task & Dataset Description
We evaluate ChatGPT on the SuperGLUE Wang et al. (2019) benchmark, which is a widely used leaderboard to evaluate the language understanding performance of NLP models.
Big-Bench Hard:
We evaluate ChatGPT on 23 hard tasks Suzgun et al. (2022) of the Beyond the Imitation Game benchmark (BIG-bench) Srivastava et al. (2022). It is a challenging benchmark that is used to evaluate the capability of LLMs.
Massive Multitask Language Understanding:
We evaluate ChatGPT on the Massive Multitask Language Understanding (MMLU) Hendrycks et al. benchmark. It is a multiple choice Question Answering (QA) benchmark, consisting of 57 different tasks, covering topics in humanities, science, technology, engineering, mathematics, etc.
Inverse Scaling Challenge:
We use all four tasks (Hindsight Neglect, Quote Repetition, Negation QA, and Redefined Math) from the Inverse Scaling Perez and McKenzie ; Wei et al. (2022b) challenge. There are a total of 11 tasks from 4 main categories.
Hindsight Neglect: This task assesses whether a bet is worth taking based on its expected value.
Quote Repetition: This task contains a sequence of a famous quote where the objective is to assess whether an altered ending of this famous quote can confuse the model into finishing the sequence with the well-known ending rather than the expected ending given in the prompt.
Negation QA: This task negates a part of a question in an existing multiple-choice dataset to see if language models are properly following instructions in the prompt or if they are sensitive to negation.
Redefine Math: This task aims to evaluate if language models can still perform proper reasoning when the mathematical symbols are redefined to mean something else. It has 8 sub-tasks.
Ethics Evaluation Benchmark:
We use the Ethics Benchmark dataset Hendrycks et al. (2021a) to assess ChatGPT in terms of basic concepts of morality and ethical judgments. This dataset covers concepts of justice, well-being, duties, virtues, and commonsense. This dataset has two test sets (Test and Hard Test). We use both versions of the test sets and evaluate ChatGPT in the following 5 categories: (i) Justice, (ii) Deontology, (iii) Virtue, (iv) Utilitarianism, and (v) Commonsense.
C.2 Task-based Evaluation
To investigate the open domain knowledge of ChatGPT, we evaluate its performance on the TriviaQA datasetJoshi et al. (2017), the NQ-Open dataset Kwiatkowski et al. (2019) and the WebQuestions Berant et al. (2013) dataset. In these datasets, the task is to answer a question asked in English by leveraging the contents of Wikipedia or the Web. Moreover, we also conduct a comprehensive human evaluation on the EfficientQA dataset Min et al. (2021), which is also derived from the NQ-Open dataset. Based on our extensive analysis, we observe several key findings in the EfficientQA dataset, such as many questions are time-sensitive, while many answers contain outdated gold answers.
Reading Comprehension:
We use the RACE dataset (both Middle and Hard versions) Lai et al. (2017) to evaluate ChatGPT for the reading comprehension task. The Race dataset is constructed from English reading comprehension exams designed for middle and high school students in China. In addition, we use the SQuAD 2.0 dataset Rajpurkar et al. (2018) for this task.
Commonsense Reasoning:
To evaluate the reasoning capability of ChatGPT, we use the following datasets: PIQA Bisk et al. (2020), SIQA Sap et al. (2019), HellaSwag Zellers et al. (2019), WinoGrande Sakaguchi et al. (2020), ARC easy and challenge Clark et al. (2018), and OBQA Mihaylov et al. (2018). Tasks in these datasets include Cloze and Winograd style challenges, multiple choice QA, etc.
Mathematical Reasoning:
We evaluate the mathematical reasoning capability of ChatGPT on the MATH dataset Hendrycks et al. (2021b) and the GSM-8K dataset Cobbe et al. (2021). In addition, we use the recently proposed Multilingual Grade School Math (MGSM) Shi et al. (2022) dataset to evaluate its mathematical capability in multilingual settings.
Natural Language Inference:
To evaluate the Natural Language Inference (NLI) capability of ChatGPT, we use the Adversarial NLI (ANLI) Nie et al. (2020) benchmark datasets.
Text Summarization:
We use various datasets to evaluate the text summarization performance of ChatGPT. The datasets we used are CNN-DM See et al. (2017); Hermann et al. (2015) and XSUM Narayan et al. (2018) to summarize articles in the news domain, while the DialogSUM Chen et al. (2021b) and SAMSum Gliwa et al. (2019) datasets for dialogue summarization.
Neural Machine Translation:
We select various languages (English (en), French (fr), German (de), Romanian (rn), Kazakh (kk)) based on different scenarios to evaluate the performance of ChatGPT in language translation. Similar to Chowdhery et al. (2022), for English-centric language pairs, we use the WMT’14 Bojar et al. (2014) for English-French translation in high-resource scenarios, WMT’16 Bojar et al. (2016) English-German in medium-resource while English-Romanian for low-resource scenarios; WMT’19 Barrault et al. (2019) for direct translation between non-English languages: German-French and for extremely low-resource language pairs: English-Kazakh.
Code Generation:
We evaluate the coding ability of ChatGPT on the MBPP Austin et al. (2021) and the HumanEval Chen et al. (2021a) datasets.
Bias and Misinformation:
To investigate whether ChatGPT has any potential biases, we evaluate its performance the WinoBias dataset Zhao et al. (2018). In WinoBias, we use both Type 1 and Type 2 versions of the datasets. The Type 1 version of the data requires the co-reference decisions to be made using the world knowledge of the model based on the given circumstances, whereas the syntactic information and proper understanding of the pronoun in the given input are enough to answer the Type 2 version of the data.
We evaluate ChatGPT in terms of misinformation generation on the TruthfulQA dataset Lin et al. (2022).
Ethical Dillemma:
A potential use of ChatGPT-like models (e.g., text-davinci-003 series models) can be to integrate them into the decision-making process of other AI agents (i.e., autonomous industry, exploratory research). For the fundamental decision-making process, geographical, cultural, and/or racial differences may play a role in some ethical and psychological dilemmas, that may vary from person to person. While it is easily possible to fool a dialogue system with complex multimodal queries, in this work we take a different approach to evaluate ChatGPT on decision problems. We evaluate the well-known Trolley Problem Thomson (2020), which is a series of thought experiments to identify decision patterns in problems related to ethics and philosophy. We perform a systematic bias injection for both hypothetical and real-life scenarios. Response to each of the questions is generated three times for a rigorous evaluation.
Sentiment Analysis:
We use the IMDB Movie Review dataset Maas et al. (2011) for the binary sentiment classification task.
Named Entity Recognition (NER):
For NER, we use the WNUT 17 Derczynski et al. (2017) dataset.
Appendix D Importance of Evaluating with Human in the Loop
Due to ChatGPT being a generative model, it is difficult to directly compare many of the ChatGPT-generated responses against the gold labels, especially in discriminative tasks, for performance evaluation. For this reason, in many datasets, we require human intervention to evaluate the ChatGPT responses. In some of these discriminative datasets, we directly evaluate the performance via humans. While in some others, we evaluate ChatGPT using an evaluation script written by us that first checks whether the generated response is correct or not (via lexical or fuzzy word matching). Afterward, we select some responses for human evaluation that could not be evaluated by our evaluation script. We denote this process as Evaluation Script + Human in the Loop. In Table 16, we demonstrate the importance of this technique for evaluation by comparing the score achieved directly by the evaluation script vs the score achieved directly by the evaluation script + Human in the Lopp.
We find that based on the average across all tasks for both Test and Hard Test versions, the average difference in performance is 3.0 in the Ethics Benchmark. While in the Big-Bench Hard and the MMLU benchmarks, the average difference is 0.8 and 0.3, respectively. For Reading Comprehension, we did not notice any difference in Race datasets, while we observe a difference of 7.0 for SQuAD-V2. Moreover, we notice a high difference in the Open-Domain QA datasets, as in the NQ-Open and the WebQuestion datasets, the differences are 6.6 and 10.9, respectively. The average difference in the Open-Domain QA datasets (NQ-Open, WebQuestions, TriviaQA) is 6.6. While in Commonsense Reasoning, the average difference is 1.1. Moreover, our Evaluation Script was perfect in the NLI datasets, while nearly perfect (with a small difference of 0.4) for Sentiment Analysis in the IMDB dataset.
It is quite clear from our analysis that in some datasets (e.g., NQ-Open, WebQuestions, PIQA, etc.), human involvement has made a great difference in results. While in some datasets, it was possible to get accurate results with just our evaluation script (e.g., ANLI datasets). It should be noted that when we designed our input prompts for ChatGPT, we added the following in our prompts for some datasets: Answer without any explanation. This is done such that the response generated by ChatGPT can be easily parsed and evaluated using our evaluation script.
Appendix E Human Evaluation of ChatGPT-generated summaries
We randomly collected 100 samples (50 for CNN/DM and 50 for XSUM) to conduct a human evaluation of the summaries generated by ChatGPT and the SummaReranker model from Ravaut et al. (2022). Two human annotators who were unaware of the source of the summaries (whether generated by ChatGPT or by the SummaReranker model) were asked to select their preferred summary. The annotation task was designed as follows: they were provided with the input document, followed by the summaries generated by ChatGPT and the SummaReranker model. To ensure a fair evaluation by avoiding any unintentional biases, the summaries of these models are shown to the annotators in a random order: sometimes the summary generated by ChatGPT is shown at first, followed by the summary generated by the SummaReranker model; or vice versa. While selecting one summary over another, the annotators were encouraged to choose based on the following criteria: factual correctness, informativeness, coherence, and fluency.
We find that our annotators prefer ChatGPT-generated summaries 92% times in XSUM and 78% times in CNN/DM. This suggests the need for a new evaluation metric to evaluate LLM-generated summaries.
Appendix F Analyzing the effect of Restricted Prompts for Text Summarization
We prompted ChatGPT to generate summaries in two scenarios: (i) Restricted Prompting: Writing a summary in not more than X words, and (ii) Unrestricted Prompting: Writing a summary without any word-limit restrictions in the summary.
In Table 17, we find that ChatGPT-generated responses are on average quite longer than the average length of gold summaries. However, restricted prompting indeed helps ChatGPT to generate smaller summaries. More specifically, it reduces the average length for CNN/DM, XSUM, SAMSum, and DialogSUM by 7.2, 18.5, 17.4, and 27.9, respectively, in comparison to unrestricted prompting. However, even using restricted prompting, on average, the generated summaries are longer by about 22 words in CNN/DM and 32 words in XSUM (in comparison to the word length restriction mentioned in our prompts). Meanwhile, we observe that this difference is quite low (not more than 4 words on average) in SAMSum and DialogSum. Thus, ChatGPT following instructions related to word limit restrictions in summarization datasets may vary across datasets. We further investigate how often ChatGPT exceeds the word limit restrictions in restricted prompting settings. We show our findings in Table 18. We find that ChatGPT exceeded the word limit restrictions by 73.5% times based on average across all datasets (word limit is exceeded at least more than 50% times in each dataset). The rate of exceeding the word limit restriction is much higher in CNN/DM and XSUM in comparison to SAMSum and DialogSum datasets. This creates a research question to investigate whether LLMs can properly follow the word limit restrictions given on their prompts for response generation.
Appendix G Example of ChatGPT Responses in the EfficientQA Dataset
Here, we discuss some ChatGPT responses in the Efficient QA dataset in the following scenarios:
Generating misinformation (see Table 19 (a)).
Generating the correct answer but the gold answer is outdated (see Table 19 (b)).
Unable to answer time-sensitive questions due to not having the knowledge about the current events (see Table 19 (c)).
Appendix H Example of ChatGPT Responses in Ethical Dilemma Evaluation
We show some example ChatGPT responses to ethical queries in the ethical dilemma evaluation in Table 20.
Appendix I Examples of ChatGPT and other models’ responses to multiple queries in a single input
Here, we show some examples of ChatGPT and other models’ responses to multiple queries in a single input sample (see Table 21 for the responses of InstructGPT series models while Table 22 for the responses of non-InstructGPT series models).
Appendix J Example of wrong responses of ChatGPT in Inverse Scaling sub-tasks
We show some examples of ChatGPT response in the following Redefine Math subtasks: (÷ as digit) and (÷ as digit instead) in Table 23.
Appendix K Detailed Evaluation Results
In this section, we demonstrate a more detailed evaluation result of different datasets:
See Table 26 for the Big-Bench Benchmark.
Appendix L Sample prompts
We show some sample prompts we used for evaluation in some of our datasets in Table 27. Our prompts along with ChatGPT-generated responses in all the datasets that we used for evaluation will be made publicly available.
Appendix M Annotator Experience Survey
The annotator who performed various queries may have a better intuitive understanding of the true limitations and power of ChatGPT. We also conducted a short survey to study the experience of the human annotators of this paper. The annotator experience on ChatGPT can be found in Table 28.