Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, Xipeng Qiu
Introduction
Srivastava et al.introduced the Beyond the Imitation Game benchmark (BIG-bench), a comprehensive evaluation framework designed to assess tasks deemed beyond the reach of contemporary language models. The benchmark encompasses 204 subtasks spanning various domains such as linguistics, child development, mathematics, commonsense reasoning, biology, physics, social bias, and software development, among others. However, due to the absence of real-world test samples, individuals often underestimate the complexity of these tasks, particularly in the context of an ongoing "arms race" among large language models. Consequently, there is a need for an intuitive and practical method to evaluate model performance.
Recent researchers have proposed different Benchmarks originating from standardized testing, including SAT ( Scholastic Assessment Test), AP exams (Advanced Placement examinations) and GRE (Graduate Record Examination), to evaluate the capability of LLMsOpenAI (2023). The benchmarks mentioned above are all English exams, and with the prosperous development of Chinese LLMs, it becomes urgent to propose a benchmark based on Chinese examinations.
We propose using the Chinese college entrance examination (GAOKAO) questions as a suitable alternative. These questions are characterized by their level of difficulty, abstraction, and wide-ranging subject matter, including computational, reasoning, knowledge assessment and writing tasksTan et al. (2021). To this end, we introduce the GAOKAO-Benchmark (GAOKAO-Bench), a benchmark specifically tailored for artificial intelligence evaluation that compiles college entrance examination questions from 2010 to 2022. The GAOKAO-Bench serves a dual purpose. Firstly, it provides a performance metric that aligns with societal understanding, offering a short-term measure of artificial intelligence capabilities. Secondly, the benchmark boasts data of consistent quality, closely mirroring real-world conditions and featuring high-quality annotations, thereby enabling researchers to scrutinize the shortcomings of model performance. Ultimately, these insights can inform and inspire the development of more effective and advanced methodologies.
GAOKAO Benchmark
The Chinese college entrance examination, also known as the GAOKAO, is a comprehensive examination designed to assess the knowledge and skills of high school students applying to universities in China. It is a rigorous and comprehensive examination. Exam subjects are extensive, including Chinese, mathematics, English, physics, chemistry, biology, politics, history and geography. Among them, the mathematics examination questions are divided into two versions for liberal arts and science students. There is a wide range of questions involving logical reasoning, calculation, knowledge quizzes, writing and other types of questions. However, previous benchmarks based on GAOKAO mainly focus on English Yuan and Liu (2022), especially English Reading and Comprehension Questions Zhang et al. (2022).
The GAOKAO-Benchmark established in this paper includes the content of all national exams in the Chinese college entrance examination of all subjects in the past 13 years, providing a intuitive evaluation benchmark for LLMs.
2 Data preprocess
We obtained the questions and answers from all college entrance examinations and transformed them into JSON file format with a one-to-one correspondence using a combination of automated scripting and manual annotation. Mathematical formulas within the questions were converted into LaTex format. Figure 3 provides an example of a mathematical multiple choice question. To generate multiple output results from our model, we utilized a zero-shot prompting strategy Ouyang et al. (2022) and created prompts tailored to different question types. By matching the questions with their corresponding prompts, we obtained a variety of model outputs. The specific prompt examples we used are illustrated in Figure 1.
3 Statistics
All college entrance examination questions from the national examination papers (A & B) between 2010 and 2022 were included in our dataset, covering both science and liberal arts subjects. These questions were mixed and divided into subjective and objective categories, depending on whether they require human scoring. In total, we included 2811 questions, including 1781 objective questions. Table 1 provides a breakdown of the specific types of questions and the corresponding number of questions in each category.
Experiments
We evaluated the performance of GPT-3.5 Turbo on the GAOKAO-Benchmark dataset by scoring the model’s answers against the standard answers. To analyze the model’s effectiveness on college entrance examination questions, we excluded questions containing images since most current LLM models lack visual capabilities.
Following the technical report for OpenAI’s GPT-4 OpenAI (2023), we scored the objective part of the questions using regular matching and corrected the subjective part with the help of human experts. To ensure fairness, high school teachers from Shanghai Caoyang No. 2 Middle School completed this task so that we could compare the results with human benchmark alignments. We used the converted scores as the final result, which represents the average score out of the total possible points for each topic. We calculated the scores of the model for each subject’s objective and subjective questions, and these results are presented in Table 2.
Analysis
We analyzed the scoring rate of subjective questions and objective questions in different subjects of the model, and found that there are large differences in the ability of the model in different subjects
According to the scoring rate of objective questions in each subject, the model performed best in English subjects, and performed worse in physics, chemistry and Math_I. If it is subdivided into various question types, the top three question types in terms of scoring rate are English_Reading_Comp, English_MCQs and English_Fill_in_Blanks. Their scoring rates are 88.3%, 78.1%, and 73.8%, while less than 40% in Math_I_MCQs, Physics_MCQs and Chemistry_MCQs.
2 Subjective questions
The scores of subjective questions are similar to those of objective questions, with the highest score in English and lower scores in Mathematics, Physics and Chemistry. In addition to these common features, we also found that scores on subjective questions were significantly lower than those on objective questions in Physics, Chemistry, Biology and Mathematics, while the other subjects were similar.This may be because the subjective questions in these subjects require more calculation and reasoning. The scoring rates for subjective questions were obtained by human evaluators. And based on the feedback from the human evaluators, we summarize the following flaws.
Complex equations in math problems are difficult to solve correctly, and wrong formulas are used during the problem solving process
Insufficient comprehension and generalization skills for longer reading material
Conclusion
In this paper, we introduced the GAOKAO-Benchmark dataset, which serves as an evaluation standard for large language models. This dataset includes college entrance examination questions from the past two decades, covering various subjects and question types, with an overall high level of difficulty. By testing large language models on the GAOKAO-Benchmark, we can analyze the gap and advantages of these models compared to humans in a reasonable and intuitive manner.
In addition, we evaluated the ability of large language models to answer Chinese college entrance examination questions using a zero-shot prediction approach. Our results showed that the model performed well on objective questions, particularly in the subject of politics, but struggled with certain types of logical reasoning and mathematical problems, as well as with reading comprehension of longer texts in Chinese. However, we found that the model’s performance on knowledge-based questions was better, and by introducing prior knowledge and enhancing input, the model’s performance could be improved to 70%. These findings suggest that large language models have potential applications in education and language assessment, but there is still room for improvement in certain areas. Future work could focus on developing approaches to enhance the model’s performance on longer text reading comprehension tasks and logical reasoning problems.