Mathematical Capabilities of ChatGPT
Simon Frieder, Luca Pinchetti, Alexis Chevalier, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, Julius Berner
Introduction
Since its release in November 2022, the language model Chat Generative Pre-trained Transformer (ChatGPT) has rapidly become a widely known question-and-answer dialogue system. ChatGPT has been referenced in traditional media across the globe and across all major internet platforms . With similar reactions, the release of ChatGPT’s successor, GPT-4, followed in March 2023 .
The performance of ChatGPT has been analyzed in a large number of exam-related use cases, with varying degrees of scientific rigor, ranging from detailed studies to anecdotal evidence. Use cases include passing the United States Medical Licensing Examination (USMLE) , scoring highly on the Psychology Today Verbal-Linguistic Intelligence IQ Test , and answering (and generating) Operations Management exam questions that were deemed to be within the scope of a typical MBA curriculum , all with a performance that elicited a positive sense of surprise from the authors. In turn, the performance of GPT-4 even surpasses that of ChatGPT on a large batch of academic and professional exams [6, Table 1]. Such strong task-related performance indicates that large language models (LLMs) could be frequently used as assistants in many domains.
In this article, we will focus on introducing a new dataset, called GHOSTS, which measures advanced mathematical abilities of LLMs. Using this dataset, we will perform a detailed analysis of the mathematical capabilities of ChatGPT on two of its versions, the 9-January-2023 version and the 30-January-2023 version. Note that, according to the release notes, the 30-January-2023 version should possess “improved factuality and mathematical capabilities” . We further examine the performance of GPT-4 on a smaller dataset, called miniGHOSTS, which exhibits statistics similar to the larger GHOSTS dataset. Our analysis includes but is not limited to testing how many of the skills necessary to do professional mathematics can be emulated by these models. Examples of such skills are the ability to answer computational questions, the ability to complete mathematical proofs that have gaps or missing steps, the ability to solve questions that are more focused on deep insights and original solutions, such as those of mathematical olympiads, and the ability to survey the literature and think across domains. None of the previous benchmarks (see Section 2) cover such a broad range of mathematical abilities.
To achieve the goals outlined above, GHOSTS consists of carefully composed prompts aimed at testing different aspects of LLMs related to mathematical comprehension, see Section 3. This includes both hand-crafted prompts as well as samples from existing datasets that were devised to test models specifically trained for mathematical comprehension .
For brevity, we will use the expression “(Chat)GPT” to refer collectively to both the ChatGPT and GPT-4 language models. We refer to Appendix C for further details on (Chat)GPT versions.
To evaluate the output of (Chat)GPT, we designed a thorough testing methodology, including warning and error codes that represent various possible failure modes of (Chat)GPT. We score (Chat)GPT’s responses, report on the results using this methodology, and compare (Chat)GPT to a selection of state-of-the-art models trained for mathematical comprehension. In summary, the contributions of this article are threefold:
Benchmark for testing the mathematical capabilities of LLMs: We introduce a new natural-language mathematics dataset, called GHOSTSgithub.com/xyfrieder/science-GHOSTS, to test the capabilities of LLMs across a range of aspects regarding advanced mathematical comprehension, see Section 3. It consists of two subdatasets derived from state-of-the-art datasets of mathematical queries for language models. Additionally, we devise four hand-crafted subdatasets covering further mathematical tasks. Parts of our dataset consist of problems that were selected to have a high probability of not being in the data on which (Chat)GPT was trained.
Insight for mathematical use of (Chat)GPT: Based on our benchmark, we show for which types of questions and which domains of mathematics, (Chat)GPT may be useful and how it could be integrated into the workflow of a mathematician. On the other hand, we identify the failure modes, as well as the limits of its capabilities. This can aid future efforts to develop LLMs that perform better in mathematics. Our analysis is akin to a mathematical model card, where the mathematical strengths and weaknesses are summarized, see Section 4.
Evaluation of improvements of (Chat)GPT: We can further use our benchmark to track the mathematical capabilities of (Chat)GPT variants over time. As a first step, we analyze the impact of the upgrade from the 9-January-2023 to the 30-January-2023 version of ChatGPT, which promises “improved factuality and mathematical capabilities”. Then, we proceed to investigate what performance increases the successor GPT-4 brings; see Section 4.1.
Related Work
As a language model, (Chat)GPT can be universally employed to perform mathematical reasoning and therefore has to compete with technologies in this space that are sometimes decades old. Performing mathematical reasoning in an automated way has a long history and can be traced back to 1959 , the most focus being devoted to proving theorems . Presently, there is a realization that classical approaches, using a symbolic encoding of mathematics, have reached a plateau .
On the other hand, there is now a growing body of literature on learning mathematical relationships directly in a supervised-learning manner or by using LLMs to perform mathematical reasoning directly on mathematics encoded in natural language . Sometimes, the distinction is blurred because architectures of LLMs can also be used in a supervised-learning setting and have been employed successfully in learning mathematical relationships .
Among the supervised approaches, we mention , where a Transformer architecture was used to generate symbolic, closed-form solutions to integrals and first and second-order differential equations, which outperformed classical solversFor a given prompt, the computer algebra system is considered to have failed if it does not provide a closed-form solution or times out after 30 seconds (in case of Mathematica)., such as Mathematica, MATLAB, and Maple by at least 14% on a test set of integration problems. On the task of solving differential equations, the Transformer-based approach still exceeds the classical approach, but by a smaller margin (at least 4% in the case of first-order differential equations and with more varied results for second-order equations).
Recent LLMs, for instance, PaLM (released in 2022), are tested only on elementary-level mathematical reasoning datasets, such as the MathQA or GSM8K datasets . We suspect that this is due to a lack of advanced-level natural language mathematics datasets. Moreover, the results obtained indicate that the models at that time had difficulty with much simpler datasets than ours. For example, the version of PaLM with 540 billion parameters only correctly solves 58% of the problems of the GSM8K dataset, even with chain-of-thought prompting and access to an external calculator [22, Table 10]. This model nonetheless outperforms GPT-3 , which only achieves 54% on the same dataset. Variations of BERT have been shown to only solve between 28% and 37% of the problems when fine-tuned and tested on the Algebra Question Answering with Rationales (AQuA-RAT) dataset , which is the direct predecessor of MathQA. For some models, such as BLOOM or the LaMDA model (both released in 2022), an evaluation of the mathematical reasoning capability is entirely missing. An up-to-date survey on mathematical datasets and the performance of various LLMs can be found in .
Among the aforementioned LLMs, Minerva , based on PaLM, stands out, being trained in equal parts on websites that contain MathJax elements and arXiv preprints (additionally to general natural language data on which PaLM was trained). It achieves a score of roughly 50% on the significantly harder Mathematics Aptitude Test of Heuristics (MATH) dataset , which was sourced from various mathematical competitions. One distinguishing feature of the MATH dataset is that its problems admit a unique answer that can be condensed within a few characters (a number, for example). This is beneficial for the automatic evaluation of a model on such a dataset since one can simply check the final answer, ignoring the step-by-step solution.
Most similar to our dataset is the NaturalProofs dataset and the NaturalProofs-Gen dataset . In this paragraph, we illustrate the similarities and differences between these datasets and ours. NaturalProofs and NaturalProofs-Gen are similar among themselves and cover graduate-level mathematics by focusing on data from ProofWikihttps://proofwiki.org/ (the latter dataset), as well as on the Stacks Projecthttps://github.com/stacks/stacks-project and two open-source textbooks (the former dataset). Using the LaTeX source code, which is available for all these resources, annotated theorems and their proof graphs are extracted. The annotations consist of reference graphs highlighting references to other theorems or definitions, the idea being that these references capture the “skeleton” of a proof. This task resembles the mathematical abilities that the Named Theorem Proof Completion subdataset from the GHOSTS dataset evaluates (see Table 1), although 1) we only retrieve a single reference and 2) (Chat)GPT, as far as known, does not use training objectives that make use of information from data annotation, in contrast to models evaluated in . Our framework pertains to general language model evaluation, which may be presented in a black-box manner (as is the case for (Chat)GPT), and therefore does not allow to leverage any additional information, such as reference graphs. This is also reflected in the human evaluation schema introduced in (see Table 24), which classifies common model mistakes. As reference graphs form the foundation of how the mathematical proofs are engineered, many elements of the evaluation schema are strongly tailored toward this representation of mathematical data. Our benchmark is not reference-centric and therefore allows evaluations of any type of proof (including computations, as featured in the Symbolic-Integration subdataset, which we consider to be a particular kind of proof). Therefore, our methodology includes further and more general failure modes to make for a more fine-grained evaluation that explains the nature of the errors. We refer to Appendix A for further related works.
GHOSTS and miniGHOSTS Dataset
We assess the mathematical reasoning capabilities of two ChatGPT versions, 9-January-2023 and 30-January-2023, and of GPT-4 by first creating a collection of prompts from various sources, and subsequently evaluating the models on (subsets of) these data points. We rate the corresponding outputs provided by the models and collect statistics, such as error types, output lengths, or the stability of the answer under prompt engineering, see Sections 3.2 and 4 and Appendices B and D. This yields a total of ratings by human experts.
We divide our dataset, the entire collection of prompts, into six subdatasets, called
which, in turn, consists of multiple files, see Table 1. The boldface letters make up the GHOSTS acronym. Details on motivation, composition, collection process, and intended uses of the GHOSTS dataset are summarized in our datasheet in Appendix E, Sections E.1, E.2, E.3 and E.5, respectively.
GPT-4 was evaluated on a subset of prompts, which we call the miniGHOSTS dataset. Specifically, after having created the GHOSTS dataset, we heuristically selected a subset of prompts from each file of the subdatasets included in GHOSTS, having the same mean rating and the same standard deviation (of ChatGPT’s output) as the original file; see also our datasheet in Appendix E for more information. In this sense, these subsets can be considered to have the most relevance by capturing the “essence” of the model performance in the respective file.
The subdatasets that make up our GHOSTS dataset are summarized in Table 1. In the following, we describe each subdataset in more detail.
This subdataset consists of a collection of books that are widely used in universities to teach upper undergraduate or first-year graduate courses in a degree in mathematics. We have used most of the exercises from these books’ first and second chapters, except for the book , where we only used exercises from the first chapter, which was longer than the other books’ chapters.
This subdataset consists of a number of proofs sourced from math.stackexchange.com, a collection of books , and the MATH dataset , where parts of the proofs were intentionally deleted and the LLM was prompted to fill in the gaps: This was done either by (1) using a MISSING token, (2) finishing the proof early and prompting the LLM to complete it, or (3) explicitly asking for certain conditions or results.
This subdataset consists of a selection of exercises from the book Problem-Solving Strategies , that is often used to prepare for mathematical competitions. We selected and graded the LLM outputs on one hundred exercises drawn from all chapters.
This subdataset consists of random samples of integrals from the test set of . There are three ways in which integrals are generated in : Forward generation (FWD), Backward generation (BWD), and Backward generation with integration by parts (IBP). We sample integrals from FWD test set, integrals from the BWD test set, and integrals from the IBP test set. As these integrals are given in Polish/prefix notation, a natural-language prompt conversion of them is unlikely to be witnessed in the training dataset of (Chat)GPT. The assessment was done by verifying the correctness of the output both by using Mathematica, as well as making use of the provided solutions (in Polish notation), which generated using SymPy. In particular, we notice that all integrals in this dataset have solutions that can be expressed using elementary functions.
This subdataset consists of a random sample of problems from the MATH dataset . The latter dataset attaches a level of difficulty to each problem. We focused on two domains, Algebra and Probability Theory, and sampled an equal number of problems at each level of difficulty.
This subdataset consists of problems that were not sampled from a particular source but generated by a human expert in the field. In the file Named Theorem Proof Completion, we focused on prompting the LLM to provide proof outlines of various theorems that are sufficiently well-known within Functional Analysis to have names. In the Definition Retrieval file, we asked the LLM to correctly state various definitions centered around Functional Analysis and Topology. In contrast, in the Reverse Definition Retrieval file, we verified whether the LLM was able to deduce the name of a mathematical object by describing its properties.
Because input to (Chat)GPT is purely textual (at the time of writing), certain types of questions that might be stated and solved in a non-text-based fashion (e.g., questions involving graphical diagrams, without text explaining the diagramSee, e.g., Exercise 15 in [38, Chapter 2], which asked the reader to inspect a figure on which the problem is based., as occasionally occur in ), have been excluded. Our subdatasets can be categorized along the following dimensions (see Appendix B.1 for more details):
Mathematical difficulty (ascending): (M1) Elementary arithmetic problems, (M2) Symbolic problems, (M3) (Under)graduate-level exercises, (M4) Mathematical olympiad problems.
Question type: (Q1) Stating mathematical facts, (Q2) Overview-type review questions, (Q3) Computational questions, (Q4) Theorem proofs or puzzle solutions, (Q5) Proof-completion questions.
Types of high out-of-distribution likelihood: (D1) Nontrivial problem encoding, (D2) Succinct solution, (D3) Spoken dialogue.
The existing datasets of natural-language mathematics are far from covering all possible combinations across these dimensions. In our well-crafted GHOSTS datasets, we have striven to cover each of these aspects individually, as can be seen in Table 1. The next section specifies the format of our dataset and the methodology for analyzing (Chat)GPT’s output.
2 Format
The format of each of the subdatasets that make up our GHOSTS dataset follows the same convention. Each subdataset consists of JSON-formatted files, and our format is similar to, e.g., the AQuA-RAT dataset . A single data pointThe JSON object of an output of the 30-January-2023 version of ChatGPT, as identifiable by the timestamp, is shown. The prompt comes from the “W. Rudin, Functional Analysis (ch. 1)” file from the Grad-Text subdataset. in a JSON file has the following form:
We require each data point to have the same JSON keys as in this example, some of which may be empty depending on the prompt. Among the listed keys, the rating key stands out as the most fundamental one. Its value serves as a condensed representation of the mathematical capability of the tested language model, compressed into a one-dimensional measure ranging from 1 (lowest) to 5 (highest). A more nuanced and fine-grained perspective on the mathematical capabilities is provided by the errorcodes and warningcodes keys. The msc key denotes the mathematics subject classification. We explain each JSON key in Appendix B.2. For end-users of (Chat)GPT, it is desirable to avoid having a long-winded dialogue to arrive at a solution. Therefore, we require that (Chat)GPT provides us with the correct solution given only the input prompt without any subsequent interaction.
3 Human Effort in Dataset Creation and Mathematical Evaluation
For all data points, the values of the keys rating, errorcodes, warningcodes, comment, and confidence were manually labeled, without any automation. The msc, ref, and timestamp keys were populated in a semi-automatic way since their values change only slightly within the same subdataset.
Two of the subdatasets, the MATH subdataset and the Symbolic-Integration subdataset, use prompts taken from existing datasets, and , respectively. This was done to compare how (Chat)GPT performs against existing state-of-the-art models that use these datasets, see Section 4. Nonetheless, significant additional annotation effort was involved since, in both cases, the authors rated the output. Furthermore, in the second case, the data is publicly presented in a Polish notation format, and manual conversion was necessaryThe authors of were not reachable at the time of our initial contact to provide us with an automatic parser..
The prompts of the other subdataset were hand-crafted by the authors. We note that it is neither possible to outsource the creation of these subdatasets to a crowdsourcing service, such as Amazon Mechanical Turk, nor is it possible to generate them automatically from code because advanced mathematical insight is required for the creation of each prompt (where applicable) and for providing the fine-grained evaluation of the mathematical capabilities. Furthermore, unlike in the case of the MATH dataset by (see Section 2), the answer to most of our prompts cannot be condensed into a few tokens (such as a number or a function), e.g., when the answer is a mathematical proof.
This raises the difficulty of the creation of more data since graduate-level (and in some cases, PhD-level) mathematics is required. The combined effort of devising mathematically insightful prompts and carefully rating the output of (Chat)GPT amounts to prompt evaluations, totaling several hundreds of person-hours, see Appendix B.6. However, as a result of these efforts, our dataset goes beyond all the mentioned mathematical datasets for LLMs in Section 2 in terms of the different aspects of mathematical reasoning that are being tested.
Results
Will ChatGPT get you through a university math class? No, you would be better off copying from your average peer—unless it is undergraduate mathematics, for which GPT-4 can offer sufficient (but not perfect) performance.
If we take a rating of 3.5, the average between the lowest and highest rating, to be the threshold between success and failure, then Figure 1 shows that for a majority of subdatasets, both versions of ChatGPT will not pass. However, for GPT-4, the situation is different, and, on miniGHOSTS, it passes (sometimes barely) on all subdatasets files, except W. Rudin, Functional Analysis (ch. 2), which tests graduate-level mathematical knowledge and the Olympiad Problem Solving file, which tests mathematical problem-solving skills. We note that, unless otherwise stated, we do not use prompt-engineered questions in the results presented here (see Appendix D.1).
We first focus on the results of the 9-January-2023 version of ChatGPT and note that the results for the 30-January-2023 are very similar, as can be inferred from the figures. On average, the 9-January-2023 version achieves a rating of 3.20 with a standard deviationWe use Bessel’s correction term to obtain an unbiased estimate of the variance. of 1.23. It performs particularly poorly on proof-based questions in the style of graduate-level exercises or mathematical olympiads, as well as more complicated symbolic calculations. We note that prompt engineering only slightly improved the results for such complex questions; see Appendix D.1. However, in tasks that only required filling in gaps or stating mathematical facts, ChatGPT was mostly able to achieve a score above 3.5. In particular, ChatGPT was strong at recognizing the context of the question, and the notation of the output almost always matched the one given in the prompt, see Figure 5 in the appendix. Generally, Figure 1 indicates that the ratings closely correspond to how mathematicians would rank the difficulty of the exercises. In this context, we note that the length of the prompt does not have a clear effect on the rating; see Figure 9 in the appendix. We present results for different mathematical fields in Figure 4 in the appendix. For a detailed qualitative analysis of the results on the different subdatasets, we refer to Appendix D.2. Finally, we note that (Chat)GPT almost never expressed any form of uncertainty, even if its output has been completely wrong. This is different from other LLMs we have experimented with; see also Appendix D.3.
Comparing ChatGPT to the performance obtained by , who correctly solved nearly of the integrals in a collection of test equations [12, Table 3], the 9-January-2023 version of ChatGPT achieves an average rating of 2.51 (standard deviation: 0.87) on our random sample of their dataset (after conversion from Polish notation to LaTeX). Specifically, a rating of 2 is dominating of the time, followed by a rating of 3 and 4 for of the prompts each; see also Figure 7 in the appendix. GPT-4 achieves an average of 3.50 (standard deviation: 1.43), barely a passing grade, on the corresponding subset from miniGHOSTS. These scores trail far behind the performance achieved by the model in . The situation is similar when comparing ChatGPT to Minerva [19, Table 3]. Their best model achieved an accuracy of 50% on the MATH dataset . However, the 9-January-2023 version of ChatGPT achieves a perfect score only on of our random samples from the MATH dataset (which is above the total average of of data points across all subdatasets in which this version achieves a perfect score), see Figures 6 and 7 in the appendix. In contrast, GPT-4 performs substantially better and obtains a score of 5 on of the corresponding questions within the miniGHOSTS dataset, see Figure 7 in the appendix.
The ensuing model version, 30-January-2023, overall performed similarly with an average rating of 3.29 (standard deviation: 1.28), although performance was inconsistent across subdatasets and on some subdatasets marginally better, see Figure 1. A significant jump in performance could only be observed for GPT-4, which achieved a substantially higher average rating of 4.15 (standard deviation: 1.12). We note that the evaluation of GPT-4 is only on the miniGHOSTS dataset, i.e., a subset of GHOSTS. Nonetheless, these preliminary findings send a clear message that the performance of GPT-4 dominates the performance of ChatGPT (both versions), see Figure 1.
Figure 2 shows how the ratings change between the different versions of (Chat)GPT. Surprisingly, one can see a shuffling of the grades for the two ChatGPT versions, even though the counts in each grade bracket stay approximately the same. For instance, there are roughly the same amount of outputs that received grade 4, yet less than half of the prompts were the same between model changes. Appendix D.5 provides different perspectives on this and reinforces the mixed performance increase the 30-January-2023 model brings. For GPT-4, we see that the percentage of perfect ratings almost doubles, while the percentage of prompts, which are not understood or completely wrong (i.e., ratings of 1 or 2), approximately halves as compared to the ChatGPT versions.
Analysis of (Chat)GPT’s output and our warning codes reveal that GPT-4 provides even longer (“rambling”) answers, whereas ChatGPT usually answered the question without giving any additional context about the topic, see Figures 6 and 8 in the appendix. The answer style of GPT-4 was often beneficial (resulting in better overall scores) but sometimes reduced the readability of the output. Furthermore, we found the behavior of GPT-4, compared to ChatGPT, to be more opinionated. Finally, despite its better overall performance, GPT-4 still seems to be vulnerable to mistakes in seemingly simple calculations. We refer the reader to Appendix D for further results on the models’ performance.
Conclusion
We have examined the behavior of (Chat)GPT across various tasks that test different aspects of mathematical skill. Contrary to the media sensation that (Chat)GPT has caused, (Chat)GPT is not yet ready to deliver high-quality proofs or calculations consistently. At the same time, the quality of the answers can be positively surprising. Moreover, our preliminary evaluation of GPT-4 on the miniGHOSTS dataset reveals promising improvements over ChatGPT’s performance. In Appendix D.6, we collect the best and worst results for a number of selected subdatasets. The best responses can be seen to justify the media sensation. It thus seems fair to say that (Chat)GPT is inconsistently bad at advanced mathematics: While its capabilities generally drop with the mathematical difficulty of a prompt, it does give insightful proofs in a few cases.
However, (Chat)GPT falls short of achieving the same performance as models specifically trained for single tasks. These models, in contrast, lack the flexibility of (Chat)GPT, which is a universal tool suitable for any area of mathematics. In fact, (Chat)GPT’s ability to search for mathematical objects, given information about them, is where it shines. For a user that is already sufficiently mathematically proficient to discern the correctness of (Chat)GPT’s output, (Chat)GPT can be integrated as an assistant in the user’s workflow. It can function as a search engine or knowledge base to speed up various lookup tasks, as they often occur at certain stages of mathematical research.
Due to the prohibitive annotation effort, the GHOSTS dataset is not yet large enough to significantly improve the mathematical capabilities of LLMs by fine-tuning them on GHOSTS; though we believe it is sufficiently comprehensive to allow an evaluation and comparison of LLMs. As a first step, we want to extend the evaluation of GPT-4 to the full GHOSTS dataset, considering its promising performance on miniGHOST. We also encourage other researchers to mine our dataset beyond the descriptive statistics we computed in order to gain a deeper understanding of how LLMs behave on mathematical tasks. Finally, we hope that our work motivates other mathematicians to contribute to the GHOSTS dataset in order to establish a thorough benchmark for assessing the mathematical abilities of LLMs.
References
Appendix A Further Related Works
In this section, we present further related works. For (Chat)GPT, most investigations related to mathematical reasoning consist of anecdotal evidence concerning its performance and its failure modes. Notable mentions on social media can, for instance, be found in . Unfortunately, a clear methodology is missing, as most of the results are scattered on various internet platforms and cannot be easily reproduced. To the best of our knowledge, the only investigations into the mathematical capabilities prior to the appearance of our first preprint were undertaken by . However, these works only report a small number of qualitative results, often on rather simple mathematical tasks and without specifying the precise versions of (Chat)GPT. The latter reference reports results only on a few selected examples, while the former reference investigates ChatGPT’sUsing an unknown version of ChatGPT that predates the 9-January-2023 version. ability to compute irrational numbers as well as to solve some elementary math word problems. Recently, the dataset by appeared, which contains a systematic evaluation of ChatGPT on the GSM8K dataset , the MATH dataset , and the MMMLU-STEM dataset . These datasets allow for an automatic evaluation using only accuracy as an evaluation metric. Additionally, a few further anecdotal examples of mathematical performance are presented in .
Finally, we would also like to mention the field of formalized mathematics, where large databases that encode advanced mathematical concepts exist, e.g., the Lean Mathematical Library . Some of the ideas that we have used in this article, such as using prompts that formulate a task to fill in gaps in proofs, are echoed in for datasets for formal mathematics, consisting of expression trees. Yet, for the purpose of doing mathematics with large language models, these formal datasets cannot be leveraged since no straightforward way exists to convert them to natural language.
Appendix B Dataset Creation
Our subdatasets can be categorized along the following dimensions, see Table 1:
Elementary arithmetic problems, as found in the MATH dataset at lower levels of difficulty.
Symbolic problems (integration of functions) that can be also solved via a supervised-learning, data-driven approach to mathematics .
(Under)graduate-level exercises from well-known textbooks as well as questions from math.stackexchange.com, spanning diverse domains of mathematics.
Exercises that are in the style of mathematical olympiad problems, such as those taken from Engel’s Problem-Solving Strategies book .
Review questions, which ask to state or name certain mathematical facts correctly.
Overview-type review questions, which cut through an entire field of mathematics.
Proof-based questions, which ask for a theorem proof or for a puzzle solution.
Proof-completion questions, where a proof has gaps or is incomplete, and needs to be completed.
Types of high out-of-distribution likelihood:
Nontrivial problem encoding: The data points from the Symbolic Integration subdataset come from and are publicly availablegithub.com/facebookresearch/SymbolicMathematics. Since the online training set uses Polish notation, it is very unlikely that (Chat)GPT has seen the corresponding prompts in LaTeX before.
Succinct solution: The solutions for the Olympiad-Problem-Solving subdataset are included in the book by Engel . But the solutions are extremely concise, and simply repeating them would not show an immediate understanding of the problem.
Spoken dialogue: The Search-Engine-Aspects subdataset is unlikely to be well represented in the data on which (Chat)GPT has been trained since its prompts resemble word fragments that might appear in a mathematical dialogue (e.g., an oral mathematical exam), rather than in a textbook.
One could, in theory, start to investigate every combination of these attributes (e.g., for elementary arithmetic problems, in a non-trivial encoding, one could generate data to cover every possible question type listed above). However, this would lead to 60 subdatasets, which, due to the manual curation effort, is too much for a single research group.
B.2 Format
The dataset consists of a collection of UTF-8 encoded JSON files. We explain the JSON keys of each data point in our dataset in the following and also indicate whether its value is optional. If the value is optional, the key has to be present, but the value will be an empty array or string.
prompt denotes the input that we provide to (Chat)GPT through its web interface at the URL chat.openai.com/chat, see also Appendix C. We use a new session for each prompt to avoid (Chat)GPT being biased by previous prompts.
output denotes the raw output that (Chat)GTP supplies us with. In some cases, mathematical formulas were rendered in the web interface such that we copied them in LaTeX.
rating is a number from 1 to 5 that shows how many points (Chat)GPT has scored, 5 being a perfect answer and 1 being the lowest rating. A detailed explanation of the rating policy that we followed is contained in Appendix B.4.
errorcodes (optional) highlight a list of error types that illustrate the failure modes of (Chat)GPT in a more fine-grained way. Not all types of errors apply to all (sub)datasets: For example, an error code for a missing proof step would not be applicable on a dataset that tests whether (Chat)GPT can multiply numbers or find prime divisors. The detailed explanation of the error codes (and the warning codes; see below) that was provided to the annotators is contained in Appendix B.4. There, we also include a policy of how ratings and error codes have to be used together.
warningcodes (optional) highlight any problematic aspects of (Chat)GPT; for example, (Chat)GPT might be rambling and providing the user with unrelated information or use a poor (but correct) way of solving problems.
comment (optional) denotes any noteworthy commentary that an assessor of (Chat)GPT may make. This can be used to give a more detailed explanation of the output, provide reasoning behind awarding a certain error code or rating, or generally provide context. For some subdatasets, this key was used to indicate the difficulty level of the prompt, as well as an official solution, if available, see Section 3.1. It was also used to indicate whether we used prompt engineering, see Appendix D.1.
msc denotes a list of mathematics subject classificationsA complete list of MSC codes can be accessed at the URL zbmath.org/static/msc2020.pdf. (MSC) that pertain to the output. Note that we do not classify the prompt given to (Chat)GPT as there may be no proper classification; for example, when (Chat)GPT is asked what the most important theorem in all of mathematics isThe answer is Pythagoras’ theorem, according to (Chat)GPT., it is meaningless to assign an MSC code. We also note that for particularly easy mathematical questions (e.g., simple arithmetical questions), no suitable MSC codes exist to classify the output, since MSC codes typically classify more advanced mathematicsThe MSC codes starting with the numbers “97”, which at first glance might be most suitable, are solely reserved to classify content that is related to the educational process of mathematics, rather than the mathematical content itself.. Nonetheless, we have attempted to match them as well as possible and allow multiple MSC codes in order to classify the output as precisely as possible.
ref (optional) indicates a reference to where the prompt was originally taken from (for some subdatasets, such as Holes-in-Proofs, we have used excerpts from various books or math.stackexchange.com; the original source was recorded as a value of this key). This key can have an empty value if the question was formulated by the authors and no authoritative source was plausible.
confidence indicates how confident we have perceived (Chat)GPT to be when presenting us with its output. We allow values of high, medium, and low.
timestamp denotes when the prompt was entered into (Chat)GPT. This can be used to track the version of (Chat)GPT; see Section 4.1.
The values of these keys within a single data point interact in nontrivial ways: If a rating of 5 is given, then it is expected that no error code is present—though there may be warning codes that are used. The error codes and warning codes are loosely in the spirit of a compiler throwing errors and warnings if it is given incorrect or sloppy code—although we have a role reversal, where the human is now the compiler, and the machine produces the code. In this sense, for some prompts, we have used multiple error and/or warning codes, which is why the corresponding values are arrays of strings. We use these codes to collect statistics on the behavior of (Chat)GPT; see Section 4.
For most of the subdatasets that make up our GHOSTS dataset, we have used LaTeX to encode mathematical formulas in our prompts. Our experiments have shown that (Chat)GPT can process LaTeX-encoded mathematics well.
The usage of MSC codes can be useful for mathematicians who want to integrate (Chat)GPT in their daily workflow, as it allows them to know in which areas the model performs better and can hence be trusted more. Our dataset is very diverse, having a total of MSC codes. The top short versions of these codes (first two digits) are 26 (“Real functions”, occurrences) followed by 00 (“General”, occurrences) and 46 (“Functional analysis”, occurrences), see also Figure 4. An exhaustive survey of (Chat)GPT’s performance across every MSC code would necessitate a large, community-driven effort to set up an extensive database. Due to the high cost of rating each output, requiring specialized skills, this is something that no individual research group could reasonably do—but we hope that our approach is a starting point for such an effort.
B.3 Copyright and Licensing Terms
Some of the subdatasets contain prompts that may be protected under copyright, i.e., from the Grad-Text and Olympiad-Problem-Solving dataset. In these cases, the publicly released dataset does not contain the prompt. The ref key includes a detailed reference to the page where the original theorem or exercise that was presented, so a reader can easily retrieve the prompt. All other prompts are either created by us or released under licenses that allow us to include the prompt.
For the prompts that are not created by us, the following applies: We license the entire data point (i.e., the content of all JSON keys except the prompt key, i.e., the content created by the authors) under the same license as the prompt. The following licenses, therefore, apply in the cases of data points using prompts from external sources:
The MATH subdataset is distributed under an MIT license.
The Symbolic-Integration subdataset is distributed under a Creative Commons Attribution-NonCommercial license.
Prompts originating from user contributions on math.stackexchange.com, see the ref key for such occurrences (e.g., in the Proofs Collection A file), are licensed under Creative Commons Attribution-ShareAlike license, in different versions, see https://math.stackexchange.com/help/licensing.
We release prompts from the GHOSTS and miniGHOSTS datasets that are created by us under the following Creative Commons license: Attribution-NonCommercial 4.0 International (CC BY-NC 4.0); see https://creativecommons.org/licenses/by-nc/4.0/ for the detailed terms of the license. By this license, one may not use the dataset for commercial purposes, and one must give appropriate credit; if users are building on the GHOSTS dataset, they need to indicate the changes that were made and distribute their contributions under the same license as the original.
B.4 Data Collection and Labeling Policies
Prompts from books were transcribed into LaTeX. The output from (Chat)GPT’s web interface was copied as-is, even if the output was not valid LaTeX code. If the output contains rendered mathematical expressions, our policy was to transcribe it to LaTeX. Below are the policies that were followed by each assessor of (Chat)GPT’s output regarding the rating, the error codes, and the warning codes:
1 failure to understand the query (e.g., the user asks it something about number theory, and it responds with information about differential equations);
2 query was understood, but the answer was entirely wrong (e.g., the user asks what the prime divisors of 111 areThey are 37 and 3., and it responds with 8 and 6);
3 query was understood, but the answer was only partially correct (e.g., the user asks it what the prime divisors of 111 are, and it responds with 3 and 6);
4 query was understood, and the answer was mostly correct (e.g., the user asks it what the prime divisors of 222 areThey are 2, 37, and 3. and it responds with 3 and 37);
5 query was understood and answer was completely correct.
e1 missing examples, or information (e.g., the user asks it what the prime divisors of 111 are, and it responds with 3, missing 37); this also applies, if (Chat)GPT ignores a part of the prompt (e.g., an equivalence needs to be shown, but (Chat)GPT shows only one direction);
e2 a few wrong/vague statements (e.g., the user asks it what the prime divisors of 30030 areThey are 2, 3, 5, 7, and 11. and it responds with 2, 3, 5, 7, 13 (wrong); or says that 2, 3, 5, and some other numbers are prime divisors (vague)); it can also denote a single statement, that is slightly vague;
e3 a lot of wrong/too vague statements (e.g., the user asks it what the prime divisors of 30030 are, and it responds with 2, 5, 8, 12, 13, 15 (wrong); or says that 2 and many other numbers are prime divisors (vague)); it can also denote a single statement, that is highly vague;
e4 wrong computations (i.e., an additional error flag to disambiguate between statements that are of computational nature or not);
e5 denotes wrong logic or wrong flow of arguments, which we further subdivide into specific flags, as we prohibit the use of e5 on its own (since it would be uninformative):
e5_1 (Chat)GPT claims that to complete a proof, statements need to be shown that are unrelated to the claim;
e5_2 a proof step is missing;
e5_3 an edge case has not been considered;
e5_4 an inference step is not supported (e.g., (Chat)GPT claims that from A follows B, but this claim is not true);
e5_5 circular logical argument (i.e., using the hypothesis to prove the hypothesis);
e6 the general set-up is understood, but the legal operations are not respected or misunderstood (e.g., we are given a puzzle where we are only allowed to add even integers, but (Chat)GPT changes the rules and motivates the solution by allowing the addition of odd integers; or (Chat)GPT misunderstands an adjective that has multiple mathematical meanings, such as “dual”, which can mean either topological dual space or algebraic dual space).
The following policy applies for error codes: If a rating with 1\ <r<\5 has been given, then an error code is mandatory to explain the type of error that occurred. For a perfect score of 5, no error codes should be assigned (but warning codes can be assigned). If the score is lowest, i.e., a rating of 1, error codes can be assigned, but do not have to: In the case where (Chat)GPT has not understood the prompt, there typically is no reason to further detail the type of error.
w1 (Chat)GPT is withholding essential information related to the prompt (e.g., the user asked it something about the integral , and it answers correctly but does not tell the user that the integral was actually a famous, named integral, i.e., the Gaussian integral);
w2 (Chat)GPT is rambling (i.e., after answering, correctly or incorrectly, (Chat)GPT tells the user much more details than the user wanted to know);
w3 (Chat)GPT is hallucinating (i.e., after answering, correctly or incorrectly, (Chat)GPT tells the user unrelated information);
w4 (Chat)GPT behaves weirdly (e.g., by using a weird proof structure (where applicable), using strange mathematical formulations, or by adopting a strange tone of the conversation or making opinionated statements);
B.5 Mitigating Human Errors
Any assessment procedure that has a human component is prone to introducing bias—in particular, a procedure involving manual work such as rating the model outputs. The following safeguards help to mitigate bias as well as human error (such as typos):
Various typographical errors may appear due to incorrect LaTeX formatting. In this case, we noticed that (Chat)GPT was able to correctly infer what was intended (e.g., was correctly interpreted as ), and therefore provided a safeguard against these types of errors.
We presented clear instructions to each author who prompted (Chat)GPT on how to record and save the data in order to avoid any file encoding issues. In the end, all JSON files were inspected and streamlined to Unicode.
Clear instructions were given to all authors that used (Chat)GPT to ensure that the language model has, to the extent possible, an identical state and starts from a blank chat.
Guarding against missing data and copy-paste errors:
Given a lack of API access in the early stages of our investigation (see Appendix C), there was a fair amount of data being copied from (Chat)GPT. To mitigate any copy-paste errors, several passes over the entire dataset, as well as automatic checks, were made to look, e.g., for potential inconsistencies, missing timestamps, and outputs not matching the prompts.
Guarding against other unforeseen errors:
Random samples: Random samples () were drawn from each dataset, and a second assessor reviewed the rating. If deemed problematic, the original assessors were asked to re-evaluate.
Statistical checks: Additional statistical checks were carried out as plausibility checks to make sure no other unforeseen errors occurred: If prompts deviated from the average length on that dataset, they were flagged and the output was manually inspected, and, if deemed necessary, a re-evaluation was carried out.
We are aware that these measures are not exhaustive, but given a fixed time budget, we considered them the most feasible.
B.6 Labeling Effort
The evaluation was carried out by a subset of the authors of this paper who have substantial mathematical expertise, ranging from master’s degrees in mathematics to postdoc-level and professor-level positions at departments of mathematics. Assignment of prompts was done based on difficulty, with more senior mathematicians having received more difficult prompts. No third parties were involved.
Each of the prompts of the GHOSTS dataset was evaluated on both the 9-January-2023 and 30-January-2023 version of ChatGPT; an additional prompts were used to test the effect of prompt engineering on a single type of subdataset, see Appendix D.1. We further evaluate GPT-4 on the prompts of the miniGHOSTS dataset. This amounts to a total of prompt evaluations of advanced mathematics, performed by graduate-level researchers.
We like to mention that our effort has occasionally unearthed small inconsistencies in existing datasets: For example, the “MATH Counting and Probability” file, which was sourced from the larger MATH dataset , contains the prompt “What is the value of ?”, which is neither about counting, nor about probability, but arithmetic (our MSC codes allow users to find such examples).
Appendix C Details on (Chat)GPT
GPT-4, launched on 1st March 2023, is the latest model of the GPT lineage , being the successor of various versions of ChatGPT, the first of which was launched on 30 November 2022 . These are all based on InstructGPT, which in turn is based on a trained GPT-3 , and fine-tuned using reinforcement learning with human feedback .
We note that already for models that predate (Chat)GPT, such as InstructGPT, where research articles and model cards have been released, full reproducibility is not possible since the code and exact datasets have not been released. Furthermore, it was confirmed by OpenAI employees that for some of their models, launched prior to 30 November, a slight mismatch exists between the trained model that is accessible via the OpenAI web interface and the model referred to in the official paper . This indicates how essential it is to document carefully which model our analysis pertains to and how we have accessed it. In our dataset, we have accordingly included time stamps for each prompt in order to be able to track, based on information provided by OpenAI, any changes in (Chat)GPT’s version that have occurred.
We have exclusively used the GUI web interface to carry out the evaluation. This was necessary for consistency reasons, since at the beginning of our evaluation, API access was not yet widely available. At the time of writing, API access to GPT-4 is still limited, and a waitlist is employed, which made the use of the GUI web interface a necessity for GPT-4 ). We note that there exist no official documents that link the GUI web interface to the different model versions and possible model settings from the API. The 9-January-2023 and 30-January-2023 ChatGPT versions we evaluated are likely to be earlier instances of the newer model gpt-3.5-turbo-0301. This model itself is a “snapshot of gpt-3.5-turbo from March 1st, 2023” that will not receive further updates . For GPT-4, at the time of writing, there exist four models gpt-4, gpt-4-0314, gpt-4-32k, gpt-4-32k-0314 that can be used for the chat completion API endpoint . Additionally, for all models, there exist various settings when using the chat completion API endpoint, such as temperature or presence_penalty that influence the models’ output, which cannot be controlled via the GUI web interface. It is also not known which values of these settings are used for the GUI web version. The only version identifier in the GUI web version is a generic “model version” link at the bottom of the page that links to the release notes . The 9-January-2023 and 30-January-2023 model versions that we evaluated are the ones presented in the release notes at the respective time.
Appendix D Further Results
One interesting finding of our study is related to performing prompt engineering on mathematical questions. Prompt engineering was solely carried out on questions from the Olympiad-Problem-Solving subdataset, and prompt-engineered questions consist of lists consisting of two JSON objects. These lists contain the original question, that was not prompt-engineered, as well as the prompt-engineered question. The latter question is identified as it contains the string
About 20% of the questions were prompt-engineered: ChatGPT was additionally instructed to proceed either step-by-step or the mathematical task was formulated in a more explicit way, i.e., by addingSome prompts, e.g., the ones taken from the book by Engel , only contain a mathematical statement, without a clear instruction; for example, “An rectangle can be covered by rectangles iff or ”. From the context, one must conclude that this statement is correct and should be proven. “Prove that…” or “Show that…” to the prompt. Instructing ChatGPT to proceed step-by-step is a type of engineering that is recommended by OpenAI in their cookbook to improve reliabilitygithub.com/openai/openai-cookbook/blob/main/techniques_to_improve_reliability.md. As a result of prompt engineering, for the 9-January-2023 version of ChatGPT, the number of wrong statements and computations (i.e., error codes e2, e3, and e4) decreased, while the number of errors rooted in faulty logic (i.e., error code e5) actually increased. Overall, prompt engineering improves the average rating only slightly, see Figure 3.
For the questions from Olympiad-Problem-Solving that were selected for the miniGHOSTS dataset, we allow to sample from the entire Olympiad-Problem-Solving subdataset, since the goal of miniGHOSTS is not to measure prompt-engineering effects. Therefore, some of the questions in the miniGHOSTS version of the Olympiad-Problem-Solving subdataset contain prompt-engineered questions. The
D.2 Qualitative Analysis of Subdatasets on ChatGPT 9-January-2023
In this section, we go through common mistakes performed by ChatGPT, as well as notable observations regarding the output, one subdataset at a time. We focus on the 9-January-2023 version, see Section D.5 for more information regarding the other version. We note that the output of (Chat)GPT (and, generally, LLMs) is stochastic and therefore may differ on the same prompt. Nonetheless, clear trends can be observed, which we describe here. Individual outputs can be found in Appendix D.6.
ChatGPT, version 9-January-2023, performed best on simple set-theory and logic questions (the first chapter from the book Topology by J. Munkres ), which is reflected in its rating, see Figure 1. On the rest of the books, it performed substantially worse. Because of the confidence (high) with which it outputs the answer, the use of ChatGPT, version 9-January-2023, is particularly deceiving in this use-case since it may be intensively used by students studying these subjects.
ChatGPT, version 9-January-2023, correctly recognized most well-known results or concepts (e.g., filling in the mean-value theorem, given a proof that lacked a reference to it). However, the ability of ChatGPT to execute algebraic manipulations is surprisingly inconsistent. In some cases, ChatGPT executes complicated symbolic tasks with ease; in other cases, it fails on simple arithmetic operations or rearranging terms. The mistakes do not seem to correlate with the complexity of the algebraic expression. When ChatGPT makes an algebraic mistake, it mostly carries over this mistake reliably to the rest of the computation.
On this subdataset, ChatGPT, version 9-January-2023, performed the poorest. From a mathematical point of view, these questions were also by far the most difficult, as they can pose difficulties even to professional mathematicians. A score of 3 was awarded when the answer started to show promise. However, of the scores are 2 because the answer does not show any promise. No rating of 5 was awarded, and only one rating of 4 was achieved. This version of ChatGPT had a tendency to try and solve many questions using induction arguments. While this is not necessarily false, this was very far from the solutions given in the book, and this version’s inductive proofs were easily seen to contain mistakes. In addition, ChatGPT often had difficulty understanding unconventional puzzles. For example, in the questions involving changing the color of squares on a chessboard, the solution offered by ChatGPT did not cover an chessboard. Sometimes it tried to solve the problem by changing only squares, far from the required. Similarly, the 9-January-2023 version of ChatGPT struggled to respect unusual constraints in the questions, resulting in e6 errors, the highest number of e6 errors out of all subdatasets. In some cases where the problem seemed to require complicated mathematics but was actually solvable by elementary techniques, ChatGPT did not spot this but instead referred to the general theory of, e.g., diophantine equations. Interestingly, ChatGPT would sometimes say, e.g., that the question could be solved with these means but that this was hard, so the confidence score was downgraded in these cases to medium or low.
The 9-January-2023 version of ChatGPT was dominated by systems that were trained specifically to solve integration problems . In a number of instances, this version got the structure of terms right (for example, the number of summands in the output, as well as where factors had to be placed before summands), but it failed at concrete computations. Even very simple examples were not correct. For example, the antiderivative of is evaluated to , where is a constant of integration (the correct answer being ). For a number of prompts, this version claims there is no closed-form solution for the integral with complete confidence when, in fact, there is a solution; only integrals that have an elementary antiderivative are in this dataset.
On the questions related to Algebra and Probability theory, the 9-January-2023 version of ChatGPT got the reasoning often correctly. However, the most common type of error was e4, occurring of the time (in total times). This version of ChatGPT may struggle when confronted with standard operations, such as inverting fractions, least common multiples, and changing the sign of numbers when moving them from one side of the equal sign to the other. Often, in these questions, a correct solution requires performing multiple operations in sequence. In such cases, most often, at least one operation was wrong. This prevented the model from getting a rating of 5 on the output, which was only achieved for of the questions.
On the Search-Engine-Aspects file, the 9-Januar-2023 version of ChatGPT knew almost all the theorems that it was asked at a basic level but made mistakes when stating them. When it came to listing other results required for the proofs, this version typically requested way more than the necessary theory—occasionally even results that only follow from the theorem which was asked for (error code e5_5). On the Definition Retrieval file, this version had quite a good performance: it recited most definitions correctly. It sometimes got confused when being asked about distributions in the sense of elements of the dual space of test functions. ChatGPT, version 9-January-2023, strongly favors the notion of distributions in the stochastic sense. Similarly, for the adjective “closed”, where it chose to pick the context of algebra (instead of topology) and interpreted it to mean “algebraically closed”. On the Reverse Definition Retrieval file, this version had the strongest performance, being able to recover most definitions from their descriptions, with an average rating of 4.30 (standard deviation 1.14). This indicates the usefulness of ChatGPT as a general-purpose mathematical search engine. This subdataset is also the simplest from a mathematical point of view since no logical thinking is required, but only a name needs to be found.
D.3 (Chat)GPT’s Confidence
(Chat)GPT is usually very confident, unlike other LLMs that we have experimented with. As an illustrative example, consider the following prompt testing the sensitivity to LaTeX-encoded mathematics vs. Unicode-encoded mathematics:
The response by ChatGPT is not phrased in order to show any nuance in terms of confidence (which is typical, even if ChatGPT is wrong):
The response by Codex , another model that we briefly tested (but whose scope would have exceeded that of a single conference article) gives a cautions response and, unlike ChatGPT, is capable of voicing doubt:
D.4 Figures of ChatGPT’s Performance (version 9-January-2023)
In this section, we collect figures that extend the discussion in the main body and provide further views on the data and descriptive statistics.
D.5 Comparison of (Chat)GPT Versions
In this section, we collect figures which illustrate the differences and similarities between versions of (Chat)GPT. We note that even though the 30-January-2023 version performs very similarly to the 9-January-2023 version, there are some differences in the distribution of ratings, error codes, and warning codes, see Figure 6.
On the other hand, GPT-4 strictly dominates the ChatGPT versions in terms of performance. It always provides context around the question (whether that was asked for or not) and often gives useful (and correct) pointers that, for example, highlight the importance of a particular theorem. Figure 8 depicts the verbosity of different (Chat)GPT versions and the achieved rating. However, we also note that the optimal level of verbosity can depend on the mathematical background of the user. As a result, there have been significantly more warning codes of type w2 (i.e., rambling) for GPT-4, see Figure 6.
D.6 Best-3 and Worst-3 Across Selected Subdatasets
We list below the best and worst answers of ChatGPT, version 9-January-2023, over a selection of subdatasets. For readability, the prompts and answers are lightly modified so that the LaTeX-based formulas are correctly displayed, and whitespace is removed.
Examples from the Grad-Text subdataset, comprising the books .
D.6.2 Holes-in-Proofs (Proofs Collection A)
Examples from the Holes-in-Proofs subdataset, Proofs Collection A file, based on the books and questions from math.stackexchange.com
D.6.3 Holes-in-Proofs (Proofs Collection B Prealgebra and Precalculus)
Examples from the Holes-in-Proofs subdataset, Proofs Collection B Prealgebra and Proofs Collection B Precalculus files, based on .
D.6.4 Olympiad-Problem-Solving
Examples from the Olympiad-Problem-Solving subdataset based on the book .
D.6.5 Symbolic-Integration
Examples from our Symbolic-Integration subdataset based on .
Appendix E Datasheet for the GHOSTS Dataset
This appendix provides a datasheet for the GHOSTS dataset. The format of this datasheet was introduced in and consolidates the motivation, creation process, composition, and intended uses of our dataset as a series of questions and answers.
For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled? Please provide a description.
The existing datasets of natural-language mathematics are far from covering all the typical tasks professional mathematicians encounter in daily life, making it unclear whether language models can be of any help in this regard. Existing datasets mostly cover elementary mathematics or resemble standard tests like SATs (see Sections 2 and 3). Hence, they do not offer any insight into the usage of ChatGPT as a tool for mathematicians. In this work, we have made the first attempt towards filling this gap, going beyond math problems that are yes-no rated, and proposed a benchmark made and curated by working researchers in the field that tests different dimensions of mathematical reasoning.
Who created this dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)?
The authors of this work created GHOSTS; see Appendix B.6 for more information.
Who funded the creation of the dataset? If there is an associated grant, please provide the name of the grantor and the grant name and number.
There is no associated grant or funding which has been used to create the GHOSTS dataset.
E.2 Composition
What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)? Are there multiple types of instances (e.g., movies, users, and ratings; people and interactions between them; nodes and edges)? Please provide a description.
GHOSTS consists of textual prompts, in natural language, representing mathematical questions. For each prompt, GHOSTS contains one or more instances of outputs of (Chat)GPT and corresponding fine-grained evaluation by the authors.
How many instances are there in total (of each type, if appropriate)?
There are prompts in GHOSTS; a selection of of these makes up miniGHOSTS. For of the questions, light prompt engineering variations have been carried out. Each of the questions from GHOSTS has been evaluated on ChatGPT, version 9-January-2023 and 30-January-2023, and questions from miniGHOSTS have been evaluated on GPT-4. Thus, in total outputs and evaluations have been carried out. See also Appendix B.6 for more information.
Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances because instances were withheld or unavailable).
GHOSTS tries to cover a wide range of mathematical questions from different MSC codes; see Appendix B.1 and B.2. However, due to the prohibitive cost of human evaluation, which cannot be fully automated away (see Section 3.3), it is not feasible to represent all mathematical fields across all dimensions of “mathematical behavior” and all types of mathematical questions (overview questions, fact-stating questions, etc.).
What data does each instance consist of? “Raw” data (e.g., unprocessed text or images) or features? In either case, please provide a description.
GHOSTS and miniGHOSTS consist of a collection of JSON objects (one for each data point), and each JSON object consists of key-values pairs as detailed in Appendix B.2.
Is there a label or target associated with each instance? If so, please provide a description.
No, we do not explicitly define a label or target for the instances. However, the rating of the output can potentially be used to select good and bad mathematical conversations of (Chat)GPT in order to fine-tune models and the errorcodes and warningcodes can be used to make a more fine-grained classification possible.
Is any information missing from individual instances? If so, please provide a description, explaining why this information is missing (e.g., because it was unavailable). This does not include intentionally removed information but might include, e.g., redacted text.
Are relationships between individual instances made explicit (e.g., users’ movie ratings, social network links)? If so, please describe how these relationships are made explicit.
Relations between instances are explicitly given by the same values on (subsets) of the fields, e.g., the same prompt, the same model version, or the same MSC code. Prompt-engineered variations of the same question are represented as an array of JSON objects, one object for each variation.
Are there recommended data splits (e.g., training, development/validation, testing)? If so, please provide a description of these splits, explaining the rationale behind them.
Are there any errors, sources of noise, or redundancies in the dataset? If so, please provide a description.
The evaluation of the prompts included in GHOSTS underlies human errors. However, we tried to mitigate these errors; see Appendix B.5.
Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)? If it links to or relies on external resources,
Are there guarantees that they will exist, and remain constant, over time?
Are there official archival versions of the complete dataset (i.e., including the external resources as they existed at the time the dataset was created)?
Are there any restrictions (e.g., licenses, fees) associated with any of the external resources that might apply to a future user? Please provide descriptions of all external resources and any restrictions associated with them, as well as links or other access points, as appropriate.
The dataset is self-contained. However, of the prompts from the Grad-Text subdataset cannot be publicly released since they are taken or adapted from sources that are protected by copyright; see Appendix B.3; though we do release the output of the models on these prompts, which make up human expert evaluation.
Does the dataset contain data that might be considered confidential (e.g., data that is protected by legal privilege or by doctor-patient confidentiality, data that includes the content of individuals non-public communications)? If so, please provide a description.
Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety? If so, please describe why.
Does the dataset relate to people? If not, you may skip remaining questions in this section.
Does the dataset identify any subpopulations (e.g., by age, gender)? If so, please describe how these subpopulations are identified and provide a description of their respective distributions within the dataset.
Is it possible to identify one or more natural persons, either directly or indirectly (i.e., in combination with other data) from the dataset? If so, please describe how.
Does the dataset contain data that might be considered sensitive in any way (e.g., data that reveals racial or ethnic origins, sexual orientations, religious beliefs, political opinions or union memberships, or locations; financial or health data; biometric or genetic data; forms of government identification, such as social security numbers; criminal history)? If so, please provide a description.
E.3 Collection Process
How was the data associated with each instance acquired? Was the data directly observable (e.g., raw text, movie ratings), reported by subjects (e.g., survey responses), or indirectly inferred/derived from other data (e.g., part-of-speech tags, model-based guesses for age or language)? If data was reported by subjects or indirectly inferred/derived from other data, was the data validated/verified? If so, please describe how.
We collected and constructed prompts from various sources, see Table 1 and Section 3. For the evaluation, we captured the corresponding outputs of (Chat)GPT and rated them according to the instructions in Appendix B.2 and B.4.
What mechanisms or procedures were used to collect the data (e.g., hardware apparatus or sensor, manual human curation, software program, software API)? How were these mechanisms or procedures validated?
To query (Chat)GPT, we used the GUI web interface at the URL chat.openai.com/chat; see Appendix B.6 for detailed reasons for using the GUI interface.
If the dataset is a sample from a larger set, what was the sampling strategy?
The prompts of the MATH and Symbolic-Integration subdatasets have been randomly sampled from and , across different files from those datasets. For our miniGHOSTS dataset, we sampled prompts from each of the files in GHOSTS in the following way: Our results in Section 4 indicate that the 9-January-2023 and the 30-January-2023 ChatGPT versions have similar overall performance; however, the behavior differs on a more fine-grained level and was marginally better for the 30-January-2023 version. Hence, we assembled miniGHOSTS by computing all subsets of prompts having approximately the same mean rating and standard deviation as the original file from GHOSTS, rated on the 30-January-2023 version of ChatGPT. A manual inspection of these subsets, in order to pick a subset with appropriate mathematical content (we want to have a mathematically diverse dataset), then led to the final selection of the miniGHOSTS dataset.
Who was involved in data collection process (e.g., students, crowd-workers, contractors) and how were they compensated (e.g., how much were crowd-workers paid)?
Only we have been involved in the data collection process. No payment (other than one made through regular employment) in relation to creating this dataset and writing this article was made.
Over what timeframe was the data collected? Does this timeframe match the creation timeframe of the data associated with the instances (e.g., recent crawl of old news articles)? If not, please provide a description of the timeframe.
The collection date matches the creation time. It is specified in the timestamp key in each data point from GHOSTS and spans a timeframe from January 9, 2023, to now. Using the timestamp, the version of ChatGPT that was used can be inferred, see Appendix C.
Were any ethical review processes conducted (e.g., by an institutional review board)? If so, please provide a description of these review processes, including the outcomes, as well as a link or other access point to any supporting documentation.
Does the dataset relate to people? If not, you may skip remaining questions in this section.
Did you collect the data from the individuals in question directly, or obtain it via third parties or other sources (e.g., websites)?
Were the individuals in question notified about the data collection? If so, please describe (or show with screenshots or other information) how notice was provided, and provide a link or other access point to, or otherwise reproduce, the exact language of the notification itself.
Did the individuals in question consent to the collection and use of their data? If so, please describe (or show with screenshots or other information) how consent was requested and provided, and provide a link or other access point to, or otherwise reproduce, the exact language to which the individuals consented.
If consent was obtained, were the consenting individuals provided with a mechanism to revoke their consent in the future or for certain uses? If so, please provide a description, as well as a link or other access point to the mechanism (if appropriate).
Has an analysis of the potential impact of the dataset and its use on data subjects (e.g., a data protection impact analysis) been conducted? If so, please provide a description of this analysis, including the outcomes, as well as a link or other access point to any supporting documentation.
E.4 Preprocessing, Cleaning, and/or Labeling
Was any preprocessing/cleaning/labeling of the data done (e.g., discretization or bucketing, tokenization, part-of-speech tagging, SIFT feature extraction, removal of instances, processing of missing values)? If so, please provide a description. If not, you may skip the remainder of the questions in this section.
We corrected various minor issues and inconsistencies that could arise in the process of manual evaluation, see Appendix B.5.
Was the “raw” data saved in addition to the preprocessed/cleaned/labeled data (e.g., to support unanticipated future uses)? If so, please provide a link or other access point to the “raw” data.
The output key in each JSON object contains the raw output from (Chat)GPT—unless ChatGPT used rendered LaTeX in which case our policy was to transcribe it.
Is the software used to preprocess/clean/label the instances available? If so, please provide a link or other access point.
The raw output of (Chat)GPT in the output key has not been cleaned, see Q36. Cleaning of the other values has been done first using Python scripts, in an automated way, and subsequently by hand, to correct any further, unforeseen mistakes, see Appendix B.5. The Python scripts are available upon request.
E.5 Uses
Has the dataset been used for any tasks already? If so, please provide a description.
We have used the GHOSTS dataset to evaluate and compare the mathematical capabilities of different LLMs, in particular, different (Chat)GPT versions; see Section 4.
Is there a repository that links to any or all papers or systems that use the dataset? If so, please provide a link or other access point.
Future work citing the GHOSTS dataset will be listed by citation trackers such as Google Scholar and Semantic Scholar.
What (other) tasks could the dataset be used for?
If the dataset is growing further, we anticipate that GHOSTS can be used as training data for fine-tuning LLMs.
Is there anything about the composition of the dataset or the way it was collected and preprocessed/cleaned/labeled that might impact future uses? For example, is there anything that a future user might need to know to avoid uses that could result in unfair treatment of individuals or groups (e.g., stereotyping, quality of service issues) or other undesirable harms (e.g., financial harms, legal risks) If so, please provide a description. Is there anything a future user could do to mitigate these undesirable harms?
Are there any tasks for which the dataset should not be used? If so, please provide a description.
E.6 Distribution
Will the dataset be distributed to third parties outside of the entity (e.g., company, institution, organization) on behalf of which the dataset was created? If so, please provide a description.
Yes, the GHOSTS dataset will be made publicly available. Some prompts will not be available due to copyright issues (see Appendix B.3), but a precise reference where the original prompt can be found will be included instead.
How will the dataset be distributed (e.g., tarball on website, API, GitHub) Does the dataset have a digital object identifier (DOI)?
The dataset will be made available on GitHub in the public repository github.com/xyfrieder/science-GHOSTS as a collection of JSON files.
Will the dataset be distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)? If so, please describe this license and/or ToU, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms or ToU, as well as any fees associated with these restrictions.
We release the GHOSTS and miniGHOSTS datasets under the following Creative Commons license: Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), unless we are bound by licenses of individual prompts or files from various subdatasets to release those prompts or files under more restrictive licenses; see Appendix B.3 for more information.
Have any third parties imposed IP-based or other restrictions on the data associated with the instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms, as well as any fees associated with these restrictions.
IP-restrictions apply only to those prompts that were not solely created by the authors (which are under the CC BY-NC 4.0, as explained above), see Appendix B.3 for these cases.
Do any export controls or other regulatory restrictions apply to the dataset or to individual instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any supporting documentation.
E.7 Maintenance
Who will be supporting/hosting/maintaining the dataset?
The dataset will be hosted on a GitHub repository; see Q46. All the information about the dataset, including links to the paper and future announcements, will be written in the README file of the GitHub repository.
How can the owner/curator/manager of the dataset be contacted (e.g., email address)?
The email addresses of the authors are publicly available. Moreover, it is possible to raise an issue on GitHub.
Is there an erratum? If so, please provide a link or other access point.
Future changes will be documented in the README file of the GitHub repository. Differences in single files can be tracked in the Git history.
Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)? If so, please describe how often, by whom, and how updates will be communicated to users (e.g., mailing list, GitHub)?
We will continue creating new prompts and evaluating future versions of (Chat)GPT and other LLMs. We are considering either allowing pull requests in order to encourage the community to contribute to our dataset (these requests would be carefully reviewed by us) or setting up a website to accommodate future updates. In the case of proceeding with GitHub hosting, after a significant amount of changes to the dataset, we intend to release new versions (potentially based on Git tags) and document them in the README file of the GitHub repository. By default, subscribers would then receive notifications when new releases are published in the repository.
If the dataset relates to people, are there applicable limits on the retention of the data associated with the instances (e.g., were individuals in question told that their data would be retained for a fixed period of time and then deleted)? If so, please describe these limits and explain how they will be enforced.
Will older versions of the dataset continue to be supported/hosted/maintained? If so, please describe how. If not, please describe how its obsolescence will be communicated to users.
Yes, older versions will be available in the GitHub history, and corresponding commits will be documented in the README file.
If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so? If so, please provide a description. Will these contributions be verified? If so, please describe how. If not, why not? Is there a process for communicating/distributing these contributions to other users? If so, please provide a description.
Any external contribution to our dataset is strongly encouraged. Every addition to the dataset will be carefully reviewed by the authors. For other details, please see Q55.