ChatGPT or Grammarly? Evaluating ChatGPT on Grammatical Error Correction Benchmark
Haoran Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, Michael Lyu
Introduction
ChatGPThttps://chat.openai.com/chat, the current “super-star” in artificial intelligence (AI) area, has attracted millions of registered users within just a week since its launch by OpenAI. One of the reasons for ChatGPT being so popular is its surprisingly strong performance on various natural language processing (NLP) tasks Bang et al. (2023), including question answering Omar et al. (2023), text summarization Yang et al. (2023), machine translation Jiao et al. (2023), logic reasoning Frieder et al. (2023), code debugging Xia and Zhang (2023), etc. There is also a trend of using ChatGPT as a writing assistant for text polishing.
Despite the widespread use of ChatGPT, it remains unclear to the NLP community that to what extent ChatGPT is capable of revising the text and correcting grammatical errors. To fill this research gap, we empirically study the Grammatical Error Correction (GEC) ability of ChatGPT by evaluating on the CoNLL2014 benchmark dataset Ng et al. (2014), and comparing its performance to Grammarly, a prevalent cloud-based English typing assistant with 30 million users daily Grammarly (2023) and GECToR Omelianchuk et al. (2020), a state-of-the-art GEC model. With this study, we aim to answer a research question:
To the best of our knowledge, this is the first study on ChatGPT’s ability in GEC.
We present the major insights gained from this evaluation as below:
ChatGPT performs worse than the baseline systems in terms of the automatic evaluation metrics (e.g., score), particularly on long sentences.
ChatGPT goes beyond one-by-one corrections by introducing more changes to the surface expression of certain phrases or sentence structure while maintaining the grammatical correctness.
Human evaluation quantitatively demonstrates that ChatGPT produces less under-correction or mis-correction issues but more over-corrections.
Our evaluation indicates the limitation of relying solely on automatic evaluation metrics to assess the performance of GEC models and suggests that ChatGPT is a promising tool for GEC.
Background
ChatGPT is an intelligent chatbot powered by large language models developed by OpenAI. It has attracted great attention from industry, academia, and the general public due to its strong ability in answering various follow-up questions, correcting inappropriate questions Zhong et al. (2023), and even refusing illegal questions. While the technical details of ChatGPT have not been released systematically, it is known to be built upon InstructGPT Ouyang et al. (2022) which is trained using instruction tuning Wei et al. (2022a) and reinforcement learning from human feedback (RLHF, Christiano et al., 2017).
2 Grammatical Error Correction
Grammatical Error Correction (GEC) is a task of correcting different kinds of errors in text such as spelling, punctuation, grammatical, and word choice errors Ruder (2022). It is highly demanded as writing plays an important role in academics, work, and daily life. Table 1 presents the illustration of different grammatical errors borrowed from Bryant et al. (2022) in a comprehensive survey on grammatical error correction. In general, grammatical errors can be roughly classified into three categories: omission errors, such as "on" in the first example; replacement errors, such as "dreamed" for "dreamt" in the second example; and insertion errors, such as "the" in the third example.
To evaluate the performance of GEC, researchers have built various benchmark datasets, which include but are not limited to:
CoNLL-2014: Given the short English texts written by non-native speakers, the task requires a participating system to correct all errors present in each text.
BEA-2019: It is similar to CoNLL-2014 but introduces a new dataset, namely, the Write&Improve+LOCNESS corpus, which represents a wider range of native and learner English levels and abilities Bryant et al. (2019).
JFLEG: It represents a broad range of language proficiency levels and uses holistic fluency edits to not only correct grammatical errors but also make the original text more native sounding Tetreault et al. (2017).
ChatGPT for GEC
We evaluate the ability of ChatGPT in grammatical error correction on the CoNLL2014 task Ng et al. (2014) dataset. The dataset is composed by short paragraphs that are written by non-native speakers of English, accompanied with the corresponding annotations on the grammatical errors. We pulled 100 sentences from the official-combined test set in the alternate folder of the dataset sequentially.
Evaluation Metric.
To evaluate the performance of GEC, we adopt three metrics that are widely used in literature, namely, Precision, Recall, and score. Among them, score combines both Precision and Recall, where Precision is assigned a higher weight Wikipedia contributors (2023a). Specifically, the three metrics are expressed as:
where , and represent the true positives, false positives and false negatives of the predictions, respectively. We use the scoring program provided by CoNLL2014 official but adapt it to be compatible with the latest Python environment.
Baselines.
In this report, we perform the GEC task on three systems, including:
ChatGPT: We query ChatGPT manually rather than using some API due to the instability of ChatGPT. For example, when a query sentence resembles a question or demand, ChatGPT may stop the process of GEC but respond to the “demand” instead. After a few trials, we find a prompt that works well for ChatGPT:
Do grammatical error correction on all the following sentences I type in the conversation.
We query ChatGPT with this prompt for each test sample.
Grammarly: Grammarly is a prevalent cloud-based English typing assistant. It reviews spelling, grammar, punctuation, clarity, engagement, and delivery mistakes in English texts, detects plagiarism and suggests replacements for the identified errors Wikipedia contributors (2023b). As stated by Grammarly, every day, 30 million people and 50,000 teams around the world use Grammarly with their writing Grammarly (2023). When querying Grammarly, we open a text file and paste all the test samples into separate paragraphs. We enable all the grammar correction in the setting and only ask it to correct the ones with correctness problems (red underline), while leaving the clarity (blue underline), engagement (green underline) and delivery (purple underline) unchanged. We iterate this process several times until there is no error detected by Grammarly.
GECToR: Besides Grammarly, we also compare ChatGPT with GECToR Omelianchuk et al. (2020), a state-of-the-art model on GEC in research, which also exhibits good performance on the CoNLL2014 task. We adopt the implementation based on the pre-trained RoBERTa model.
2 Results and Analysis
Table 2 presents the overall performance of the three systems. As seen, ChatGPT obtains the highest recall value, GECToR obtains the highest precision value, while Grammarly achieves a better balance between the two metrics and results in the highest score. These results suggest that ChatGPT tends to correct as many errors as possible, which may lead to more overcorrections. Instead, GECToR corrects only those it is confident about, which leaves many errors uncorrected. Grammarly combines the advantages of both such that it performs more stably.
ChatGPT Performs Worse on Long Sentences?
To understand which kind of sentences ChatGPT are good at, we divide the 100 test sentences into three equally sized categories, namely, Short, Medium and Long. Table 3 shows the results with respect to sentence length. As seen, the gap between ChatGPT and Grammarly is significantly bridged on short sentences. In contrast, ChatGPT performs much worse on those longer sentences, at least in terms of the existing evaluation metrics.
ChatGPT Goes Beyond One-by-One Corrections.
We inspect the output of the three systems, especially those for long sentences, and find that ChatGPT is not limited to correcting the errors in the one-by-one fashion. Instead, it is more willing to change the superficial expression of some phrases or the sentence structure. For example, in Table 4, GECToR and Grammarly make minor changes to the source sentence (i.e., “an example” to “example”, “family potential disease” to “a family ’s potential disease”), while ChatGPT modifies the sentence structure (i.e., “for family potential disease” to “in preventing potential family diseases”) and word choice (i.e., “chances” to “opportunities”). It indicates that the outputs by ChatGPT maintain the grammatical correctness, although they do not follow the original expression of the source sentences.
To validate our hypothesis, we let Grammarly to further correct the grammatical errors in the outputs of GECToR and ChatGPT. Table 5 lists the results. We can observe that Grammarly introduces a negligible improvement to the output of ChatGPT, demonstrating that ChatGPT indeed generates correct sentences. On the contrary, Grammarly further improves the performance of GECToR noticeably (i.e., +2.1 , +16.5 Recall), suggesting that there are still many errors in the output of GECToR.
Human Evaluation.
We conduct a human evaluation to further demonstrate the potential of ChatGPT for the GEC task. Specifically, we follow Wang et al. (2022) to manually annotate the issues in the outputs of the three systems, including 1) Under-correction, which is the grammatical errors that are not found; 2) Mis-correction, which is the grammatical errors that are found but modified incorrectly; it can be either grammatically incorrect or semantically incorrect; 3) Over-correction, which is the other modifications beyond the changes in the reference. We sample 20 sentences out of the 100 test sentences and ask two annotators to identify the issues. Table 6 shows the results. Obviously, ChatGPT has the least number of under-corrections among the three systems and fewer number of mis-corrections compared with GECToR, which suggests its great potential in grammatical error correction. Meanwhile, ChatGPT produces more over-corrections, which may come from the diverse generation ability as a large language model. While this usually leads to a lower score, it also allows more flexible language expressions in GEC.
Discussions.
We have checked the outputs corresponding to the results of Table 5, and observed different behaviors of ChatGPT and Grammarly. The slight improvement (i.e., +0.5 ) by Grammarly mainly comes from punctuation problems. ChatGPT is not sensitive to punctuation problems but Grammarly is, though the modifications are not always correct. For example, when we manually undo the corrections on punctuation, the score increases by +0.0015. Other than punctuation problems, Grammarly also corrects a few grammatical errors on articles, prepositions, and plurals. However, these corrections usually require Grammarly to repeat the process twice. Take the following sentence as an example,
... constructs of the family and kinship are a social construct, ...
... constructs of the family and kinship are a social constructs, ...
... constructs of the family and kinship are social constructs, ...
Nonetheless, it does correct some errors that ChatGPT fails to correct.
Conclusion
This paper evaluates ChatGPT on the task of Grammatical Error Correction (GEC). By testing on the CoNLL2014 benchmark dataset, we find that ChatGPT performs worse than a commercial product Grammarly and a state-of-the-art model GECToR in terms of automatic evaluation metrics. By examining the outputs, we find that ChatGPT displays a unique ability to go beyond one-by-one corrections by changing surface expressions and sentence structure while maintaining grammatical correctness. Human evaluation results confirm this finding and reveals that ChatGPT produces fewer under-correction or mis-correction issues but more over-corrections. These results demonstrate the limitation of relying solely on automatic evaluation metrics to assess the performance of GEC models and suggest that ChatGPT has the potential to be a valuable tool for GEC.
Limitations and Future Works
There are several limitations in this version, which we leave for future work:
More Datasets: In this version, we only use the CoNLL-2014 test set and only randomly select 100 sentences to conduct the evaluation. In our future work, we will conduct experiments on more datasets.
More Prompt and In-context Learning: In this version, we only use one prompt to query ChatGPT and do not utilize the advanced technology from the in-context learning field, such as providing demonstration examples Brown et al. (2020) or providing chain-of-thought Wei et al. (2022b), which may under-estimate the full potential of ChatGPT. In our future work, we will explore the in-context learning methods for GEC to improve its performance.
More Evaluation Metrics: In this version, we only adopt Precision, Recall and as evaluation metrics. In our future work, we will utilize more metrics, such as pretraining-based metrics Gong et al. (2022) to evaluate the performance comprehensively.