Differentiate ChatGPT-generated and Human-written Medical Texts
Wenxiong Liao, Zhengliang Liu, Haixing Dai, Shaochen Xu, Zihao Wu, Yiyang Zhang, Xiaoke Huang, Dajiang Zhu, Hongmin Cai, Tianming Liu, Xiang Li
Introduction
Since the advent of pre-trained language models such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers) in 2018, transformer-based language models have revolutionized and popularized NLP. More recently, (very) large language models (LLM) have demonstrated superior performance on zero-shot and few-shot tasks. Among large language models, ChatGPT is favored by users due to its accessibility, as well as its ability to produce grammatically correct and human-level answers in different domains. Since the release of ChatGPT in November 2022 by OpenAI, it has quickly gained significant attention within a few months and has been widely discussed in the natural language processing (NLP) community and other fields.
To balance the cost and efficiency of data annotation, and train a large language model which that better aligns with user intent in a helpful and safe manner, researchers used reinforcement learning from human feedback (RLHF) to develop ChatGPT. The RLHF uses a ranking-based human preference dataset to train a reward model and with this reward model, ChatGPT can be fine-tuned by proximal policy optimization (PPO) . As a result, ChatGPT can understand the meaning and intent behind user queries, which empowers ChatGPT to respond to queries in the most relevant and useful way. In addition to aligning with user intent, another factor that makes ChatGPT popular is its ability to handle a variety of tasks in different domains. The massive training corpus from the world wide web endows ChatGPT with the ability to learn the nuances of human language patterns. ChatGPT seems to be able to successfully generate human-level text content in all domains .
However, ChatGPT is a double-edged sword . Misusing ChatGPT to generate human-like content can easily mislead users, resulting in wrong and potentially detrimental decisions. For example, malicious actors can use ChatGPT to generate a large number of fake reviews that damage the reputation of high-quality restaurants while falsely boosting the reputation of low-quality competitors. This is an example that can potentially harm consumers .
ChatGPT has also demonstrated a strong understanding of high-stake domains such as medicine , including specialties such as radiation oncology. . Medical information typically requires rigorous validation. Indeed, false medical-related information generated by ChatGPT can easily lead to misjudgment of the developmental trend of diseases, delay the treatment process, or negatively affect the life and health of patients .
2 Development of Language Models
The transformer-based language models have demonstrated a strong language modeling ability. Generally speaking, transformer-based language models are divided into 3 categories: encoder-based models (e.g., BERT , Roberta , Albert ), decoder-based models (eg: GPT , GPT2 ), encoder-decoder-based models (e.g. Transformers , BART , T5 ). In order to combine biomedical knowledge with language models, many researchers have added biomedical corpus for training . Alsentzer et al. fine-tuned the publicly release BERT model on the MIMIC dataset , and demonstrated good performance on natural language inference and named entity recognition tasks. Lee et al. fine-tuned BERT on PubMed dataset and perform well on biomedical named entity recognition, biomedical relation extraction, and biomedical question-answering tasks. Based on the backbone of GPT2 , Luo et al. continue pre-training on the bio-medical dataset and show superior performance on six biomedical NLP tasks. Other innovative applications include AgriBERT for agriculture, ClinicalRadioBERT for radiation oncology and SciEdBERT for science education .
In recent years, decoder-based LLM has demonstrated excellent performance on a variety of tasks . Compared with previous language models, LLM contains a large number of trainable parameters, such as GPT3 contains 175 billion parameters. The increased model size of GPT-3 makes it more powerful than previous models; boosting its language ability to near human levels . The ChatGPT belongs to the GPT-3.5 series, which fine-tuned its base on RLHF. Research shows that ChatGPT achieves a passing score equivalent to that of a third-year medical student on a medical question-answering task.
3 Potential risks of using ChatGPT
There is more and more content generated by ChatGPT on the Internet. However, when using ChatGPT, some potential risks need to be considered. First of all, ChatGPT may limit human creativity. ChatGPT has the ability to debug code or write essays for college students. It is important to consider whether ChatGPT will generate unique creative work, or simply copy content from their training set. New York City public schools have banned ChatGPT.
Secondly, ChatGPT with the ability to produce a text of surprising quality which can deceive readers, and the end result is a dangerous accumulation of misinformation . StackOverflow, a popular platform for coders and programmers, banned the use of ChatGPT-generated content. Because the average rate of correct answers from ChatGPT is too low and could cause significant harm to the site and the users who rely on it for accurate answers.
Thirdly, ChatGPT lacks the knowledge and expertise necessary to accurately and adequately convey complex scientific concepts and information. For example, human medical writers cannot yet be fully replaced because ChatGPT do not have the same level of understanding and expertise in the medical field . Additionally, human medical writers will be responsible for ensuring the accuracy and completeness of the information communicated and for complying with ethical and regulatory guidelines, however, ChatGPT cannot be held responsible.
In addition, the application of ChatGPT to generate medical texts must consider some ethical issues. First of all, training a large language model requires a huge amount of data, but the high quality of the data is difficult to guarantee, so the trained ChatGPT is biased. For example, ChatGPT can provide biased output and perpetuate sexist stereotypes . Secondly, ChatGPT may lead to private information leakage. This may be because the large language model remembers the personal privacy information in the training set . Thirdly, it involves the legal framework. Who is to be held accountable when an AI doctor makes an inevitable mistake? ChatGPT cannot be held accountable for its work, and there is no legal framework to determine who owns the rights to AI-generated work .
To prevent the misuse use of ChatGPT to generate medical texts and avoid the potential ethical risks of using ChatGPT, in this paper, we focus on the detection of ChatGPT-generated text for the medical domain. We collect both publicly-available expert-generated medical content and ChatGPT-generated content through OpenAI API. The aim of this study is twofold: 1) What is the difference between medical content written by humans and generated by ChatGPT? Can we use machine learning methods to detect whether medical content is written by human experts or ChatGPT?
In this work, we make the following contributions to academia and industry:
We construct two datasets to analyze the difference between ChatGPT and human-generated medical text. We will release these two datasets to facilitate further analysis and research on ChatGPT for researchers.
In this paper, we conducted a language analysis of the medical content written by humans and the medical content generated by ChatGPT. From the analysis results, we can grasp the difference between ChatGPT and humans in constructing medical content.
We built a variety of machine learning models to detect samples generated by humans and ChatGPT and explained and visualized the model structures.
In summary, this study is among the first efforts to qualitatively and quantitatively analyze and categorize differences between medical text from human experts and AIGC. We believe this work can spur further research in this direction and provide pathways toward responsible AIGC in medicine.
Methods
To analyze and discriminate human and ChatGPT-generated medical texts, we constructed two datasets:
medical abstract: This original dataset comes from kagglehttps://www.kaggle.com/datasets/chaitanyakck/medical-text. The medical dataset involves 5 different conditions: digestive system diseases, cardiovascular diseases, neoplasms, nervous system diseases, and general pathological conditions.
radiology report: This original dataset comes from the work of Johnson et al. , and we only select a part of radiology reports to build our radiology report dataset.
We sampled 2200 samples from the medical abstract and radiology report datasets as medical texts written by humans. In order to guide ChatGPT to generate medical content, we adopt the method of text continuation with demonstration instead of rephrase or query , with in-context learning, because text continuation can produce more human-like text. The prompts of medical abstract and radiology report datasets are shown in Figure 1. We randomly select a sample (except the sample itself) from the dataset as a demonstration. Finally, we obtained medical abstract and radiology report datasets containing 4400 samples respectively.
2 Linguistic analysis
we will perform linguistic analysis of the medical content generated by humans and ChatGPT, including vocabulary and sentence feature analysis, part-of-speech (POS) analysis, dependency parsing, sentiment analysis, and text perplexity.
The vocabulary and sentence feature analysis illuminates the differences in the statistical characteristics of the words and sentences constructed by humans and ChatGPT when generating medical texts. We use NLTK (Natural Language Toolkit) to perform POS analysis. Dependency parsing is a technique that analyzes the grammatical structure of a sentence by identifying the dependencies between the words of the sentence. We apply stanford-corenlp for dependency parsing and compare the proportions of different dependency relationships and their corresponding dependency distances. We apply a pre-trained sentiment analysis model https://huggingface.co/cardiffnlp/twitter-roberta-base-sentiment to conduct sentiment analysis for both medical abstract and radiology report datasets. Perplexity is often used as a metric to evaluate the performance of a language model, with lower perplexity indicating that the language model is more confident in its predictions. We use the BioGPT model to compute the perplexity of the human-written and ChatGPT-generated medical text.
3 Detect ChatGPT-generated Texts
The text content generated by the LLM has become popular on the Internet. Since most of the content generated by the LLM is text with a fixed language pattern and language style, when a large number of generated text content appears, it will not be conducive to human active creation, and It can also cause panic if the incorrect medical text is generated. We use a variety of methods to detect medical texts generated by ChatGPT to reduce the potential risks to society caused by improper or malicious use of language models.
First, we divide the medical abstract and radiology report datasets into a training set, test set, and validation set at a ratio of 7:2:1, respectively. Then we use a variety of algorithms to model in the training set, select the best model parameters through the validation set, and finally calculate the metrics on the test set.
Perplexity-CLS: As shown in Figure 6, the medical text written by humans with higher text perplexity than the medical text generated by ChatGPT. An intuitive idea is to find an optimal perplexity threshold to detect the medical text generated by ChatGPT. The idea is the same as GPTZero https://gptzero.me/, but our data is medical-related text, so we use BioGPT as a language model to calculate text perplexity. We find the optimal perplexity threshold in the validation set and calculate the metrics on the test set.
CART (Classification and Regression Trees): CART is a classic decision tree algorithm, the classification tree uses gini as the measure of feature division. We vectorize the samples through TF-IDF (term frequency–inverse document frequency), and for the convenience of visualization, we set the maximum depth of the tree to 4.
XGBoost : XGBoost is an ensemble learning method, and we set the maximum depth for base learners as 4 and vectorize the samples by TF-IDF.
BERT : BERT is a pre-trained language model, we fine-tune our medical text based on pre-trained BERT to detect the text generated by ChatGPT. The version of BERT we use is bert-base-cased https://huggingface.co/bert-base-cased.
In addition, we will analyze the models of CART, XGBoost, and BERT to explore what features of the text help to detect the text generated by ChatGPT.
Results
Vocabulary and sentence analysis: As shown in Table 1, From the perspective of statistical characteristics, the main difference between the human written medical text and the medical text generated by ChatGPT exists in the vocabulary and stem. Human-written medical text vocabulary size and the number of stems are significantly larger than those of ChatGPT. This suggests that the content and expression of medical texts written by humans are more diverse, which is more in line with the actual patient situation, while the texts generated by ChatGPT are more inclined to use commonly used words to express common situations.
Part-of-speech analysis: The results of POS are shown in Figure 2. The ChatGPT uses more words about NN(noun), DT(determiner), NNS(noun plural), and CC(coordinating conjunction), while using less CD(cardinal-digit) and RB(adverb).
Frequent use of NN and NNS tends to indicate that the text is more argumentative, showing information and objectivity . The high proportion of CC and DT indicates that the structure of the medical text and the relationship between causality, progression, or contrast is clear. At the same time, a large number of CDs and RBs appear in medical texts written by humans, indicating that the expressions are more specific rather than general. For example, doctors will use specific numbers to describe the size of tumors.
Dependency parsing: The results of dependency parsing are shown in Figure 3 and Figure 4. As shown in Figure 3, the comparison of dependencies exhibits similar characteristics to POS analysis, where ChatGPT uses more det (determiner), conj (conjunct), and cc(coordination ) relations while using less nummod (numeric modifier) and advmod(adverbial modifier). For dependency distance, ChatGPT with obviously shorter conj(conjunct), cc(coordination), and nsubi(nominal subject) which makes the text generated by chatGPT more logical and fluent.
Sentiment analysis: The results of sentiment analysis are shown in Figure 5. most of the medical texts written by humans or the texts generated by ChatGPT with neutral sentiments. It should be noted that the proportion of negative sentiments in humans is significantly higher than that in ChatGPT, while the proportion of positive sentiments in humans is significantly lower than that in ChatGPT. This may be because ChatGPT has added a special mechanism to carefully filter the original training dataset to ensure any violent or sexual content is removed, making the generated text more neutral or positive.
Text perplexity: The results of text perplexity are shown in Figure 6. It can be observed that whether it is a medical abstract or a radiation report dataset, the text perplexity generated by ChatGPT is significantly lower than that written by humans. ChatGPT captures common patterns and structures in the training corpus and is very good at replicating them. Therefore, the text generated by ChatGPT has relatively low perplexity. Humans can express themselves in a variety of ways, depending on the intellectual context, the condition of the patient, etc., which may make BioGPT more difficult to predict. Therefore, human-written text with a higher perplexity and wider distribution.
Through the above analysis, we can get the main differences between the human-written and ChatGPT-generated medical text, including:
Medical texts written by humans are more diverse, while medical texts generated by ChatGPT are more common.
Medical texts generated by ChatGPT have better logic and fluency.
Medical texts written by humans contain more specific values and text content is more specific.
Medical texts generated by ChatGPT are more neutral and positive.
ChatGPT has lower text perplexity because it is good at replicating common expression patterns and sentence structures.
Detect ChatGPT-generated texts
The results of detecting ChatGPT-generated medical text are shown in Table 2. Since Perplexity-CLS is an unsupervised learning method, it is less effective than other methods. The XGBoost integrates the results of multiple decision trees, so it works better than CART with a single decision tree. The pre-trained BERT model can easily recognize the differences in the logical structure and language style of medical texts written by humans and generated by ChatGPT, thus achieving the best performance.
Figure 7 is the visualization of the CART model of the two data sets. It can be seen that through the decision tree with depth 4, the text generated by ChatGPT can be detected well. We define the feature importance of XGBoost as the times of the feature appears in the node of the base learner, and the top-20 important features are shown in Figure 8. Comparing Figure 7 and Figure 8, we can see that their decision tree nodes are similar. For example, in the medical abstract data set, "outcomes", "study", "potential", "suggest", etc. are used as nodes in the CART and XGBoost models.
In addition to visualizing the global features of CART and XGBoost, we also use the toolkit of transformers-interpret https://github.com/cdpierse/transformers-interpret to visualize the local features of the samples, and the results are shown in Figure 9. It can be seen that for BERT, conjuncts are important features for detecting ChatGPT, for example: "due to", "however", "or", etc. In addition, the important features of BERT are similar to CART and XGboost. For example, "significant" and "acute" in the radiology report dataset are important features for detecting medical text generated by ChatGPT.
Discussion
In this paper, we focus on analyzing the differences between medical texts written by humans and generated by ChatGPT and design machine learning algorithms to detect medical texts generated by ChatGPT. The results show that medical texts generated by ChatGPT are more fluent and logical but more general in content and language style, while medical texts written by humans are more diverse and specific. Although ChatGPT can generate human-like text, due to the differences in their language style and content, the text written by ChatGPT can still be accurately detected by designing machine learning algorithms, and the F1 exceeds 95% .
2 Limitations
We only use ChatGPT as an example to analyze the difference between medical texts generated by LLM and medical texts written by humans. However, more advanced LLMs have emerged. It will be part of our future work to analyze more language styles generated by LLM and summarize the language construction rules of LLM.
3 Conclusions
In general, for AI to realize its full potential in medicine, we should not rush into its implementation, but advocate its careful introduction and open debate about risks and benefits. The medical field is a field related to human health and life. We provide a simple demonstration to identify ChatGPT-generated medical content, which can help reduce the harm caused to humans by ChatGPT-generated erroneous and incomplete information. Assessing and mitigating the risks associated with LLM and its potential harm is a complex and interdisciplinary challenge that requires combining knowledge from various fields to avoid its risks and drives the healthy development of LLM.