HuatuoGPT, towards Taming Language Model to Be a Doctor

Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, Xiang Wan, Benyou Wang, Haizhou Li

Introduction

Medicine stands as a paramount pillar in human existence, encompassing profound significance. Medicine relies heavily on experiential knowledge, wherein seasoned physicians outperform their novice counterparts. However, the advent of generative artificial intelligence (AI) systems, such as ChatGPT and DALLE, which also can learn from past experiences and external sources, heralds a transformative era for experience-driven professions. It is increasingly evident that intelligent (or, say, ‘data-driven’) medicine is an inexorable trend destined to materialize soon, albeit with ethical quandaries that demand consideration.

Medicine is a profoundly human endeavor where language plays a crucial role in facilitating interactions among clinicians, researchers, and patients. Coincidentally, the emergence of large language models (LLMs) in artificial intelligence is language-driven. This presents a remarkable opportunity for LLMs to contribute significantly to medicine. By bridging the gap between Medicine and LLMs, referred to as LLM for Medicine or LLM4Med, large language models can bring about transformative changes in human lives. One such impact is the ability to provide equitable access to high-quality medical resources to people worldwide through online means. This aligns with the original vision of the internet era and fulfills the aspirations of AI.

It is distressing to envision the thousands of lives lost each day, particularly in underdeveloped areas, due to the unavailability of medical resources, untimely medical care, or exorbitant medical costs. Given the substantial disparities in medical resources across countries and even within a single country, LLMs for Medicine have the potential to address these imbalances and promote equality among all human beings.

The short answer is ‘NO’. According to the recent study , it has been observed that ChatGPT, and even GPT-4, exhibit relatively poorer performance in vertical domains such as medicine. One contributing factor to this phenomenon is the potential lack of proficiency in medical knowledge among annotators. Consequently, there exist significant opportunities for further exploration and improvement in this domain.

On the other hand, online medicine often presents customized and localized challenges. For instance, Chinese medicine differs fundamentally from Western medicine, as does Indian medicine and many others. However, ChatGPT, being a general language model, lacks the capability for extensive customization. Additionally, entrusting private companies with users’ medical data raises concerns, emphasizing the need for private deployment to ensure local data storage. Developing a medical ChatGPT that is fully open-sourced and commercially viable would be advantageous for the well-being of individuals.

The intended purposes of LLM4Med could be medical and health advice, triage, diagnosis, prescribing drugs, interpretation of medical reports, etc. In general, any medical or health information could be consolidated into an online chat process, similar to utilizing ChatGPT. Online medical consultation offers numerous advantages, including:

Cost-effectiveness: The marginal cost of serving multiple users in an online manner is not linearly proportional to that of serving a single user. This scalability allows for cost-efficient expansion once the model is trained.

Reducing hospital crowding: The recent pandemic highlighted the risks associated with overcrowded hospitals, as many individuals sought offline consultations even when not requiring immediate medical treatment. By providing online alternatives, the strain on hospitals can be alleviated in order to mitigate the risks of future pandemics.

Addressing psychological barriers: Some individuals may refrain from seeking medical help or treatment due to fear or superstition, a phenomenon known as ‘讳病忌医’ in Chinese. Online chatting platforms may provide a more comfortable environment for such individuals to discuss their concerns.

As widely recognized, healthcare inequality in China is a significant issue. Disparities in medical conditions between residents in first-tier cities and those in small cities and rural areas are striking. For instance, the average life expectancy in Shanghai stands at approximately 82 years, whereas in regions such as Guizhou, characterized by relative economic disadvantage, life expectancy drops significantly to 73 years.https://en.wikipedia.org/wiki/List_of_Chinese_administrative_divisions_by_life_expectancy

Here, we present a new Chinese medical LLM called ‘HuatuoGPT’ to commemorate the renowned Chinese physician Hua Tuohttps://en.wikipedia.org/wiki/Hua_Tuo. Rather than training from real-world medical data as many previous language models did, a straightforward way is to distill from ChatGPT as it could quickly equip a language model with fluent chat and well-formatted responses. However, distilling from ChatGPT in medical domain is problematic, since the teacher model (i.e. ChatGPT) has the following issues:

ChatGPT does not perform well in medical domain, especially in Chinese .

ChatGPT refuses to diagnose and prescribe drugs due to ethical and safety issues.

ChatGPT does not perform as a doctor does. For example, it never ask questions even though the patients’ situation is incomplete for medical decision-making while doctors usually ask for further details. In this case, ChatGPT gives a general response instead of a specialized one.

ChatGPT struggles with hallucination due to the auto-regressive fashion.

To overcome the above issues, the core recipe of HuatuoGPT is to leverage both real-world data from doctors and distilled data from ChatGPT in the Supervised Fine-Tuned (SFT) stage; both data consist of medical instruction data and medical conversation data . The distilled data from ChatGPT is used to tame language models to follow medical instructions and talk fluently. The additional real-world medical data not only inject medical knowledge into language models but also tame the language models to perform medical diagnoses or prescribe medications, act like a doctor and provide accurate information. The complementarity between real-world medical data and distilled data is further discussed in Sec. 2.

To leverage the strengths of both data (i.e., the real-world and distilled data) and meanwhile mitigate their weaknesses, we design a well-defined RL from AI Feedback (RLAIF) method after the SFT stage. It is used to reward the generated responses that are not only patient-friendly (learned from ChatGPT with better presentation quality, lengthy and informative contents, instruction-following abilities and fluent chat), but also doctor-like (learned from doctors with professional and interactive diagnosis.). Technically, we employ LLMs to score generated responses based on their correctness, richness, logical consistency, and diagnostic ability to align our model with the both merits of ChatGPT and doctors.

In assessing the performance of our model in the medical consultations, we meticulously crafted an evaluation schema encompassing both automated and manual assessments. HuatuoGPT, when assessed using GPT-4 in automatic evaluations on a series of 100 questions sourced from CBLUE with ten distinct medical intents, consistently outperformed incumbent Chinese medical models. More impressively, our model surpassed the performance of GPT-3.5-turbo in a majority of the evaluated cases. For the more complex multi-turn conversation evaluations, our HuatuoGPT model notably outshone ChatGPT in over 60% of the instances in 20 departments, showcasing our proficiency in fusing real-world and distilled data and effectively applying reinforcement learning techniques to them. Furthermore, HuatuoGPT also achieved state-of-the-art (SOTA) performance in several medical benchmarks such as CmedQA, webmedQA, and Huatuo26M datasets.

To ensure the integrity and precision of our assessment, we incorporated manual evaluations of our model’s performance in both single-turn and multi-turn conversation scenarios. The results from these manual evaluations corroborated the findings from our automated evaluations, thus reinforcing the reliability and consistency of our model’s performance.

The contributions of HuatuoGPT are manyfold:

HuatuoGPT is the first medical language model to use RLAIF to leverage the merits of both real data and distilled data (including instruction and conversation data).

This is among the first work that conducts systematic evaluation in medical LLMs.

Human evaluation shows that HuatuoGPT outperforms existing open-sourced LLMs and ChatGPT(GPT-3.5-turbo). Its performance is most similar to that of a doctor.

We open-source our training data, code, HuatuoGPT model and the reward model at https://github.com/FreedomIntelligence/HuatuoGPT.

Motivations

Training language models from purely real-world conversation was a common practice . However, this suffers from low-quality data. For example, the responses in real-world conversations might be uninformative, short, and poorly presented. More importantly, the values in these data are not aligned and even contradictory. Learning from purely humans usually result in an unsatisfied chat-based language model compared to ChatGPT.

Recent work tends to distill a language model from ChatGPT, either imitating ChatGPT responses from single-turn instructions or learning the ChatGPT responses when interactively chatting with humans . By distilling output from ChatGPT, a model can quickly acquire impressive instruction-following capabilities and seamless dialogue skills. In addtion, characterized by its diversity and rapid generation, ChatGPT-distilled data can span various medical dialogues, encompassing various diseases, symptoms, and treatment modalities. This breadth and diversity substantially enhance the predictive performance and generalizability of the model.

2 Learning From Both Doctors and Chatgpt in Medicine

However, distillation from ChatGPT might not work for medical LLMs since there exists a fundamental gap between ChatGPT responses and doctor responses, as shown in Figure 1 and Table 1. The quality of distilled data can fluctuate, manifesting as incorrect or ambiguous information in the generated conversations. Contrastingly, real-world data, harvested from authentic doctor-patient interactions, provide an indispensable perspective into the complexities of actual medical scenarios. It can accurately reflect the true intention distribution of patients and has accurate diagnoses from doctors. The primary strength of real-world data lies in its high accuracy and professionalism.

When consulting with doctors about our medical conditions, their responses typically exhibit professionalism that meets the personalized consultation. They are adept at inquiring about the symptoms and providing accurate diagnoses. However, due to time constraint Interestingly, ChatGPT does not has a sense of time and life, it does not need to save time., their replies are often informal and concise in nature, and sometimes incoherent. Our preliminary study shows that training from purely patient-doctor interaction data is not satirised: 1) it cannot fluently follow diverse instructions or chats ; 2) the responses are short, poorly-presented, and sometimes uninformative, which are not patient-friendly.

On the other side, although ChatGPT usually generates informative, well-presented and logical responses, it usually tends to enumerate multiple possibilities and provides general and high-level advice. Since ChatGPT does not raise questions and guide patients to describe their symptoms, it lacks patients’ input that can be used to generate specialized responses. In general, its responses often lack the contextual understanding that a doctor possesses, resulting in abstract responses that offer little substantial help to patients. In Conclusion, ChatGPT does not perform like doctors that conduct interactive diagnosis.

3 Our Solution

Considering these challenges, we propose to combine the strengths of both distilled data (from ChatGPT) and real-world data (from Doctors), as illustrated in Table 2. The objective is to tame the medical LLM to perform like doctor. For example, it is expected to not only provide detailed, informative, and well-presented content but also conduct accurate and interactive diagnostic (usually posing clarifying questions) like doctors. To this end, our approach first mix distilled and real-world data in the Supervised Fine-Tuning stage (SFT). Furthermore, we employ RL from AI Feedback (RLAIF) to leverage the strengths of both data and meanwhile mitigate their weaknesses.

Methodology

Our approach focuses on integrating the characteristics of both doctor and ChatGPT to enhance the quality of responses in medical consultations through a two-stage training strategy: SFT with hybrid data and RL with AI feedback. We first utilize well-selected hybrid data to train the model through supervised fine-tuning and subsequently reinforce the generation of desired responses through feedback from AI, as illustrated in Figure 2.

In the first stage, we employ a blend of distilled data and real-world data, capitalizing on both strengths to endow the model with Doctor-like and Patient-friendly characteristics. Within each data category, we have collected instruction data and conversation data to imbue the model with the capacity for instruction-following and interactive diagnosis.

We follow the work of self-instruct to construct a set of medical instruction data aiming to enable the model to follow user’s medical instructions. The difference is that we have employed top-down manner to create more natural and comprehensive responses. We design a taxonomy to collect or manually create seed instructions based on the roles and use cases. Based on each role or use case, we generate instructions separately using self-instruct . This could provide a wide range of instructions and meanwhile keep enough instructions for each role or use cases. Finally, we mix all seed instructions together and conduct self-instruct; this might be helpful to generate more diverse instructions. Details refer to Appendix A.1.

Real-world instruction data are derived from question-answering between doctors and patients. Responses from doctors are expertise, with high relevance and conciseness. Therefore, we further enhance the quality and reliability of the single-turn instruction data by refining authentic doctor-patient question-answer pairs. Details refer to Appendix A.2.

Distilled conversations are generated by two ChatGPTs, each ChatGPT is associated with a role (either doctor or patient) using a well-designed prompt. First, we leverage a third-party medical diagnosis database as a valuable source of medical knowledge and expertise for generating synthetic dialogue data. Based on the basic background of patients and the final diagnosis from doctors, two ChatGPTs are asked to generate dialogue utterances one by one. In these conversations, the responses generated by LLMs usually are informative, detailed, well-presented, and adhere to a consistent style; the format and information are usually friendly to patients. Details refer to Appendix A.3

Real-world conversations are collected from genuine scenarios, where doctors’ responses often demand diverse abilities, including long-range reasoning and raising questions to guide patients in describing their symptoms. However, this type of data sometimes suffers from being overly concise and too colloquial. To address this, we utilized language models to enhance and refine the data based on the original content, which yields a high-quality real conversation dataset. Details refer to Appendix A.3

2 RL with AI Feedback

In the Supervised Fine-Tuning (SFT) phase, we introduced a diverse dataset with the aim of enabling HuatuoGPT to emulate the inquiry and diagnosing strategy of doctors, while maintaining the rich, logical, coherent characteristics of LLMs’ responses. In order to further align the model’s generation preferences to our needs, we propose reinforcement learning with AI feedback to improve the quality of models’ responses. Previously, OpenAI introduced reinforcement learning with human feedback to align LLMs with human preference but at a significant time and labor cost. demonstrated that with a carefully designed prompt, AI is able to imitate human preferences and to give relatively consistent scores on generated responses. Inspired by these alignment methods, we design a new pipeline to force the model to generate informative and logical responses without deviating from doctor’s diagnosis.

We train a reward model to align with the characteristics of doctors and LLMs. We use real instructions and conversations as training data, sampling multiple responses from our fine-tuned model. For multi-turn conversations, we provide the dialogue history to align our model’s response generation. These responses are then scored by an LLM, such as ChatGPT, considering informativeness, coherence, adherence to human preferences, and factual accuracy based on given real doctors’ diagnoses. The scoring LLM evaluates each response and assigns a score. We use this paired response data to train the reward model, using the fine-tuned model as its backbone for better generalization.

In RL process, we sample kk different responses {y1,…,yk}\{y_{1},\dots,y_{k}\} of a given query xx by current policy π\pi. Each response yiy_{i} is fed to our reward model to provide a reward score rRMr_{RM}. To ensure that the model does not deviate too far from the initial state π0\pi_{0}, we add the empirically-estimated KL penalty term, and the final reward function is as follows:

where λKL\lambda_{KL} is a hyperparameter for KL penalty, DKLD_{KL} is the KL penalty function. In our experiment, λKL\lambda_{KL} is set to 0.050.05. Input queries are de-duplicated and sampled from the remaining SFT hybrid data. This ensures a diverse range of inputs while retaining the model’s response preferences in both the single-turn instruction and the multi-turn conversation scenarios.

Experiments

In this section, we first introduce the training implementation (Section 4.1) and then present the evaluation manners and results including automatic evaluation (Section 4.2) and manual evaluation (Section 4.3).

Our model is implemented in PyTorch using the Acceleratehttps://huggingface.co/docs/accelerate/index and trlxhttps://github.com/CarperAI/trlx packages with Bloomz-7b1-mt as the base architecture.We adopt BLOOMZ-7b1-mt as our backbone, which stands out with its exceptional multilingual capabilities and suitable for open-source applications. BLOOMZ model family is trained with the PILE corpus , which contains varied medical texts, including resources like PubMed Central and PubMed Abstracts. These valuable texts significantly enrich the BLOOMZ models with an extensive body of medical knowledge, subsequently enables our models to perform better in the medical domain. We leverage ZeRO-3 to distribute the model across 8 A100 GPUs for training. In the supervised fine-tuning process, we set the learning rate, batch size, and maximum context length to 2e−52e-5, 128128, and 20482048, respectively. All models are trained for 3 epochs and weights performed the best on the validation set are saved. During the reinforcement learning process, we only update the parameters of the last two layers. The total number of steps is 16,00016,000, with a learning rate of 8e−68e-6. In addition, to enhance the model’s conversational and instruction-following capabilities in the general domain, we have incorporated Chinese instruction data (the Chinese Alpaca dataset and conversation data (ShareGPThttps://huggingface.co/datasets/philschmid/sharegpt-raw). This enhances the model’s ability to effectively understand and generate responses in various conversational scenarios and accurately follow instructions across different domains.

2 Automatic Evaluation

We select three existing Chinese medical QA datasets as examples, namely cMedQA2 , webMedQA and Huatuo-26M , and compare the results with the existing baselines. cMedQA2 is a publicly available dataset based on Chinese medical questions and answers consisting of 108,000 questions and 203,569 answers. webMedQA is a real-world Chinese medical QA dataset collected from online health consultancy websites consisting of 63,284 questions. Huatuo-26M is the largest Chinese medical QA dataset which has 26M QA pairs from online medical consultation, knowledge bases and encyclopedias.

Following the previous works , we utilize evaluation metrics such as BLEU, ROUGE, GLEU, and Distinct. BLEU computes the k-gram overlap between generated and reference sentences to measure similarity. ROUGE-N assesses the N-gram overlap, and ROUGE-L gauges the longest common subsequence of word matches. GLEU auto-evaluates sentence-level fluency. Distinct-1/2 aids in assessing textual diversity of the generated response by determining distinct n-grams count. However, these reference-based metrics may not suit medical QA scenarios due to diverse potential reference answers; more sound metrics should be paid more attention.

We compare our model to the best reported zero-shot model ChatGPT (GPT-3.5-turbo) and an in-domain fine-tuned model Chinese T5 https://huggingface.co/imxly/t5-pegasus respectively, which is continuously trained for 1 epoch on the full training set using batch-size 8, with a learning rate of 10−410^{-4} using Adam, linear scheduling with a warm-up rate of 0.1.

HuatuoGPT demonstrates impressive performance across various Chinese medical benchmarks, achieves consistently high scores across all metrics, and demonstrates a high level of accuracy, fluency, and diversity in its generated responses. In cMedQA2 and webMedQA, HuatuoGPT even outperforms fine-tuned T5, suggesting that it has a robust generalization capability and is able to effectively handle a wide range of medical question-answering tasks.

2.2 Evaluation with GPT4

We conduct an automated evaluation on single-turn questions with different intents and multi-turn conversations from different departments to observe the performance of the model in various scenarios.

For the single-turn questions, we extract 100 questions representing 10 intents (condition diagnosis, etiological analysis, treatment plan, medical advice, indicators interpretation, disease description, consequences description, precautions, efficacy, medical expenses) from the validation set of the Knowledge-based Universal Automated Knowledge Extraction for Query Intent Classification (KUAKE-QIC) in Chinese Biomedical Language Understanding Evaluation (CBLUE )https://github.com/CBLUEbenchmark/CBLUE. KUAKE-QIC is collected from search engine queries, which makes it suitable for single-turn questions. To filter the noisy data, these questions were initially scored by ChatGPT, and a manual filtering process was conducted to select higher quality candidate questions for the test set. For the multi-turn questions, we used the patient cases from . We selected 20 departments and randomly sampled 5 patient cases from each department, resulting in a total of 100 real patient cases. These cases were provided to ChatGPT, which played the role of the patient, interacting with each doctor model to obtain the diagnosis results.

We use GPT-4 as the referees to review the quality of model outputs. We prompt it to consider doctor-like language, symptom inquiry capability, the effect and reliability of the treatment recommendations and prescriptions, and the helpfulness to the patient. Given the question and the corresponding two answers from two models, GPT-4 is asked to first compare the advantages of each output and analyze the helpfulness to the patient, then it is requested to provide a score to each response respectively. In this way, we can get the evaluation scores of the 100 questions for each model comparison pair. We take the average scores over all the questions and calculate the performance ratio for each compared model (i.e. the overall score of the compared model divided by that of HuatuoGPT in a comparison pair).

We mainly compare HuatuoGPT to the two most popular general models ChatGPT and GPT4 ChatGPT and GPT-4 version is the online one on 12th May 2023, and the two most representative open-source Chinese medical large language models: BenTsao (tuned from LLaMA)https://github.com/SCIR-HI/Huatuo-Llama-Med-Chinese/tree/main, DoctorGLM (tuned from ChatGLM)https://github.com/xionghonglin/DoctorGLM. For single-turn questions evaluation, we compare to all the mentioned four models. For multi-turn conversations evaluation, we only compare our model to DoctorGLM and GPT-3.5-turbo due to the quote limit of GPT-4. We report the performance ratio of all models over all single-turn questions and multi-turn conversations respectively.

For the single-turn questions evaluation, all the model performance results are shown in Table 5. The comparison among models for each category is shown in Figure 3 and the comparison among their overall performance is shown in Figure 4, where the performance of HuatuoGPT is set to 1.0. According to GPT-4, HuatuoGPT is much better than BenTsao and DoctorGPT in all categories. Compared to GPT-3.5-turbo, HuatuoGPT outperforms it in three categories (Indicators Interpretation, Condition Diagnosis, and Medical Expenses) and performs similarly to it in two categories (Efficacy and Disease Description). However, HuatuoGPT is still worse than GPT4 in almost all categories, where it attains similar performance to GPT4 in two categories (Efficacy and Medical Expenses). Overall, HuatuoGPT achieves higher scores than DoctorGLM, BenTsao, and GPT-3.5-turbo. For the multi-turn conversations evaluation, similarly, the overall performance of HuatuoGPT surpasses GPT-3.5-turbo in over 60% of cases. The comparison for each category and for overall performance are shown in Table 6, Figure 5, and Figure 6 respectively.

3 Manual Evaluation

We utilize the 100 KUAKE-QIC questions (the same as those in automated evaluation) as the test set for single-turn question evaluation and randomly sample 50 patient cases from 100 test cases used in automated evaluation for multi-turn conversations manually evaluation.

In the manual evaluation of the HuatuoGPT, we think that the following three aspects should be considered, particularly in medical consultation and medication prescription and take them as the guidelines for evaluation:

Diagnosis accuracy. This aspect evaluates the model’s accuracy and comprehensiveness in diagnosing patient symptoms. Evaluators are provided a set of medical cases or symptom descriptions and assess the correctness, relevance, and reasonableness of the model’s diagnosis. Comparisons can be made with assessments made by medical professionals to ensure the model’s accuracy.

Treatment recommendation accuracy. This aspect assesses the accuracy and appropriateness of the model’s treatment recommendations for patients. Evaluators are provided a set of medical cases or symptom descriptions and evaluate whether the model’s treatment recommendations align with medical knowledge and real-world applications that are effective and reliable to the patient’s main condition and problem.

Medication knowledge and prescription accuracy. This aspect evaluates the model’s understanding of medications and the accuracy of its prescription recommendations. Evaluators are provided a set of medical cases or symptom descriptions and assess the accuracy and reliability of the medication recommendations based on medical knowledge and guidelines.

We provide physicians with above considerations, enabling them to align their evaluation guidelines. This allows for a meticulous comparison of the good and bad outputs of different models for the same scenario.

During evaluation, medical experts are asked to provide assessments on different responses. Each physician is solely responsible for evaluating the output of a single pair of models, ensuring that each response data is scrambled and anonymized with the utmost strictness. Consistent with automatic evaluation, we take BenTsao, DoctorGLM, ChatGPT and GPT4 as the baselines in single-turn question evaluation and select DoctorGLM, ChatGPT as the baselines in multi-turn conversation evaluation.

3.2 Results

As shown in Table 7, HuatuoGPT performs exceptionally well against BenTsao, DoctorGLM, and it even slightly outperforms ChatGPT, highlighting its robust diagnosis accuracy, treatment recommendations, and medication knowledge. The results of multi-turn evaluation are shown in Table 8, which reveals that HuatuoGPT excels in extended dialogue contexts, evidenced by an 86% win rate against DoctorGLM and 58% against ChatGPT. It indicates that HuatuoGPT has a more prominent interactive diagnostic capability in patient consultation scenarios.

Discussion

In this section, we explore the impact of two types of data on the model. We trained two distinct models, namely HuatuoGPT (w/ real data) and HuatuoGPT (w/ distilled data), using exclusively real-world data or distilled data, respectively. We thoroughly compare the variations in responses between the two models for the same set of questions as shown in Table 9. HuatuoGPT (w/ real data) has a tendency to ask clarifying questions to patients, performing as expected, similar to a doctor. However, a minor flaw is that the response is brief and the content appears less well-organized for reading. On the other hand, HuatuoGPT (w/ distilled data) generates well-organized, detailed, and informative content. Nevertheless, its responses are more focused on providing suggestions rather than making a diagnostic decision. Thus, HuatuoGPT (w/ distilled data) resembles a "non-doctor friend" rather than a doctor.

To assess the impact of RLAIF (Reinforced Learning with Auxiliary Information Feedback), we also compare two models: the default model called HuatuoGPT and a variant called HuatuoGPT (w/o RLAIF) which does not utilize RLAIF. It is worth noting that the latter model, HuatuoGPT (w/o RLAIF), did not ask additional questions to patients. This might be attributed to the fact that its training data could be biased towards the ChatGPT data, while real-world data may have been overlooked. In contrast, our default model, HuatuoGPT with RLAIF, can function like a doctor by asking follow-up questions to patients to get more accurate diagnoses.

2 Limitation

We emphasize the potential risks associated with generation-based medical consultation. The main concern lies in the challenge of verifying the accuracy and correctness of the generated content. In the medical domain, the dissemination of misleading information can have severe ethical implications. Although generative QA has shown promise, especially with the success of models like ChatGPT, they are not yet fully prepared for real-world deployment in the biomedical domain.

While generation methods currently hold great potential, it is important to exercise caution and prudence before deploying them in real-world applications. Further research and development are necessary to refine these models, enhance their accuracy, and establish robust mechanisms for accurateness-checking and error correction. Only through careful scrutiny and continual improvement can we minimize the risks and ethical concerns associated with generation-based medical QA.

LLMs in Medicine

The language model in the medical field has always been a concern for researchers. The early models were mainly based on the GPT-2 series models to continue pre-training in the domain. BioMedLMhttps://www.mosaicml.com/blog/introducing-pubmed-gpt is a domain-specific large language model for biomedicine, trained from 2.7B GPT-2. It is trained on the PubMed Abstracts and PubMed Central portions of the Pile dataset, which contains around 50B tokens and spans a collection of 16 million abstracts and 5 million full-text articles from the biomedical literature. Similarly, BioGPT is a medium GPT-2 model pre-training in medical data collected from the official PubMed website https://ftp.ncbi.nlm.nih.gov/pubmed/. For downstream tasks, it uses the soft prompt for fine-tuning training.

Recently, many efforts have attempted to use instruction fine-tuning to enhance the ability for medical consultation on large-scale language models (>6B), as shown in Table 2. MEDALPACAhttps://github.com/kbressem/medAlpaca is a LLaMA model trained on the Medical Meadow, consisting of two main categories, a collection of established medical NLP tasks reformatted in instruction tuning formats, as well as a crawl of various internet resources. ChatDoctor https://github.com/Kent0n-Li/ChatDoctor is also a medical LLM trained on the HealthCareMagic-100k dataset based on the LLaMA model. The HealthCareMagic-100k dataset consists of 100k real-world patient-physician conversations from an online medical consultation site. ChatDoctor has autonomous knowledge retrieval capabilities by accessing real-time and authoritative information and answering patient questions based on databases such as Wikipedia to improve the accuracy of the model’s response. Baize-healthcarehttps://huggingface.co/project-baize/baize-healthcare-lora-7B is a variant of Baize that is fine-tuned on Medical data (Quora Dialogs and Medical Dialogs). The technique report associated with it has not been published, resulting in limited details being available, as only the model weights were released. Visual Med-Alpacahttps://github.com/cambridgeltl/visual-med-alpaca is fine-tuned on LLaMA-7B model using a model-generated dataset comprising of manual filtering 54,000 biomedical examples for instruction-tuning purposes, plus the fine-tuned Microsoft GIT model on the Radiology Objects in Context (ROCO) dataset to incorporate visual modality. Recently, Med-PaLM2 was published, which is based on PaLM2 and finetuned in MultiMedQA for Expert-Level Medical Question Answering.

In Chinese, DoctorGLM https://github.com/xionghonglin/DoctorGLM is a Chinese Medical LLM trained on Multiple Medical QA datasets based on ChatGLM. It utilizes the training data from ChatDoctor through translation and incorporates Chinese medical dialogues encompassing five departments’ QA and MedDialog chat data as part of the training data. BenTsao https://github.com/SCIR-HI/Huatuo-Llama-Med-Chinese is a knowledge-enhanced Chinese Medical LLM trained on over 8K instructions. The instruction is generated from CMeKG https://github.com/SCIR-HI/Huatuo-Llama-Med-Chinese by ChatGPT API. MedicalGPT-zh is a Chinese medical general model based on ChatGLM-6B LoRA with 16-bit instruction fine-tuning. The dataset for training the model was obtained from Chinese medical knowledge question-and-answer pairs and clinical guideline texts from 28 medical departments.

Conclusion

In conclusion, this paper presents a comprehensive approach to training a reliable and conversational healthcare model by leveraging complementary data sources and incorporating AI model feedback through reinforcement learning. The proposed approach addresses the limitations of relying solely on real or synthetic data and allows for the creation of a model that combines the strengths of both sources. By continuously refining its responses based on feedback, the model can improve its conversational abilities while maintaining the reliability necessary for healthcare applications. Further research in this area holds significant potential for advancing the field of AI in healthcare and improving patient outcomes.

Acknowledgements

We thank Prof. Zhi-Quan Luo and Dr. Ping Li for their support in SRIBD.

References

Appendix A Methodology details

Following previous work, we use self-instruction to generated the instructions from ChatGPT with the medical seed instructions we manually build and the prompt is shown below:

Different from the original self-instruction, we generated role-enhanced instructions and it will be used to generate the output with the following prompt.

A.2 Real-world Instructions from Doctors

In the experiment, we collect real-world question answering data from web and sample a set of high quality question-answering pairs used for training. Every pair is refined by LLMs. The prompt is shown below:

A.3 Real-world Conversations with Doctors

We show prompts used for patient LLM and doctor LLM. Prompt for patient LLM:

A.4 Prompt for AI feedback