AlpaGasus: Training A Better Alpaca with Fewer Data
Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, Hongxia Jin
Instruction fine-tuning (IFT) (Longpre et al., 2023) has been recently applied as an essential continual training stage for pre-trained large language models (LLMs) to achieve instruction-following capability (Ouyang et al., 2022b; Chen et al., 2023b), which is often attributed to aligning the models’ behavior with a diverse set of human instructions and responses (Taori et al., 2023; Askell et al., 2021). The recent series of open-sourced instruction-tuned models (Taori et al., 2023; Xu et al., 2023) reveal that the alignment of better IFT data could result in better instruction-following skills. For example, GPT-4-LLM (Peng et al., 2023) (with GPT-4 (OpenAI, 2023b) as its teacher) exhibits better reasoning and math ability than Alpaca (Taori et al., 2023) (with Text-davinci-003 as its teacher), though they share the same base model LLaMA (Touvron et al., 2023), demonstrating the importance of data quality.
Although stronger teachers can usually bring further improvement by providing better IFT data, their responses inevitably include incorrect or irrelevant answers to the corresponding instructions (see examples in Fig. 2), which can be misleading or detrimental to IFT. Moreover, these data also increase unnecessary training costs. Alpaca-cleanedhttps://github.com/gururise/AlpacaDataCleaned/ is the pioneer of filtering bad data in Alpaca dataset though it requires humans fully involved in examining and filtering the data. Nonetheless, how to automatically filter out poor-quality data from IFT datasets has not been investigated yet. A primary bottleneck is that rating the data quality usually requires expensive human labor but still may not be accurate for IFT because stronger teachers are more powerful in generating eloquent but incorrect responses that are more subtle to detect by humans. When considering datasets crafted by humans, such as the Dolly dataset (Dolly, 2023), assessing quality becomes even more intricate, given that responses stem from seasoned writers.
This paper aims to bridge the gap by proposing a novel data-filtering strategy for IFT that is efficient, automatic, and accurate. Specifically, we design a prompt applied to a powerful LLM (e.g., ChatGPT) for evaluating the quality of each (instruction, input, response) tuple and then filter out the ones with scores lower than a threshold. By applying this filter to the 52k data used to train Alpaca, we find that a majority of the data suffer from low-quality issues. Using the LLM filter, IFT on a much smaller but carefully filtered subset of 9k data produces a much better model, i.e., AlpaGasus, than the original Alpaca, as shown in Fig. 1, following exactly the same training configuration of Alpaca. This also reduces the training time from 80 minutes to merely 14 minutes on 4 NVIDIA A100 (80GB) GPUs. Moreover, we validate the versatility of our method, demonstrating its effectiveness on a range of datasets(e.g., Dolly, Alpaca, GPT4LLM), base models(e.g., LLaMA-1 and LLaMA-2), and LLM filters(e.g., ChatGPT and Claude-2). This discovery is inspiring, as it shows that the data quality in IFT can outweigh the quantity. In addition, this shift towards prioritizing data quality presents a new and more efficient paradigm that can generally improve the fine-tuning of LLMs.
Our experiments include comprehensive evaluations for our AlpaGasus, incorporating free-form instruction evaluation, various benchmarks, and human studies. We select four different human-instruction test sets for evaluating instruction-following capability, including the ones used by WizardLM (Xu et al., 2023), Vicuna (Chiang et al., 2023), Koala (Geng et al., 2023), and Self-Instruct (Wang et al., 2022). Given the notable advantages that GPT-4 judge could match with both the controlled and crowdsourced human preferences ( agreement) (Zheng et al., 2023), we employ GPT-4 as our judge for the major evaluations. In the 7B and 13B model comparisons, AlpaGasus performs significantly better than Alpaca on all four test sets. To address potential concerns regarding biases in model-based evaluations, we conduct human studies and benchmark evaluations, both of which corroborate the superiority of our model compared to baseline counterparts. Furthermore, we present a fine-grained evaluation of AlpaGasus on individual tasks including Generic, Roleplay, Knowledge, and Commonsense from the Vicuna test set. The results indicate AlpaGasus exhibits advantages on a majority of the tasks.
To sum up, our data-filtering approach exhibits significant benefits in terms of scalability and automation. We also demonstrate that prudent management of training data quality can lead to substantial performance improvement and computation savings of IFT. In addition, our data selection and evaluation strategies can generalize to other instruction finetuning datasets and LLMs, thereby paving the way for a promising new research trajectory aimed at pragmatic LLM deployment.
Methodology
Unlike the recent work (Zhou et al., 2023), which relies on human labor to curate 1k high-quality instruction data that leads to a better finetuned model, we aim to avoid the expensive and time-consuming human annotations. Hence, we exploit the potential of strong LLMs to be auto-graders of the training data and then filter out the data with lower scores.
In particular, we prompt a strong API LLM, i.e., ChatGPT, to produce a score for each triplet of (instruction, input, response). The prompt is given in Fig. 3, where “dimension” denotes a user-preferred property such as helpfulness and accuracy. We then only select the triplets with scores higher than a certain threshold to fine-tune a LLaMA-series model following an existing IFT pipeline. Fig. 2 illustrates the data selection and training pipeline.
2 Data Rating and Filtering
Given an IFT dataset of triplets (instruction, input, response) with and an open-sourced LLM (e.g., LLaMA), let denote the finetuned on , our overarching goal is to select a subset such that IFT on results in a better model than .
In order to select from , we prompt an API LLM (e.g., ChatGPTWe also use claude-2 as our response quality evaluator, which can be found in Section A.2 ) as an auto-grader rating each sample by a score wherein is the rating prompt in Fig. 3. We then select whose score is above a certain threshold , i.e.,
We achieve by finetuning on using an existing IFT framework.
3 AlpaGasus: 9k Training Data Filtered from Alpaca
For “dimension” in the rating prompt shown in Fig. 3, given that “accuracy” closely aligns with human expectations of LLMs’ responses, we designate “accuracy” as the dimension for rating purposes.We defer the experiment of other dimensions, e.g., helpfulness, to the Section A.5. Correspondingly, we establish in Eq. 1 as an accuracy threshold for the subsequent experiments. The distribution of scores in relation to the 52k Alpaca dataset is presented in Fig. 4.
In particular, we choose the threshold according to the score histogram. For the Alpaca dataset with 52,002 samples, this filtering criterion leads to a subset of 9,229 samples 52k denotes 52002 samples from the original Alpaca training set and 9k represents 9229 data samples. (either randomly sampled or filtered in our experiments).
Experimental Setup
Most instruction-tuned models are evaluated on one test set that might not cover sufficient diverse instructions and thus leads to a risk of biased evaluation (Chia et al., 2023). To conduct a holistic evaluation of AlpaGasus, we curate our test sets from Self-instruct (Wang et al., 2022), Vicuna (Chiang et al., 2023), WizardLM (Xu et al., 2023), and Koala (Geng et al., 2023), which together can cover more types of instructions and reduce the evaluation bias. Details of these four test sets are provided in Table 1.
2 Baseline Models
We compare our AlpaGasus with the following four recent LLMs.
(Taori et al., 2023) is an open-sourced model developed by Stanford University through IFT of LLaMA on a training dataset of 52,002 (instruction, input, response) samples with the responses generated by Text-Davinci-003 (teacher).
is an OpenAI LLM trained with an increased emphasis on contextual understanding and response accuracy. Its proficiency in capturing complex linguistic patterns makes it a powerful teacher LLM for generating high-quality training data for finetuning LLMs such as Alpaca.
(OpenAI, 2023a) is an AI chatbot finetuned via reinforcement learning with human feedback (RLHF). It exhibits exceptional capability across a wide range of tasks and might be the most popular chatbot recently. Hence, it would be interesting to study to what extent AlpaGasus can match its performance.
(Bai et al., 2022) is an AI chatbot developed by Anthropic. It was finetuned by RLHF to align with humans’ preference on three dimensions, i.e., helpful, honest, and harmless. We use Claude-v1.1 for comparison, which is comparable to ChatGPT on the AlpacaEval (Li et al., 2023).
3 Evaluation Metrics
The evaluation of the instruction-following capability of LLMs is usually challenging due to the existence of multiple eligible responses to one instruction and the difficulty of reproducing human evaluations. In light of the recent advancements in automated evaluation (Dubois et al., 2023; Zheng et al., 2023; Chiang et al., 2023), which offer superior scalability and explainability than human studies, we also apply an API LLM (e.g., GPT-4) as the judge to evaluate and compare it with . In particular, we apply to compare the responses of and to each instruction drawn from a test set . Let and denote the two models’ responses to instruction , the judge outputs a score for each response and we aim to achieve a higher score on , i.e.,
for most . In our experiments, we include both models’ responses in the input to the judge (e.g., GPT-4), followed by an instruction to the judge, which aims to rate the responses with a score between 1 and 10. Details of the input and prompt to the judge can be found in Appendix C To address potential concerns regarding bias in the evaluation prompts, we also present results of using alternative evaluation prompts in Section A.1.
Since there exists position bias within LLM judges, which refers to a phenomenon where LLM judges have tendencies to prefer specific positions over others (Wang et al., 2018; Ko et al., 2020; Wang et al., 2023), to mitigate it, we try both orders (i.e., placing AlpaGasus’s response before/after the baseline model’s response) and define the final judge of “Win-Tie-Lose” to be:(1) Win: AlpaGasus wins twice, or wins once and draws once. (2) Tie: AlpaGasus draws twice, or wins once and loses once. (3) Lose: AlpaGasus loses twice, or loses once and draws once. To avoid cut-off responses, we allow models to generate up to 1024 tokens. For ChatGPT, Claude, and Text-Davinci-003, we set the temperature to 0.0, respectively, to reduce randomness and ensure a fair comparison.
Experimental Results
We compare AlpaGasus and Alpaca on two sizes of models in Fig. 5. They only differ in the training data: Alpaca uses all the 52k data while AlpaGasus only uses 9k data selected from the 52k. Their hyperparameters and training scripts are the same. As shown in the evaluation results, AlpaGasus significantly outperforms the original Alpaca across all four test sets. Moreover, when using LLaMA-2 as the base model, we observe consistent outcomes (See Section A.3). This consistency underscores the universality of our data filtering method, irrespective of the model choices. These findings also confirm that our training data selection approach leads to superior performance even when the selected training data are only 17.75% of the original dataset.
To investigate the efficacy of our data selection strategy, we compare AlpaGasus with LLaMA models fine-tuned on a randomly sampled subset of the Alpaca 52k data, denoted by Alpaca-9k-random in Fig. 6. Both models start from the same initial model (i.e., LLaMA) and are then finetuned on the same number of samples (i.e., 9k). They only differ in terms of the data selection criteria. In Fig. 6, we compare the two types of models under two model sizes, i.e., 7B and 13B. AlpaGasus-9k significantly outperforms Alpaca-9k-random, showing the high quality of our selected data and their importance to the performance of IFT.
2 How Much Data Should Be Filtered?
In Eq. 1, we select data with score and we set in our main experiments, which results in 9k out of the 52k data to finetune AlpaGasus. To study the impact of the threshold on IFT, we compare AlpaGasus with LLaMA finetuned on 39k data selected by applying a lower threshold of . We report the comparison results in Fig. 7. When tested on the Koala and WizardLM test sets, Alpaca-39k model outperforms the original Alpaca-52k model. However, when using the Vicuna and Self-Instruct as test sets, Alpaca-39k does not exhibit advantages over the original Alpaca-52k model. Hence, a loose criterion (a lower threshold) includes more data in the selected data and a model with comparable performance as the original Alpaca. However, it still performs poorer than AlpaGasus trained on much fewer but higher-quality data, indicating the negative impact of low-quality data to IFT.
On the other hand, high-quality data show a positive impact on IFT. To verify this, we randomly draw 3k and 6k data from the 9k data selected for training AlpaGasus and finetune two variants of AlpaGasus from LLaMA using the same training script. Fig. 8 reports the evaluation results of these variants: AlpaGasus trained on 9k data performs the best on all four test sets, indicating that more high-quality data leads to better IFT models.
According to Fig. 1, 6k high-quality data suffices to finetune LLaMA achieving similar performance as the original Alpaca.
3 Human Study
We further undertake human studies by enlisting three participants tasked with labeling the question/answer pairs. To be specific, we select 40 prompts from each test set, resulting in a total of 160 prompts. These are then presented to the participants alongside the corresponding responses generated by both AlpaGasus-13B and Alpaca-13B. The final answers are determined by majority voting. There are 63/160 wins for AlpaGasus-13B, 64/160 ties and 33/160 loses, which indicates the superiority of our AlpaGasus. Comprehensive results on each test set and user guidelines could be found in Appendix J.
4 Comparison with ChatGPT/Claude/Davinci003.
In Fig. 9, we compare AlpaGasus with text-Davinci-003, ChatGPT, and Claude. The results show that AlpaGasus-13B can achieve capacity of its teacher model, text-Davinci-003, which is used to generate the Alpaca-52k instruction data.
5 Benchmark Performance
Following InstructEval (Chia et al., 2023), we also evaluate our models on benchmark datasets, i.e., MMLU (Hendrycks et al., 2020), DROP (Dua et al., 2019) Humaneval (Chen et al., 2021), BBH (Suzgun et al., 2022), to evaluate the models’ performance. The details of the benchmark setting can be found in Appendix B. Benchmark results of our AlpaGasus are shown in Table 2, where higher values indicate better performance. AlpaGasus-7B, 13B show superiority on the 3/4 datasets, which demonstrates the effectiveness of our filtering algorithm. Another interesting finding is that the models trained with our filtered data can be better on all the benchmarks than training with randomly selected data.We observe similar performance gains of the 7B model on Dolly, and our 13B (3k) model consistently outperforms baselines, i.e., 13B(random-3k) and 13B(15k), on all four benchmark datasets, which are deferred to the Appendix B.
Human-written instruction set filtering
In addition to filtering machine-generated datasets, our approach is capable of filtering human-written datasets. Specifically, we investigate the Databricks-dolly-15k dataset (Dolly, 2023), a seminal collection of 15,000 high-quality human-generated prompt/response pairs. Notably, this unparalleled dataset is a product of the collective efforts of more than 5,000 Databricks contributors and the included prompts and responses are more than just simple text; they embody a comprehensive spectrum of human cognition, covering activities from inventive brainstorming to succinct summarization.
We also applied a threshold of for data filtration, resulting in a filtered dataset of 2,996 samples. (Score distribution can be found in Appendix B) A comparison between the 7B/13B LLaMA trained on our filtered 3k dataset and the one trained on the entire Dolly 15k dataset is illustrated in Fig. 10 and Fig. 21. Our evaluation suggests that the model trained on our filtered data exhibits superior performance, thus underscoring the efficacy of our filtering method on human-composed datasets. Comprehensive details regarding training hyperparameters are provided in the Appendix D. The result in Section A.4 (GPT4LLM dataset) shows the potential of applying our ChatGPT-based response quality evaluator to filter GPT-4’s responses, which is considered as the most powerful model.
Case Study & Analysis
Fig. 11 shows two case studies of 13B models trained on 52k data (Alpaca), 9k selected data (AlpaGasus), and 9k randomly selected data (Alpaca-9k-random). The left case study focuses on the math capability, where AlpaGasus can produce a correct answer while Alpaca-9k-random cannot. As the judge, GPT-4 rates the answer of AlpaGasus by a score of 10.0 while Alpaca-9k-random receives a score of 2.0. The right case study focuses on coding skills, Alpaca-52k cannot follow the instructions but produces a regular expression to validate the website address while AlpaGasus directly generates the correct code.
We also conduct a fine-grained evaluation of AlpaGasus on each skill/category in the WizardLM and Vicuna test sets, whose samples are split into a list of skill sets/categories and thus facilitate detailed analyses of the capabilities achieved by IFT (Appendix H). We compare two 7B models on the WizardLM test set and report the results in Fig. 25. Our AlpaGasus achieves better or equally good performance than Alpaca on 22/29 skills but does not show advantages on the remaining 7 skills such as coding (e.g., code generation). To investigate the reasons, we notice that the coding categories include “python”, “Java”, “C++”, and “C#”, which indicate that we can allocate training samples regarding coding skills based on these related keywords (Appendix E). We find that our data selection/filtering, without specifying the proportions of skill categories, leads to a much higher filtering ratio of coding-related data than the average filtering ratio . Hence, the resulting coding skill is weaker than other skills. This indicates the importance of keeping the training data diverse and balanced across different categories in IFT.
Cost Saving
We compare the training cost of AlpaGasus and Alpaca in terms of the estimated expenses for the required computation on AWS. Notably, the training time is reduced from 80m to 14m for the 7B model and 5.5h to 1h for the 13B model. Such training time reduction not only substantially enhances model iteration speed, but also reduces the cost from 4.78 for the 7B model and 40.96The hyperparameters for IFT and the projected costs calculation method are deferred in Table 5. for the 13B model. It’s noteworthy that instruction-tuning 65B LLaMA models require a greater number of GPUs and an extended training duration. Consequently, as the model size scales up, our data selection method yields progressively pronounced cost savings.
Related Work
Instruction-tuning datasets can be gathered in two ways. A number of studies (Köpf et al., 2023; Dolly, 2023; Zhou et al., 2023) utilize crowdsourcing to produce human-generated pairs of instructions and responses. This approach, while effective, can be laborious and costly. Alternatively, Alpaca (Taori et al., 2023) opens the door to create machine-generated IFT sets from the distillation of the “teacher” LLM, i.e., Text-Davinci-003. Peng et al. (2023) keep the instructions from Alpaca intact but using GPT-4 as the “teacher” LLM, which enhances model on 3H (Helpfulness, Honesty and Harmlessness) (Askell et al., 2021) alignment criteria. Vicuna (Chiang et al., 2023) is the first to adopt ShareGPT (ShareGPT, 2023) data, which is the realistic dialogue data chatting with ChatGPT shared by users. Xu et al. (2023) and Luo et al. (2023) evolve the original Alpaca instruction set and obtain more complex instructions which help better elicit the instruction-following ability of LLMs. There also exists concurrent work like Koala (Geng et al., 2023) and UltraChat (Ding et al., 2023), using dialogue & preference data as well as the adversarial prompts to conduct safe alignment.
Over the last decade, the realm of data-centric AI (Chu et al., 2016; Motamedi et al., 2021) has witnessed substantial progress. Central to this concept is the belief that the quality of data (Hajij et al., 2021; Zha et al., 2023; Chen et al., 2023a; c; d) warrants the same level of importance as algorithms within the AI/ML lifecycle. As noted by Chu et al. (2016), for an effective engagement with diverse types of data across various domains, data cleaning processes should exhibit a higher degree of automation and adaptability. With the advent of the Transformer architecture (Vaswani et al., 2017b), a shift in the paradigm of language models has occurred. Models such as RoBERTa (Liu et al., 2019), BERT (Vaswani et al., 2017a), and Bard https://bard.google.com/ all have incorporated this effective structure, stacking varying quantities of transformer blocks to create more potent models. This marked a turning point in NLP research, signifying a heightened emphasis on data as opposed to model structure. Presently, SOTA LLMs like ChatGPT also underscore this shift toward data. They employ user data to conduct Reinforcement Learning from Human Feedback (RLHF) (Ouyang et al., 2022a; Gao et al., 2022), which further aligns with the Data-centric AI philosophy.
Evaluating the open-ended instruction-following ability of LLMs is often neglected by previous works (Chung et al., 2022; Anil et al., 2023), though they conduct a series of benchmark evaluations centered around factuality (Hendrycks et al., 2020) and reasoning (Bisk et al., 2020) for their pre-training models. Similarly, the frameworks proposed by Liang et al. (2022) and Gao et al. (2021) focus more on the evaluation of the base models but not on the evaluation of the IFT models, where open-ended instruction-following capability are supposed to be prioritized. Since instruction-following is a general ability but the scope of benchmarks is limited, the recent works such as Koala (Geng et al., 2023), Vicuna (Chiang et al., 2023), Self-Instruct (Wang et al., 2022), and WizardLM (Xu et al., 2023) all provide the instruction sets they collected and some of them also include the categories of the instructions for the evaluation of instruction-tuned LLMs. There are also some leaderboards like Alpaca-Eval (Li et al., 2023) measuring the model’s instruction-following ability. Leveraging these recent advancements, we evaluate our models on human instruction sets.
Conclusion
In conclusion, our study reveals significant insights about the influence of data quality over quantity in IFT. Through our proposed data-filtering method, we have demonstrated that relying on a small subset of high-quality IFT data can lead to LLMs that exhibit enhanced instruction-following capabilities, while also offering substantial computational advantages. Notably, our method proves versatile across different rating dimensions (e.g., Accuracy and helpfulness), LLM filters (e.g., ChatGPT and Claude-2), base model families (e.g., LLaMA-1 and LLaMA-2), model sizes (e.g., 7B and 13B), dataset types(e.g., machine-generated and human-written). By emphasizing the importance of data quality, we advocate for a transition in the existing paradigm where data accumulation has been a primary focus. This perspective transition can lead to more meaningful advancements in the field of LLMs, making models more aligned with human intentions and less prone to errors induced by poor-quality data.
acknowledge
Lichang Chen and Heng Huang were partially supported by U.S. NSF IIS 2347592, 2347604, 2348159, 2348169, DBI 2405416, CCF 2348306, CNS 2347617.
References
Appendix
We also explore alternate evaluation prompts such as the prompts provided by Zheng et al. (2023), which are shown in Table 3. We apply the same rules to calculate the “Win-Tie-Lose” and show the results in Fig. 12. Notably, AlpaGasus consistently outperforms across all test sets.
A.2 Have you tried other LLM filter?
Yes, we also try to use Claude-2https://www.anthropic.com/index/claude-2 as our response quality evaluator (LLM filter). Fig. 13 and Fig. 14 demonstrate the score distribution and evaluation results on the four testsets, respectively. Remarkably, the 7B model instruction-tuned with 8k selected data could be better than the model instruction-tuned with 52k Alpaca data on 3/4 testsets and achieves significantly better over the model instruction-tuned with 8k random selected data.
As Fig. 13 shows, the interval between two scores is 1, which is different from the ChatGPT-based filter, where the interval is . Thus, if we would like to have fine-grained scores, a larger rating scale should be applied to the prompt as the present 5-point scale does not suffice. We leave the exploration of the rating scales to future work.
A.3 What about the results on other base models, e.g., LLaMA-2?
We also have the results of LLaMA2 in Fig. 15, which shows the superiority of our method.
A.4 Can your LLM filter evaluate the stronger model’s responses, e.g., filtering the responses given by GPT-4?
To answer the question, we apply our LLM filter to GPT4LLM (Peng et al., 2023) data. According to the score distribution, we use 4.5 as the threshold and select 13721 data samples from the GPT4LLM dataset for IFT LLaMA-7B.
The results presented in Fig. 17 demonstrate the superiority of our method on the Vicuna and WizardLM test sets. Even though the responses from GPT4LLM are generated by GPT-4, recognized as the most advanced LLM globally, our approach attains comparable outcomes using merely 25% of the original data. Notably, the performance of our method markedly surpasses that of randomly selected counterparts. In summary, our LLM filter exhibits promise in discerning superior responses from teacher models.
A.5 Results on other rating dimensions, e.g., helpfulness?
We also use “helpfulness” as our rating dimension and find that we only need 2k data to train the base model that can surpass the base model trained with 52k Alpaca data. The score distributions are shown in Fig. 18.
From Figure 19, it is evident that the models trained using our filtered Alpaca dataset outperform those trained on randomly selected datasets across all instruction test sets. Furthermore, our model outperforms the model trained on the complete Alpaca set in 3 out of 4 test sets. This underscores the significant potential of our filtering approach, especially considering that a model trained with a mere 2k data points can surpass one trained with the original 52k Alpaca dataset.
Appendix B Additional Results on Dolly Dataset
We show the score distribution of Dolly dataset(rated by ChatGPT) in Fig. 20.
B.2 Benchmark results
We use the code provided by Chia et al. (2023) to conduct benchmark evaluation. For MMLU, BBH, Drop, and humaneval, we also use 5-shot, 3-shot, 3-shot, and 0-shot settings, respectively. We show the benchmark results in Table 4 of Dolly and the filtered set.
Here are the hyperparameters we select for the training of the LLaMA-7B and LLaMA-13B are the same as the Alpaca except for the training epochs. To avoid the under-train issue, we train 10 epochs, instead of 3 in Alpaca, for all the 7B models and 15 epochs, instead of 5 in Alpaca, for all the 13B models.
B.3 Dolly-13B Results
We show the dolly-13B results. As Fig. 21 shows, our filtered Dolly dataset is better than the original Dolly dataset since it can achieve stronger instruction-following capacity of the instruction-tuned LLaMA-7B models via ours. (See the results on the four tests)
Appendix C Details of GPT-4 Evaluation Prompt
We provide the detailed form of the prompt to GPT-4 used for evaluation in Fig. 22. It is the prompt for evaluation used in the original Vicuna blog https://lmsys.org/blog/2023-03-30-vicuna/
Appendix D Training Hyperparameter Details
We show the training hyperparameters and costs in Table 5. https://aws.amazon.com/ec2/instance-types/p4/ a p4de.24xlarge(preview) node has 8 80GB A100 and it costs $40.96/h.*we assume training time of using 8 GPUs is half of using 4 GPUs
D.2 Dolly Dataset
We show the training hyperparameters in Table 6.
Appendix E Keywords set for detailed analysis
We use the keyword set of and count the number of (instruction, input, output) tuples which contain the keyword in this set.
Appendix F Rated examples in Alpaca Dataset
We include more examples rated by the response quality evaluator, i.e., ChatGPT, in this section. The examples of Score 5.0, Score 4.5, Score 4.0, Score 3.5, Score 3.0, Score 2.5, Score 2.0 are shown in Table 7, Table 8, Table 9, and Table 10, respectively.
Appendix G Rated examples in Dolly Dataset
Appendix H Analysis
We conduct a fine-grained evaluation of AlpaGasus on each skill/category in the WizardLM and Vicuna test sets, whose samples are split into a list of skill sets/categories and thus facilitate detailed analyses of the capabilities achieved by IFT.
We compare these two 7B models on the WizardLM test set and report the results in Fig. 25. Our AlpaGasus achieves better or equally good performance than Alpaca on 22/29 skills but does not show advantages on the remaining 7 skills such as coding (e.g., code generation). To investigate the reasons, we notice that the coding categories include “python”, “Java”, “C++”, and “C#”, which indicate that we can allocate training samples regarding coding skills based on these related keywords (Appendix E). We find that our data selection/filtering, without specifying the proportions of skill categories, leads to a much higher filtering ratio of coding-related data than the average filtering ratio . Hence, the resulting coding skill is weaker than other skills. This indicates the importance of keeping the training data diverse and balanced across different categories in IFT.
H.2 Analysis on Vicuna Test Set
Fig. 23 demonstrates the detailed analysis on Vicuna testset. AlpaGasus-7B is better than the Alpaca-7B in the majority of the categories, including Counterfactual, Roleplay, Knowledge, and Generic, etc. Another strong point is that when the base model scales up, the conclusion still holds. (See right part of the Fig. 23)
Appendix I Detailed Analysis on the WizardLM testset
In Fig. 26, Fig. 27, and Fig. 28, we compare AlpaGasus with text-Davinci-003, ChatGPT, and Claude, respectively. The results show that AlpaGasus-13B can achieve capacity of its “teacher” model, text-Davinci-003 (all the responses in the Alpaca-52k dataset are generated by text-Davinci-003 so we call it “teacher” LLM). The results also show that our model could achieve pretty good performance on tasks like Writing, RolePlay, Toxicity, Art, etc., while it still needs improvement on coding and math capacity when compared with stronger LLMs.
Appendix J Human Study
We conduct the human study among three different users. The evaluation interface is shown as Table 15:
We show more detailed results of human evaluations in Fig. 29.
Appendix K Limitations
In our experiments, we evaluated our IFT strategy by training models of two different sizes, 7B and 13B, since they are the most common sizes for recent open-source LLMs. We plan to extend this study to larger model sizes such as 33B, 65B, or even 175B, and verify whether the same conclusion still holds, i.e., a small subset of high-quality data selected by our method can improve the instruction-finetuned model. We leave analysis on the IFT of larger models as future work.