From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao
Introduction
Large Language Models (LLMs) have revolutionized the landscape of artificial intelligence Touvron et al. (2023a, b); Penedo et al. (2023); Scao et al. (2022). Notable models such as GPT-3 Brown et al. (2020) and GPT-4 OpenAI (2023) leverage extensive datasets and advanced training methodologies to exhibit high-level text understanding and generation capabilities. The applications of these models extend across diverse domains, including interactive systems, automated content generation, and support for scientific inquiries.
Instruction fine-tuning Wei et al. (2022); Longpre et al. (2023) is a method employed to refine the performance of LLMs by providing specific guidelines or instructions during the model’s training phase. It operates by supplying the LLM with explicit training instructions to produce the corresponding outputs that are more congruent with the desired responses. A well-formulated instruction or prompt provides essential contextual information, refining the model’s capability to generate relevant and task-specific outputs Taori et al. (2023); Ouyang et al. (2022).
Based on the findings from Wang et al. (2022b) and Self-Instruct Wang et al. (2023b) early experiments, reducing the number of instances per task does not degrade the model’s generalization performance to unseen tasks. While conventionally, instruction tuning is predominantly relied on amassing vast datasets. A seminal revelation from the LIMA Zhou et al. (2023) highlights the art of instruction tuning: rather than a sheer volume of data, it’s the quality of the data that dictates the model’s performance. LIMA’s findings emphasize that even a limited amount of manually curated, high-quality data can elevate the model’s instruction-following prowess. While it underscores the efficacy of data overabundance, the question of how to automatically identify high-quality data from a vast ocean of available datasets remains under investigation.
In our study, we introduce a novel approach for autonomously identifying the most impactful training samples from extensive open-source datasets, which we refer to as "cherry data." These are the data fragments that are particularly effective in enhancing Large Language Model (LLM) instruction tuning. Central to our hypothesis is the idea that LLMs, through initial training with carefully selected instruction data, can inherently learn to discern and follow instructions. This ability allows them to evaluate the quality of broader datasets and estimate the difficulty of following instructions in a self-guided manner.
Our method involves a self-guided process that begins with familiarizing the model with a subset of the target dataset during the "Learning from Brief Experience" phase. This phase lays the groundwork for the subsequent "Evaluating Based on Experience" phase, where we introduce the Instruction-Following Difficulty (IFD) score. This metric, focused on minimizing cross-entropy loss, helps in isolating the instructional component from the answer’s influence by comparing the loss in model responses with and without instructional context.
We propose selecting samples with moderate IFD scores for instruction tuning, as this strikes a balance between addressing challenging points and avoiding redundancy. This approach is vital for enhancing the model’s ability to process and follow complex instructions. In the final "Retraining from Self-Guided Experience" phase, we use cherry data with notable IFD scores to refine our model, resulting in what we term "cherry models." This methodology, which emphasizes data quality over quantity, differs markedly from existing techniques that rely on external models for data curation.
Extensive experimental results validate the efficacy of our method. By applying our methodology to the Alpaca and WizardLM instruction tuning datasets, our model outperforms the official Alpaca model with only approximately data selected and outperforms the reimplemented WizardLM model with approximately data selected. The key contributions of this paper:
We propose a self-guided approach enabling models to autonomously “select cherry data” from vast open-source datasets. This innovation minimizes manual curation and optimizes the use of existing data resources, reducing costs and streamlining training.
We introduce the Instruction-Following Difficulty (IFD) metric as a tool to identify gaps in a model’s responses versus its autonomous generation capability. Using the IFD metric, we can pinpoint these cherry samples, optimizing model training efficiency.
Backed by validation on datasets like Alpaca and WizardLM, our strategy demonstrates enhanced outcomes with only 10% of the typical data input, emphasizing our approach’s efficiency and transformative impact.
We provide a different model-specific view in measuring the difficulty of new instructions, which may benefit future instruction data generation work.
Methodology
As illustrated in Figure 1, our methodology is divided into three core phases: Learning from Brief Experience, Evaluating Based on Experience, and Retraining from Self-Guided Experience. The initial phase emphasizes equipping the model with a basic instruction-following capability using select portions of the dataset. The subsequent phase introduces a novel metric to evaluate the instruction-following difficulty score of each sample based on the previously trained pre-experienced model. Finally, after obtaining difficulty scores in the target dataset, the cherry samples are defined and sampled to train our final model, which we call the cherry models. In our experiments, the underlying model used is the Meta LLaMA Touvron et al. (2023a), complemented by the target dataset.
This phase aims to equip the initial model with a basic instruction-following capability by forcing the model to first experience a subset of the target dataset. Specifically, for the initial full target dataset, contains triplets , we define the string as the complete instruction. The function is aligned with the original target dataset. Each word in and is denoted as and respectively. Let denote the LLM we use and represent the weight of LLMs, specifically, represents the pre-trained base LLM model. Then the instruction embeddings for each sample are obtained by:
where represents the word of strings of sample and represents its corresponding last hidden states.
To ensure the diversity of instructions exposed to the initial model, the basic clustering technique K-Means on these instruction embeddings is utilized. Motivated by LIMA’s finding, we are motivated to make this experience process as brief as possible by sampling only a few instances in each cluster which we call pre-experienced samples. Specifically, we generate clusters on instruction embeddings and sample instances in each cluster. Then the initial model is trained for only epoch with these samples to obtain our brief pre-experienced model.
2 Evaluating Based on Experience
In this stage, advancing from the previous phase, we introduce the Instruction-Following Difficulty (IFD) score, a metric devised to evaluate the challenge each instructional sample presents. Our primary motivation, adhering to the goal of minimizing cross-entropy loss in model training, guides the use of this metric. It specifically targets gauging the impact of training data by isolating the instructional component’s influence from that of the answer. To achieve this, we employ a method that compares the loss when the model generates responses both with and without the context provided by instruction. This comparison is crucial as it forms the basis of the IFD score, effectively quantifying the extent to which instruction aids in response generation.
In the instruction-tuning process, the loss of a sample pair is calculated by continuously predicting the next tokens given the instruction and their proceeding words:
where is the number of words of the ground-truth answer . We denote this averaged cross-entropy loss as the Conditioned Answer Score .We use different symbols to differentiate the loss when used as an objective function and the loss when used as a score. This metric evaluates the model’s capability to generate appropriate responses based on provided instructions. It measures the extent to which the model’s output aligns with both the instruction and the corresponding correct answer.
Under this circumstance, a higher does not mean a harder instruction to follow, it may simply be caused by the inherent factor of string itself. In the pre-LLM era, when models are required to learn both the knowledge and instruction-following ability during finetuning, it is reasonable to simply use as an indicator for the difficulty of a sample. However, things change a little for current LLMs, which have learned most of the knowledge in the pre-training phase and only need to learn to align and follow the instructions. To estimate the difficulty of following instructions of a given sample, we introduce the Direct Answer Score :
which measures LLM’s ability to generate this answer alone. This metric gauges the inherent difficulty or challenge posed by the answer in isolation, without the contextual guidance from its corresponding instruction. A higher direct answer score may suggest that the answer is inherently more challenging or intricate for the model to generate.
Further, analyzing the balance between a sample’s inherent challenge and the model’s capabilities in following it sheds light on the intricacies of estimating the difficulty of the instruction of a given sample. Specifically, we try to estimate the Instruction-Following Difficulty (IFD) scores on following instruction of a given pairs by calculating the ratio between and :
Under this circumstance, the influence of LLM’s intrinsic ability to fit the answer string is partially alleviated. The score measures the degree how a given instruction benefits the alignment of the corresponding response. High IFD scores infer the inability of the model to align responses to the given corresponding instructions, which in turn indicates the difficulty of an instruction. It is worth noting that this is a model-specific value, and we use our pre-experienced model to obtain all these values in the target dataset.
Experimental Setup
Training Datasets The Alpaca dataset Taori et al. (2023) from Stanford University, encompasses instruction-following samples. Developed using the self-instruct Wang et al. (2023b) approach with text-davinci-003. Though initially competitive, its dependence on text-davinci-003 posed data quality concerns. WizardLM dataset Xu et al. (2023) leverages the Evol-Instruct algorithm to improve the quality of instruction data. Furthermore, the incorporation of ChatGPT during the reformulation guarantees high fidelity of data. Of its instructions, we primarily utilized the WizardLM-7b subset, consisting of samples.
Test Datasets To ensure comprehensive and unbiased assessment, we employed diverse test sets: Vicuna Chiang et al. (2023), Koala Vu et al. (2023), WizardLM Xu et al. (2023), Self-instruct Wang et al. (2023b), and LIMA Zhou et al. (2023). These test sets contain approximately human curated instructions, closed-domain or closed-domain for different tasks from different sources. Among them, Vicuna and WizardLM further provide the specific sub-category for each instruction, making it possible for in-depth analysis. SAlthough all of these test sets are introduced to guarantee the variety of testing, we select all of these sets to offer a broader palette of instruction types than the typical one or two test sets.
2 Implementation Details
Rooted in the Llama-7b pre-trained model, our training framework aligns with protocols from Alpaca and WizardLM datasets. The Adam optimizer Kingma and Ba (2017), with a learning rate and a batch size of , steers the training across three epochs. Our pre-experienced models, however, undergo just a single epoch of training. Training on the Alpaca dataset necessitated a max input length of . For WizardLM, we opted for a input length due to hardware constraints while its original model used , which offers an inherent edge to the original model. Another challenge with WizardLM was “AI censure” instances. Taking a leaf from the Vicuna strategy, we filtered these samples, resulting in a streamlined WizardLM subset with entries. Our data selection methodology was then applied to this subset. Samples with IFD scores higher than will be filtered out before selection. For experiments on llama2-7b and llama2-13b models, we utilize the instruction prompt from Vicuna Chiang et al. (2023). Thanks to the Flash Attention mechanism Dao et al. (2022), all models on llama2 use the max length of .
3 Evaluation Metric
Assessing the instruction-following capabilities of LLMs is challenging. While extensive research is dedicated to creating automated evaluation metrics for LLMs Chang et al. (2023), human judgment remains unmatched. However, it’s both labor-intensive and potentially influenced by biases. Leveraging the recent advancements in independent LLM evaluations Zheng et al. (2023); Chiang et al. (2023); Li et al. (2023), we utilize GPT4 and ChatGPT for comparative evaluations. Following Chen et al. (2023), for each instruction in the test dataset, models that need to be compared are prompted to generate responses respectively. Then an API model, either GPT4 or ChatGPT, assigns scores for their responses. The model is regarded to be better in this dataset only if its answer is preferred by the judging model.
In the evaluation, each model’s response is rated by the judge on a scale from to , reflecting attributes like relevance and accuracy. To further address the positional bias Ko et al. (2020); Wang et al. (2023a), we send the responses of two models to the judge twice with different orders and compare their scores. Thus we define one model to be seen as winning only if it does not lose in both the ordering, specifically:
Wins: outperforms in both or wins in one and ties in the other.
Tie: ties in both or wins in one and loses in the other.
Loses: lags in both or ties in one and loses in the other.
This evaluation serves as the foundation of our experimental outcomes.
3.2 Benchmarks
The performances on two recently popular benchmarks for LLMs are also provided: Huggingface Open LLM Leaderboardhttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard and AlpacaEval Leaderboardhttps://tatsu-lab.github.io/alpaca_eval. Huggingface Open LLM Leaderboard evaluates LLMs using Gao et al. (2021), a unified framework to test generative language models on a large number of different evaluation tasks, on key benchmarks including ARC Clark et al. (2018), HellaSwag Zellers et al. (2019), MMLU Hendrycks et al. (2021) and TruthfulQA Lin et al. (2022). AlpacaEval Leaderboard provides an LLM-based automatic evaluation based on AlpacaFarm Dubois et al. (2023) evaluation set, in which the model responses are compared with responses of Davinci003 by GPT4.
3.3 Human Evaluation
To better illustrate the efficacy of our method, further human evaluation is conducted. Specifically, we randomly sampled instructions from each test set to generate a new random set containing instructions in total. Then human participants are asked to compare the responses generated by the models to be compared. For each comparison, 3 options are given (Win, Tie, and Loss) and the final results are determined by the majority voting of the participants.
Experimental Results
In this section, we first present our primary pair-wise evaluation results in Figure 2. (a) our model trained with only approximately of the original Alpaca data beats the Alpaca model trained with full data. (b) our model trained with only approximately of the original WizardLM data beats the reimplemented WizardLM model under the same training configuration which is described in the Implementation Details.
Moreover, we craft subsets containing the top , , , and of the target datasets and train models on these distinct subsets, enabling us to investigate the performance changes. As shown in Figure 3, we draw the overall winning rate changes across the data growth, which is calculated as (Num(Win)Num(Lose))Num(All) , providing a direct indicator on the comparison with the full-data trained models. A consistent observation across both datasets is that with merely of selectively chosen data, our models manage to exceed the results of models trained on the full dataset. These findings not only highlight the efficiency of our data selection strategy but also underscore the potential of training powerful models with significantly reduced data requirements. By validating our approach on the renowned Alpaca dataset and the more intricate WizardLM dataset, we emphasize the wide applicability and robustness of our proposed method.
The comparison between our cherry models with baseline models on Huggingface Open LLM Leaderboard and AlpacaEval Leaderboard are presented in Table 6 where we can see our cherry model using Alpaca data outperforms the official Alpaca on both benchmarks, our cherry model using WizardLM data has a closed performance compared with our re-implemented WizardLM. These results further showcase the effectiveness of our automatically selected data.
Moreover, the human evaluation results also showcase the usefulness of our method. When comparing the Cherry Alpaca (5%) and the Alpaca (100%), there are wins for our cherry alpaca, ties, and losses. When comparing the Cherry WizardLM (10%) and the reimplemented WizardLM (100%), there are wins for our Cherry WizardLM, ties, and losses.
2 Ablation on Cherry Data Selection
We train various LLaMA-7B models using randomly chosen data and juxtaposed their performance with that of our models, which employed a difficulty ratio. As shown in Figure 4 (labeled as Random), models trained on , , or random data consistently underperformed against the official Alpaca model. Notably, with an equivalent amount of data, our model surpasses the performance of models using randomly selected data, underlining our method’s superiority.
2.2 Data with Diversity
In this experiment, we train a series of models only considering the diversity of the data samples. Specifically, we utilize the same method for obtaining diverse samples from the full dataset, in which a k-means algorithm is first implemented, and then data is sampled from each cluster. It is a direct baseline for the situation where only the diversity of data is considered. As illustrated in Figure 4 (labeled as Diversity), these models render subpar performance and are similar to the random trained models. This result highlights the necessity of using difficult samples over pure diverse samples.
2.3 Data with Low IFD Score
In this experiment, we aim to underscore the efficacy of our proposed IFD score. We train a model using data chosen based on low IFD scores on the pre-experienced model, a direct antithesis to our primary experimental setting. As illustrated in Figure 4 (labeled as Low IFD score), models trained using low IFD scores render subpar performance. This observation highlights the prowess of our metric in sifting through high-quality data: a higher score consistently yields superior results compared to the baseline, while a lower score deteriorates the model’s intrinsic performance.
2.4 Data with High CA Scores
For this comparison, we juxtapose our model against one trained on data selected by higher Conditioned Anser scores which is equivalent to the loss or perplexity, and is a commonly accepted baseline. As Figure 4 (labeled as High CA score) elucidates, models in this group trail the official Alpaca model significantly. The salient difference between these models and ours rests on the elimination of Direct Answer scores. In models relying solely on CA scores, the underlying comprehension of the pre-trained LLM towards original answer texts isn’t factored in, rendering high CA scores ineffective in gauging the intricate nuances of the instruction following.
3 Results on Other Models
In this section, experiments on newer LLaMA2-7B and LLaMA2-13B models are conducted as shown in Table 2. In these experiments, the IFD score of each sample is calculated directly based on the corresponding LLaMA2 pre-trained models by using prompts from Vicuna Chiang et al. (2023). On both LLaMA2-7B and LLaMA2-13B models, our cherry models trained with much less data outperform the models trained with original full data. These experimental results illustrate the consistent advantages of our method and further verify the generalizability of our method.
Cherry Data Characteristics
Our goal in this section is to determine if the data selected based on the IFD scores aligns with known characteristics of high-quality training data. To this end, we randomly sample instances from data with the top scores and the least scores. Utilizing ChatGPT, we evaluate each instruction on six aspects: Scope, Complexity, Clarity, Depth, Simplicity, and Knowledge Required. The results are depicted in Figure 5. Data with a higher IFD score generally scored higher in Scope, Complexity, Depth, and Knowledge Required, but lower in Clarity and Simplicity. Simplicity, in particular, have the most pronounced discrepancy. This lends credence to our assertion that our IFD scores aptly gauge instruction complexity. Consequently, our method gravitates towards selecting more intricate samples.
As mentioned in the previous section, we try to evaluate each instruction into six aspects, Scope, Complexity, Clarity, Depth, Simplicity, and Knowledge Required. We define these aspects as follows:
Scope: The instruction encompasses the breadth and range of actions or information necessary for successful completion.
Complexity: The instruction integrates multiple steps or concepts that require careful attention and understanding.
Clarity: The instruction is articulated straightforwardly, ensuring it’s easily understood without ambiguity.
Depth: The instruction provides thorough details and nuances, ensuring a comprehensive understanding of the task at hand.
Simplicity: While thorough, the instruction avoids unnecessary jargon or convolutions, making it accessible and easy to follow.
Knowledge Required: The instruction acknowledges and, if necessary, provides the foundational knowledge or context the user needs for successful execution.
From the previous Figure 5, we can see samples selected with top IFD scores have larger scores in the aspects that reflect the difficulty of instruction, including Scope, Complexity, Depth, and Knowledge Required. These samples only underscore samples with the lowest IFD scores on the aspect of Clarity and Simplicity. This experiment detailedly illustrates the difference between samples with high or low IFD scores and verifies the effectiveness of our method in measuring the difficulty of an instruction.
2 Distributional Characteristics
In this segment, our focus is on understanding the distributional properties of the cherry data within the original dataset. Specifically, we first compute the embedding of each instruction in the Alpaca dataset and employ t-SNE for dimensionality reduction, mapping high-dimensional embeddings to 2D space. The visualized vectors, color-coded based on the top or least difficulty ratios, are showcased in Figure 6. Contrary to conventional beliefs, our cherry data isn’t uniformly scattered. Instead, a palpable demarcation exists between samples of high and low difficulty, challenging prior assumptions that selected data should span the entire instruction spectrum and maximize diversity.
To delve deeper into the distributional intricacies of instruction embeddings, we utilize naive K-means (K=) for clustering. We home in on representative clusters, half of which displayed a significant overlap with the top samples and the other half with the least samples. Clusters dominated by low IFD score samples are replete with rudimentary tasks like editing punctuation, words, or sentences. In contrast, high IFD score clusters are typified by deeper, more intricate tasks such as storytelling or elucidation of phenomena. We posit that these in-depth tasks are paramount for aligning large language models, compelling them to rearrange and access their intrinsic knowledge repositories. Our methodology lends partial credence to this hypothesis, leaving room for further exploration.
Related Work
Previous instruction tuning collections are typically handcrafted or task-related Khashabi et al. (2020); Ye et al. (2021); Wei et al. (2022); Wang et al. (2022a); Du et al. (2022); Honovich et al. (2023), Wang et al. (2023b) utilized GPT3 Brown et al. (2020) to generate k distinct instructions which do not directly relate to each task, which paves the way to generating instruction data set by distilling from teacher models. After the release of Meta LLaMATouvron et al. (2023a), the world witnessed a surge of open-sourced instruction tuning datasets and LLMs.
2 Coreset Selection
Coreset selection is pivotal in machine learning, aimed at identifying a representative subset of data points to expedite learning in various models. This approach finds its effectiveness in SVM learning Tsang et al. (2005), k-means Har-Peled and Kushal (2005), logistic regression Munteanu et al. (2018). However, these methods primarily focus on traditional machine models.
In neural network training, recent advancements, such as those by Toneva et al. Toneva et al. (2018), explore the dynamics of data point utility during training. They find that points infrequently forgotten have minimal impact on final model accuracy. Mansheej Paul et al. Paul et al. (2021) demonstrate that expected loss gradient norm scores, averaged over various weight initializations, effectively prune training data without significantly compromising accuracy. Sören Mindermann et al. Mindermann et al. (2022) use Bayesian probability theory to estimate the individual impact of training points on holdout loss, refining training efficiency.
The common thread in these works is the focus on traditional training mechanisms. Our research pivots towards optimizing core-set selection in instruction tuning. Similar to aboved methods, our score estimation is relied on the feature representation of the target model.
3 Instruction Data Selection
Though consensus has been made that "quality is all you need" Touvron et al. (2023b) for instruction tuning, finding high-quality data other than through human curation is still an under-explored topic. Two recent papers, Instruction Mining Cao et al. (2023) and ALPAGASUS Chen et al. (2023), proposed methods to bridge this gap, sharing similar motivations to our work. Instruction Mining evaluates various indicators and applies a statistical regression model for data selection. However, it lacks comparative performance analysis with models trained on full data, and its method is seen as overly complicated, involving splitting data into several bins and training models fully. In contrast, ALPAGASUS utilizes an external, fully-trained LLM (ChatGPT) to score each sample. It selected 9k Alpaca data, surpassing the official Alpaca trained on full data. While effective, this approach may neglect the intrinsic abilities of the base model, relying excessively on external models.
In contrast, our work aims to develop a methodology utilizing the representation feature of the target model to identify high-quality data for instruction tuning, advancing the field with a more simple and efficient approach.
Conclusion
This study has illuminated the potential of harnessing the innate capabilities of LLMs for selecting high-quality instruction tuning data that fit the model. Through our innovative self-guided approach, LLMs demonstrate the ability to discern and cherry-pick the most pertinent data samples, a concept we’ve aptly termed cherry data. Central to our methodology is the Instruction-Following Difficulty metric, a novel tool adept at gauging the nuanced differences between a model’s autonomous outputs and expected responses. Our findings not only emphasize the importance of data quality over quantity but also underscore the potential for cost-effective and streamlined LLM training.
Limitation
The main limitation of this method is the inconvenience of training the pre-experienced model. The concept of the Instruction-Following Difficulty score proposed by us is simple and effective, while the inconvenient pre-experienced phase makes it hard to directly put our method into usage in real-world scenarios. Though experiments on LLaMA2 models show that calculating IFD scores directly on the base LLaMA2 models also promises a good selection, we believe using the pre-experienced phase is valuable since it equips base models with the basic instruction-following ability, making the calculation of Conditioned Answer Score more reasonable. As a result, we believe the use of the pre-experienced phase could be a tradeoff: From the Research Viewpoint, using pre-experienced models is more reasonable and performs better. From the Real-world Implementation Viewpoint, directly using the base model is more efficient and at the same time effective as well.
References
Appendix A Ablation on Pre-Experienced Data Selection
Following the findings from LIMA that high-quality samples are enough to train a reasonably good model, we set the amount of data used for our pre-experienced model as . However, it is still under-investigated how many data samples are required to equip the model with basic instruction-following ability. Thus this section analyzes the significance of employing experience-augmented models and how the number of pre-experienced data affects the final performance of our cherry models. For these comparisons, we conduct the experiments where , , , and pre-experienced samples are utilized to train the pre-experienced models for epoch. Using pre-experienced samples represents direct use of the initial raw model as the pre-experienced model. We calculate the IFD scores from these different pre-experienced models and select the top , , and samples for training while keeping other experimental conditions constant.
As depicted in Figure 7 0 Pre-Experienced Samples, when no pre-experienced samples are utilized, the corresponding cherry models have the least performance. These results underline the indispensability of an experience-augmented model equipped with foundational instruction-following capabilities. Moreover, even in the absence of a pre-experienced model, our IFD score remains effective in identifying optimal training data as it outperforms the Alpaca model when using of the data. When samples are utilized as shown in 100 Pre-Experienced Samples, the corresponding cherry models are slightly better than no samples used but with a similar trend, which indicates that samples are not enough for the model to acquire the basic instruction-following ability. When adding the number of pre-experienced samples to , a distinct performance gain is discovered, and further addition of samples does not make the performance of corresponding cherry models better. We hypothesize this is when the model is equipped with the basic instruction-following capability and thus can better illustrate the instruction-following score of each instance.
A.2 Distribution of Pre-Experience Data
To better illustrate what kinds of data are required in the pre-experience process, extensive experiments are conducted where we selected pre-experienced samples by calculating the IFD scores based on the initial raw model and utilized these samples to train the pre-experienced model and further get the cherry samples and the cherry model. Different from our main method where the pre-experienced samples are selected based on the diversity of instruction distribution, this experiment is used to figure out what is the better strategy for the pre-experienced model, considering the diversity of instructions or difficulty of instructions. Another baseline method is using randomly selected data for the training of pre-experienced models. The performance of using , , and cherry data is shown in Table 3 compared with the Alpaca model. Comparing random selection or considering embedding distributions or instruction difficulties, they all surpass the Alpaca model and are comparable to each other, indicating the effectiveness of both strategies and further proving that our IFD metric is robust across different pre-experienced models. This experiment further illustrates that what matters is this pre-experience process, rather than the sampling strategies for this process.
Appendix B Performance across Sub-Categories
To evaluate the performance variations of our model, we scrutinize the capabilities across diverse instruction tasks. To accomplish this, we compare the response of our cherry models, trained with Alpaca data and WizardLM data, to their corresponding comparing models, the official Alpaca and the reimplemented WizardLM across sub-categories in the WizardLM and Vicuna test sets, as displayed in Table 4 and Table 5.
Our cherry model trained on Alpaca data exhibits superior or at least comparable performance to the official Alpaca model on most of the subcategories in the Vicuna and WizardLM test sets. Notably, exceptions are observed in the Math and Coding categories, corroborating the observations made by Chen et al. (2023). We surmise that the base 7B models inherently perform sub-optimally on these two tasks, necessitating a greater volume of data samples to effectively learn the alignment.
Our cherry model trained on WizardLM data also has a better or comparable performance compared with the reimplemented WizardLM model on most of the subcategories. Specifically, Our model underperforms in Math, Code, Complex Format, and Counterfactual. The main reason our model loses in these categories is the abundance of training data for these categories in the original dataset and the supreme abilities of the original WizardLM in these tasks, which is mentioned in Xu et al. (2023). As a consequence, when we reduce the number of data used, our model can not be trained on these data-needed categories as much as the original model, thus leading to a relatively incomparable performance.
Appendix C Results with Official WizardLM
In this section, we provide the results of using of the WizardLM data to have a comparable performance with the official WizardLM model in a relatively unfair setting. The official WizardLM is uncensored and trained with the max token size of , while our model is trained with the max token size of , representing an inherent disadvantage of our model. However, even with this situation, our model can still reach a comparable performance with the official WizardLM model, inferring the effectiveness of our method.
Appendix D Cherry Example Analysis
To illustrate the implications of our findings and demonstrate the characteristics of the data selected by our method, we provide several examples in Figure 9.
The first positive example presents the situation that both the direct answer score (DA) and the conditioned answer score (CA) are relatively high. In this situation, the high DA means that it is hard for the initial pre-trained LLM to generate this poem, and the high CA means given the instruction does not make the generation of this poem much easier. So it is valuable for LLM to learn this sample. The second positive example presents the situation that both the CA score and DA score are relatively low. The low DA score means that LLM has learned this knowledge it is easy for LLM to generate this sentence. However, providing the corresponding instruction does not change the situation much, indicating the poor ability to follow this instruction.
The first negative example presents a situation where the response is too short. Due to the intrinsic nature of next token prediction that longer texts tend to have lower perplexity, the DA score is relatively high for the response that is too short and thus causes the IFD Score large, which we believe is a good feature of our method. The second negative example presents a situation where the DA score and CA score are relatively small. In this example, the response is quoted from a book that LLM must have read, thus as a known knowledge, it is easy for LLM to reproduce this sentence. However, with an instruction included, the CA score becomes even much lower, indicating LLM has gained quite a good ability in following this instruction. The third example presents the most common situation, where the instruction is simply not difficult enough.
Appendix E Additional Discussion
In our method, efforts are conducted to keep the pre-experience process as simple as possible, however, there still exists a question of whether the fully-trained model can be the pre-experienced model for selecting the cherry samples. To better illustrate this question, the fully-trained Alpaca model is utilized as the pre-experienced model for selecting the cherry data, , , and of the cherry data are selected and the corresponding cherry models are trained. The performances are shown in Table 7, in which the models with the fully-trained Alpaca hardly surpass the Alpaca with fewer data and our models. This experiment proves that the fully-trained model is not appropriate in selecting samples for the initial raw model, which is caused by the overly distribution gap between the fully-trained models and raw models.
E.2 How Many Cherry Samples are Required?
While extensive experiments with our method on Alpaca and WizardLM prove the effectiveness of our method in selecting high-quality samples from the original target dataset automatically, it is still under-exploring how much data is optimal. Unlike Chen et al. (2023) in which the scores of target samples are scarce, the dense scores from our method provide better flexibility in deciding how much data you can use. However, this flexibility is also a curse that makes it hard to conclude the optimal number of data to select, which is influenced by various factors including the absolute values of the IFD scores, the distribution of hard examples, and the number of data in original datasets. However, from our empirical study, we think selecting samples with the top IFD scores would be a safe and reasonable choice.
Appendix F Prompt for Evaluation
In this section, we provide the detailed prompt we used for evaluating the performance of two responses for the same instruction as shown in Figure 10.
Appendix G Detailed Main Comparison
As shown in Figure 11, we present the detailed comparison between our cherry models with the official Alpaca (7B) model across different test set with different percentage of cherry data, from to , using ChatGPT as the judge. Starting from of the full data, our cherry models outperform the official Alpaca model in all these data scales.
G.2 Comparison with the Reimplemented WizardLM
As shown in Figure 12, we present the detailed comparison between our cherry models with the reimplemented WizardLM (7B) model across different test set with different percentage of cherry data, from to , using ChatGPT as the judge. Our cherry models begin outperforming the reimplemented WizardLM from the scale of of the data.
G.3 Comparison with the Official WizardLM
As shown in Figure 13, we show the detailed comparison between our cherry models with the reimplemented WizardLM (7B) model across different test set with different percentage of cherry data, from to , using ChatGPT as the judge. When compared with the official WizardLM data, our cherry model achieves a comparable performance when using of the WizardLM data, which is positive considering the inherent disadvantage of our training configuration.
Appendix H Detailed Ablation Comparison
As shown in Figure 14(a)(b)(c), we show the detailed comparison between the models trained with randomly selected data with our cherry models across different test set with different percentage of data, from to , using ChatGPT as the judge. From to of the data, our cherry models consistently outperform the random models.
H.2 Data with Low IFD Score
As shown in Figure 15, we show the detailed comparison between the models trained with data selected with low IFD scores with our cherry models across different test set with different percentage of data, from to , using ChatGPT as the judge. From to of the data, our cherry models consistently have better performances.
H.3 Data with High CA Scores
As shown in Figure 16, we show the detailed comparison between the models trained with data selected with high conditioned answer scores with our cherry models across different test set with different percentage of data, from to , using ChatGPT as the judge. From to of the data, our cherry models consistently have better performances.
H.4 Number of Pre-Experienced Data
Figure 17 shows the comparisons when different numbers of pre-experienced samples are utilized to train the pre-experienced model.
H.5 Distribution of Pre-Experience Data
Figure 18 shows the comparisons when IFD scores are used as the strategy to select pre-experienced data to train the pre-experienced model.
H.6 Fully-trained Model as Pre-Experienced Models
Figure 19 shows the detailed comparisons when the fully-trained official Alpaca is utilized as the pre-experienced model for selecting cherry data.
Appendix I More Examples
In this section, some positive examples with top IFD scores in the Alpaca dataset are presented in Figure 20 and 21. Negative examples with the least IFD scores are presented in Figure 22.