Reflection-Tuning: Data Recycling Improves LLM Instruction-Tuning

Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Heng Huang, Jiuxiang Gu, Tianyi Zhou

Introduction

Recently, the emergence and rapid advancement of Large Language Models (LLMs) have pushed the boundaries of natural language understanding and generation. These models have been applied to a variety of applications , from content generation to answering complex questions. A salient feature of LLMs is their potential to follow instructions given to them, a characteristic that has been harnessed to fine-tune and control their outputs. This process, commonly referred to as instruction tuning , holds immense promise for customizing LLMs to specific tasks or preferences.

However, instruction tuning is susceptible to the quality of training data. Introducing suboptimal data into the training process can have a cascade of adverse effects. Within the ambit of natural language generation, empirical research delineates that both the integrity and the homogeneity of training data critically modulate the fluency, pertinence, and precision of the generated linguistic content . Datasets exhibiting inconsistencies or subpar quality can precipitate models to engender erratic, prejudiced, or even specious outputs, thereby attenuating their dependability and applicability. Analogous issues permeate instruction-tuning environments. Recent research underscores that even a minuscule fraction of skewed virtual prompts can severely impinge upon a model’s operational efficacy, manifesting the susceptibility of large language models (LLMs) to inferior data. On the other hand, ALPAGASUS and Cherry LLM demonstrate that LLMs can achieve enhanced performance metrics by leveraging a select subset of high-quality data.

To address this identified challenge, we introduce a novel method engineered to enhance the quality of extant instruction-tuning datasets autonomously. Drawing inspiration from the evaluative proficiencies of LLMs and contemporary paradigms in self-enhancement , our approach hinges on employing an oracle model to introspectively assess and improve the current dataset against specific criteria. This process of data refinement, which we term “reflection-tuning”, constitutes a potent and efficacious mechanism to bolster the quality of instruction-tuning data. Crucially, this approach obviates the need for supplementary model training and boasts universal adaptability to diverse instruction-response pair architectures. While analogous methodologies have been broached in recent self-alignment literature – typified by their application of the model for its own enhancement or in aligning model outputs with preconceived critiques – our contribution is pioneering in integrating the reflection and modification paradigm to both instruction and response dimensions, thereby facilitating the genesis of superior instruction-tuning datasets.

Our extensive experiments include comprehensive evaluations of the models trained with reflection-tuning, including the instruction-following evaluations, e.g., Alpaca-Eval, some human-instruction test sets, and benchmarks. Since GPT-4 demonstrates higher agreement with human preferences than agreements between humans , we utilize it as our judge for our main instruction-following evaluations. In the comparison with the models trained with the original datasets, e.g., Alpaca , WizardLM , our reflection-tuned models achieve much better performance. Specifically, our recycled WizardLM 7B model achieves the highest win rate among other open-source 7B models in the Alpaca-Eval leaderboard. Moreover, Our recycled Alpaca achieves a win rate of 88.75%88.75\% and our recycled WizardLM achieves a win rate of 81.25%81.25\% on the Vicuna test set with the same number of training data and model size.

Related Work

The overarching goal of our work is to enhance the model’s instruction-following capability, which is consistent with the previous works . It is discovered that the cross-task generalization ability of LLMs could be enhanced by fine-tuning on NLP datasets which are structured with instruction-response pairs . More recent works have expanded instruction tuning to include open-ended generation tasks, which exhibit enhanced handling of complex human instructions.

Our method also targets generating better instruction tuning data , but it is orthogonal to the previous work since any kind of instruction-response pairs can be further reflected and improved by our method. Recent works either curate the instruction tuning datasets by human labors, e.g., Dolly , Longpre or distill the responses from SOTA LLMs like GPT4 , e.g., Alpaca , Alpaca-GPT4 , Vicuna , Koala . There is also some exploration of making the instructions more difficult through the evolution , which achieves incredible performance on Alpaca-Eval . Different from them, our method could be treated as a useful posthoc tool, which can further enhance the quality of the instruction tuning data.

Our study contributes to the expanding body of self-alignment , i.e., it proves the self-check and self-refine ability of the LLMs. Constitutional-AI first introduces the idea of using the feedback of the AI itself as the preference data to optimize the objectives of helpfulness and harmlessness. Recent works show that LLMs can generate useful signals for debugging, filtering, and finetuning with RL. These works inspire our study prompting the ChatGPT to self-reflect its own generated responses and then self-revise.

Methodology

Initially, we elucidate and formalize extant methodologies that leverage large language models for instruction-tuning. Let fθf_{\theta} denote the pre-trained LLM, e.g., Llama, with parameters θ\theta and gg the oracle LLM, e.g., ChatGPT. We use other lowercase letters x,y,z,c,..x,y,z,c,.. to denote the text segments, which could be phrases or sentences, and each token in xx is denoted as x[i]x[i]. We use uppercase letters D,..D,.. to denote the collection of language sequences or datasets, and D0D_{0} represents the initial base dataset. Since both fθf_{\theta} and gg are in auto-regressive manners, a sequence x=(x,...,x[n])x=(x,...,x[n]) can be further denoted as fθ(x)=∏i=1nf(x[i]∣x[1,...,i])f_{\theta}(x)=\prod_{i=1}^{n}f(x[i]|x[1,...,i]).

In the instruction-following setting, there will be a mapping function that turns the original raw instruction xx into the desirable format and requests models for a response yy. For simplicity, we directly notate this process as y∼f(y∣x)y\sim f(y|x). And the loss function for instruction-tuning can be denoted as L=−1n∑i=1nlog⁡fθ(y∣x)L=-\frac{1}{n}\sum_{i=1}^{n}\log f_{\theta}(y|x) where nn is the length of response yy.

2 Reflection-Tuning

There are two main phases in our method, instruction reflection and response reflection. Based on the intuition that students who reflect on the answers usually get higher scores since they can find the errors and make some reasonable changes through the reflection process, and astonished by the self-improvement and judging capability of LLMs, we propose a reflection method for improving the quality of instruction-response pairs. Given the initial base dataset, we are motivated to generate a high-quality version of each data point with an oracle model, ChatGPT for instance. However, a common problem with using LLMs as judges is the failure to obtain diverse results. To overcome this potential problem, inspired by Chain-of-Thought and Tree-of-Thought prompting , we further define several specific criteria {c1,...,ck}\{c_{1},...,c_{k}\} for the oracle model to follow, and respond to those specific criteria with critical responses {z1,...,zk}\{z_{1},...,z_{k}\}, respectively. Then the responses to these criteria can bridge the generation of new instruction-response pairs.

Specifically, in the instruction reflection phase, the oracle model gg is required to reflect on the given instruction-response pair (x0,y0)(x^{0},y^{0}) from the original dataset D0D^{0} with some specific criteria {c1ins,...,ckins}\{c_{1}^{ins},...,c_{k}^{ins}\} and then generate a better instruction-response pair (xins,yins)(x^{ins},y^{ins}) according to its reflection results. With the criteria given, the oracle model gg is able to generate critical responses:

where both original instruction and response are wrapped into the prompt rather than original instruction alone. These critical responses further serve as the guidance (chain of thought) for the generation of the new instruction and response pair:

where in practice the above process is sampled as a continuous language sequence, and the critical responses would not be decomposed from the whole outputs. The criteria used for instruction are “the Complexity of the Topic”, “the Level of Detail Required for response”, “Knowledge Required for response”, “the Ambiguity of the Instruction” and whether “Logical Reasoning or Problem-Solving Involved”.

2.2 Reflection on Response

Although both instruction and response are modified, the corresponding response yinsy^{ins} for a given modified instruction xinsx^{ins} is not optimal. Thus another reflection on the response process is further proposed. Similar to the above procedure, a new set of criteria for reflection on response is defined as {c1res,...,cmres}\{c_{1}^{res},...,c_{m}^{res}\}. The overall process can be noted as:

where ziresz_{i}^{res} represents the critical response of iith response criteria ciresc_{i}^{res}. After the above process, the instruction and response pair (xins,yres))(x^{ins},y^{res})) is regarded as the recycled data pair which will be used for instruction-tuning of model fθf_{\theta}. The criteria used for instruction are “Helpfulness”, “Relevance”, “Accuracy”, and “Level of Details”.

We name the whole above process as a recycling process, which greatly improves the quality of the previous dataset. Then the raw model fθf_{\theta} will be trained on the newly generated recycled dataset, and the newly generated models are notated as “Recycled Models”, eg. Recycled Alpaca.

Experimental Setup

The Alpaca dataset , sourced from Stanford University, offers 52,00252,002 instruction-following samples. Developed via the self-instruct paradigm , it leveraged the capabilities of the text-davinci-003 model. This dataset, while a pioneering attempt in instruction tuning for the LLaMA model, raised concerns about data quality owing to its reliance on the text-davinci-003 model.

On the other hand, the WizardLM dataset , which employs the sophisticated Evol-Instruct algorithm, is a refined collection encompassing a total of 250,000250,000 instruction samples. Two primary evolutionary trajectories, namely "In-depth Evolving" and "In-breadth Evolving", are introduced within this dataset. These trajectories are specifically designed to allow a base instruction to progress either in terms of intricate details or in its overall scope. To enhance data fidelity, ChatGPT has been meticulously integrated during the refinement process. From this extensive dataset, we predominantly focused on the WizardLM-7b subset, comprising 70,00070,000 samples. We test our method on both of these two datasets to verify the effectiveness of our method.

2 Implementation Details

Rooted in the Llama2-7b pre-trained model , we utilize the prompt and code base from Vicuna and flash attention while the overall training arguments are aligned with protocols from Alpaca and WizardLM datasets. The Adam optimizer , with a 2×10−52\times 10^{-5} learning rate and a batch size of 128128, steers the training across three epochs with a max length of 20482048. The warmup rate is set to 0.030.03.

3 Evaluation Metric

The task of quantitatively evaluating the instruction-adherence efficacy of LLMs presents considerable challenges. Despite a wealth of research endeavoring to design automated evaluation metrics for LLMs , the gold standard remains subjective human evaluation. However, such manual assessments are not only resource-intensive but are also susceptible to inherent human biases.

Incorporating methodologies from cutting-edge LLM evaluations , we operationalize GPT4 and ChatGPT as evaluation benchmarks. As delineated in , models subjected to evaluation are prompted to generate outputs for each instruction in the test corpus. Subsequent to this, an API-driven model, be it GPT4 or ChatGPT, allocates a score to each response. A model’s superiority on this dataset hinges on its endorsement by the adjudicating model.

The adjudication phase entails rating each model-generated response on a scale spanning from 11 to 1010, with scores encapsulating facets such as pertinence and precision. To mitigate the positional bias elaborated upon in , model-generated outputs are presented to the adjudicating entity in two distinct sequences and subsequently scored. Hence, a model’s dominance is ratified under the following conditions: Wins: Exhibits superiority in both sequences or prevails in one while maintaining parity in the alternate sequence. Tie: Demonstrates parity across both sequences or prevails in one while faltering in the alternate. Loses: Underperforms in both sequences or maintains parity in one while being eclipsed in the alternate. This adjudication paradigm underpins our experimental findings.

4 Benchmarks

Two prominent benchmarking platforms for LLMs are highlighted: the Huggingface Open LLM Leaderboardhttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard and the AlpacaEval Leaderboardhttps://tatsu-lab.github.io/alpaca_eval. The Huggingface Open LLM Leaderboard employs the evaluation methodology from , providing a cohesive framework for assessing generative language model capabilities across a spectrum of evaluation tasks. It focuses on 44 pivotal benchmarks: ARC , HellaSwag , MMLU , and TruthfulQA . Specifically, ARC is a specialized dataset curated for assessing the proficiency of models in answering science questions tailored for grade-school levels. The challenge employs a 25-shot learning paradigm, implying that models are exposed to 25 examples prior to evaluation. HellaSwag is Specifically designed to probe models on their commonsense inference capabilities, which utilizes a 10-shot learning setup, meaning models are trained on 10 sample instances before being tested. MMLU is a comprehensive evaluation suite designed to gauge a model’s multitasking learning capability across a diverse range of 57 tasks. These tasks span a myriad of domains including but not limited to elementary mathematics, US history, computer science, and jurisprudence. TruthfulQA is constructed to appraise a model’s susceptibility to perpetuating misinformation or falsehoods, which are ubiquitously found online.

On the other hand, the AlpacaEval Leaderboard offers an LLM-centric automatic assessment utilizing the AlpacaFarm evaluation dataset. It is an automated evaluation mechanism for LLMs that offers efficiency, cost-effectiveness, and reliability. Operating on the AlpacaFarm evaluation dataset, it gauges models’ proficiency in adhering to generic user instructions. The generated outputs are juxtaposed against benchmark responses from Davinci003. These benchmarks are subsequently auto-annotated by either GPT-4, Claude, or ChatGPT, leading to the determination of the aforementioned win rates. Empirical evidence suggests that AlpacaEval’s alignment with ground truth annotations sourced from human experts is notably high. Furthermore, model rankings on the AlpacaEval leaderboard exhibit a strong correlation with rankings derived from human annotators.

Experimental Results

As depicted in Figure 1, a juxtaposition between our recycled models and other distinguished models is presented. Remarkably, our models exhibit superior performance across the board, with GPT4 being the sole exception, underscoring the efficacy of our methodology. Notably, SelFee aligns with our motivation in leveraging an oracle model to refine dataset responses while using much more data for training including the Alpaca dataset, the ShareGPT dataset, the FLAN dataset, and extra math and code collections. However, even with much more data used, they overlook the criticality of enhancing the instruction set and neglect the deployment of granular criteria for self-enhancement. This negligence results in their suboptimal performance despite a voluminous training dataset. Importantly, our models, equipped solely with instruction tuning on the Alpaca dataset, surpass several counterparts that employ additional RLHF techniques.

2 Alpaca Eval Leaderboard

Table 1 delineates the outcomes on the AlpacaEval Leaderboard. Within this evaluation framework, GPT4 is harnessed as the adjudicating entity, contrasting the responses of the test models against the benchmark set by Davinci003. This comparison provides a direct quantification of a model’s capacity for instruction adherence and the intrinsic quality of its output. Notably, our models eclipse the performance of all extant 7B open-source counterparts, with the sole exception being Xwin-LM whose training data is unknown and extra RLHF is implemented. Remarkably, our models even surpass some of the models with a larger parameter count. The eminent positioning of our models on this leaderboard underscores the superior caliber of the responses they generate.

3 Open LLM Leaderboard

Table 2 showcases the performance comparison on the Huggingface Open LLM Leaderboard with some related models. With our Recycle mechanism, our models achieve better average performances across these four representative benchmarks and our results are comparable to llama-2-7b-chat, which is elaborately fine-tuned with extra RLHF.

Discussion

In the ensuing discourse, we delve into a quantitative juxtaposition of the instruction-response data, pre- and post-application of our recycling methodology, as delineated in Table 3. Observationally, there’s an increase in the average token length of instructions within the Alpaca dataset, whereas a decrement manifests for the WizardLM dataset, epitomizing the method’s adept adaptability. The succinctness and elementary nature of the Alpaca dataset’s instructions warrant an enhancement in intricacy through our method, thereby elongating their length. Conversely, the pre-existing complexity and intricacy in WizardLM’s instructions render our algorithm inclined towards succinctness. Pertaining to the response section, there’s a marked propensity of our approach to engender detail-rich textual content, leading to relatively long responses. Moreover, leveraging Sentence-BERT , we quantify the coherence metric between instructions and their affiliated responses. It’s discernible that our technique invariably fabricates samples with better coherence, signifying a superior alignment between modulated instructions and consequent responses. Additionally, to elucidate the metamorphosis in instructional difficulty, we employ the Instruction-Following Difficulty (IFD) score, as posited by Cherry LLM , executed on the nascent pre-trained language model. This score gauges the efficacy of instructions in bolstering response predictions. The consistent ascension in IFD scores lucidly illustrates our instruction’s progressive evolution.

2 Performances on 13B Models

We further train a Recycled Alpaca in the 13B version to further validate the efficacy of our method. With only 5252k recycled alpaca data being used for instruction-tuning, our Recycled Alpaca 13B reaches the win rate of 83.42%83.42\% in the Alpaca Eval leaderboard and reaches an average score of 58.93%58.93\% on Huggingface Open LLM leaderboard. Considering the small amount of data we used compared with other models, the results are intriguing and satisfactory. We will soon apply our recycled WizardLM data to the 13B model.

Conclusion

The evolution of Large Language Models has brought forth unparalleled capacities in natural language processing, especially in the domain of instruction tuning. However, the quality of training data remains a pivotal determinant of model performance. In this work, we introduced the reflection-tuning method, an innovative approach to autonomously improve and recycle the quality of instruction-tuning datasets by leveraging the inherent self-improvement capabilities of LLMs. Our method emphasizes a unique reflect-and-recycle mechanism, a first in the domain, applied comprehensively to both instructions and responses. Experimental results affirm the efficacy of reflection-tuning, with models trained using this method consistently outperforming those trained with traditional datasets. This paves the way for more reliable, consistent, and high-performing LLMs in the future, underscoring the importance of high-quality data recycling and innovative methods in the realm of natural language generation.

References