LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
Shizhe Diao, Rui Pan, Hanze Dong, Ka Shun Shum, Jipeng Zhang, Wei Xiong, Tong Zhang
Introduction
Large foundation models, and in particular large language models (LLMs), have demonstrated general abilities to perform different tasks beyond what was possible previously. However, for specialized domains or tasks, it is necessary to further finetune such LLMs to achieve improved performance on such domains or tasks. The typical processes to finetune such large models include:
Continuous pretraining on special domains so that a large foundation model can acquire knowledge on these domains.
Instruction tuning to teach a large foundation model the capability to follow these specialized natural language instructions and perform tasks required by such instructions.
Reinforcement learning with human feedback (RLHF) to teach a large foundation model skills to perform conversation according to human preference.
While a number of pretrained large models, including GPT-J , Bloom , LLaMA , etc., are publically available and have already been incorporated into the Hugging Face model repository , there is no publically available toolkit that can be easily used to perform finetuning tasks for these different models. The purpose of this package is to offer a simple-to-use and lightweight toolkit so that developers and researchers can perform efficient finetuning and inference of large models with limited resources.
The following key features are supported by the toolkit:
Continous pretraining, instruction tuning, and RLHF on user-defined datasets.
Simple and extensible APIs for developers.
Efficient tuning with low-rank adaptation (LoRA).
A novel RLHF algorithm RAFT (Reward rAnked FineTuning) to simply RLHF pipeline for generative models.
Based on a 7-billion-parameter LLaMA model, it only takes one Nvidia 3090 GPU and five hours to train a personalized model. We used this framework to finetune a series of 7-billion, 13-billion, 33-billion, and 65-billion parameter versions of LLaMA on a single machine and have released the model weights for academic research. The trained model weights can be immediately used for a question-and-answer service on the website lmflow.com.
Using LMFlow, anyone can train their own personalized model. Each person can choose the appropriate model according to their available resources, for tasks such as question answering, companionship, writing, translation, and expert consultations in various fields. The larger the model and data size, the longer the training time provided the better the results. Currently, we trained a 33B model and achieved comparable or even better performance than ChatGPT.
Toolkit Overview
An illustration of the LMFlow system design is shown in Figure 1. There are four stages for improving the performance of a publicly available large language model. The first stage is domain adaptation, which involves modifying the model to better handle a specific domain by training the model on that domain. The second stage is task adaptation, which involves adapting the model to perform a specific task, such as summarization, question-answering, and translation. The third stage is instruction finetuning, which involves adjusting the model’s parameters based on instructional question-answer pairs. The final stage is reinforcement learning with human feedback, which involves using human feedback to further align the model to human preference. LMFlow provides a complete finetuning workflow for these four stages, supporting large language models’ personalized training with limited computing resources.
2 Installation
LMFlow has been fully tested on Linux OS (Ubuntu 20.04) and can be installed by executing the following commands.
3 Data Format
LMFlow accepts several .json files as input. Users can provide a list of .json files under a specified dataset directory. For example,
Each json file shall have the following format (three instances with four keys for example),
where the TYPE indicates the dataset type and defines the set of keys { KEY_1, KEY_2, ... } and their corresponding interpretations. A list of supported types is detailed as follows.
This is the most common dataset type, which only contains raw texts in each sample. This type of dataset can be used as the training set for text decoder models, or the input of decoder models / encoder-decoder models. Its format is as follows (three instances, for example),
Text2Text
This is the dataset type mostly used for inferencing, which contains a pair of texts in each sample. This type of dataset can be used as the training set for text encoder-decoder models, or question-answer pair for evaluating model inferences. Its format is as follows (three instances for example),
4 Continuous Pretraining
The endeavor to bridge the divide between pretraining domains and downstream domains has led to the adoption of a prevalent approach, known as continuous pretraining , which involves the ongoing pretraining on an extensive collection of unlabeled data that is specific to a given domain. Continuous pretraining is LMFlow supports continuous pretraining natively, which is an effective way to adapt LLMs to a specific domain. Users just need to collect a set of unlabeled data and prepare them to TextOnly data format. The following process will be handled by autoregressive training.
5 Instruction Tuning
Instruction tuning , also called supervised finetuning, is an approach used to enhance the performance of language models by training them to follow natural language instructions. This involves training the model on a small set of task-specific data, most of which are in prompt-answer format, including positive or negative examples, prompts, constraints, and other elements commonly present in human language. The primary objective of instruction tuning is to improve the model’s proficiency in undertaking multiple tasks and to generalize more effectively to new or unseen tasks. This is accomplished by teaching the model to comprehend and integrate various language cues and constraints relevant to the given task. By improving the language models’ ability to comprehend and follow natural language commands, this approach can unlock new levels of performance and productivity in diverse applications. Instruction tuning enables LLMs to provide more accurate and relevant responses to user queries, making them a more effective conversational agents.
6 RLHF as Finetuning
Large language models (LLMs) are often pretrained to replicate the vast amount of text available on the internet, which unfortunately includes the generation of text that would not align with human preferences . Examples of such content include falsehoods, offensive comments, or even harmful texts. However, there is a growing need to explore alternative pretraining objectives that can guide LLMs to generate text that aligns with human preferences. By doing so, we can ensure that LLMs produce text that is more helpful, honest, and harmless for humans, which are called ‘HHH’ rules . divides the alignment process into three steps, including SFT, reward modeling, and RLHF (reward optimization). We have integrated reward modeling into our LMFlow framework. For reward optimization, PPO has been shown to be effective in various studies . However, it relies on a trial-and-error approach through interaction with the environment, making it less stable and efficient than supervised learning . A more feasible option for finetuning generative models may be to use a reward function instead of a pre-determined supervised dataset, especially when collecting high-quality samples. To address this, we propose a new alignment method for generative models called RAFT . RAFT utilizes a reward model to rank the output of the generative model, allowing us to continue training using supervised finetuning (SFT)-like techniques with the selected samples. This approach encourages the generative model to prioritize samples with higher rewards and offers significant computational advantages over PPO, resulting in substantial savings in memory and gradient computations. Moreover, due to the stability of SFT-like training, our approach demonstrates lower sample complexity and requires fewer learnable parameters, making it easily adaptable to any generative model. We believe that our novel alignment algorithm represents a competitive and innovative approach that contributes to the well-behaved behavior of generative models.
7 Efficient Tuning
LMFlow supports low-rank adaptation (LoRA) tuning based on the implementation of huggingface/peft https://github.com/huggingface/peft. LoRA is an efficient tuning method that involves freezing the weights of the pretrained model and incorporating trainable rank decomposition matrices into each layer of the Transformer architecture. This approach significantly reduces the number of trainable parameters.
8 Inference
LMFlow developed an easy-to-use inference interface for LLMs. Based on Deepspeed https://github.com/microsoft/DeepSpeed, LMFlow supports parameter partitioning with zero-offload strategies as introduced by .
In LMFlow, the inference interface is provided by an inferencer class. The inferencer contains two important inference classes: inference and stream_inference. The distinction lies in whether the output is printed word by word in real-time.
API Documentation
Please refer to https://optimalscale.github.io/LMFlow/autoapi/index.html for the details of API documentation.
Case Studies
In this section, we will provide case studies of LMFlow in task tuning, instruction tuning, and alignment tuning.
The aim of task tuning is to enhance the proficiency of a language model in a specific field, such as the medical or financial domain, by imparting domain-specific information that allows it to better adapt to the target subject matter. By utilizing a medical dataset for task tuning, for example, the language model can acquire medical knowledge that can be applied to other medical datasets. To highlight the importance of this approach, we employed task tuning on LLaMA models in medical domain to assess their performance. The evaluations on three medical datasets revealed significant enhancements in both in-domain (PubMedQA , MedMCQA ) and out-of-domain (MedQA-USMLE ) datasets. The LLaMA-33B (LoRA) performance is achieved with only about 16h finetuning on the training split of PubMedQA and MedMCQA with a single 8 * A100 server.
2 Instruction Tuning
Following previous work in instruction tuning , we finetune the model with the instruction-following data. Expanding upon the initial idea of self-instruct techniques, we incorporated several different data sources and build a new dataset called LMFlow Datasethttp://lmflow.org:5000/lmflow_data.tar.gz. The new training split is created by merging the following datasets:
ShareGPT: randomly sample 50K English data and 10K Chinese data from ShareGPT https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered.
GPT-4-LLM : 52K English data from GPT-4-LLM https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
BELLE : randomly sample 80K Chinese data from BELLE https://github.com/LianjiaTech/BELLE.
This data fusion takes the Chinese and English data balance into consideration. Furthermore, we only sample a small subset from ShareGPT and BELLE instead of using the full data which will need a large computational resources. We call our instruction-tuned model Robin Robin is a small passerine bird that belongs to the family Turdidae. Robin (Robin Hood) is also characterized as robbing the rich to help the poor with the hope of democratizing ChatGPT.. Based on LMFlow Dataset, we trained Robin-7B-v2, Robin-13B-v2, Robin-33B-v2 and Robin-65B-v2 based on the respective LLaMA base model. The delta weights of Robin are released at https://github.com/OptimalScale/LMFlow#model-zoo.
In order to evaluate the models’ instruction-following ability, we participate the Huggingface Open LLM Leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard. The performance is shown in Table 2. Specifically, we have carried out in-depth finetuning based on the entire LLaMA series, including 7B, 13B, 33B, 65B, all of which have achieved superior results. Robin-7B-v2 scored 51.7 in the OpenLLM standard test, and Robin-13B even reached as high as 59.1, ranking sixth, surpassing many 33B models. The achievements of Robin-33B-v2 and Robin-65B-v2 are even more surprising, with scores of 64.1 and 65.2 respectively, firmly securing the top positions.
In addition, we collected GPT-4 instruction data from GPT-4-LLM , which provides many instruction tuning data labeled by GPT-4 and create a test set by sampling 1,000 English data. We manually filtered examples with the following issues, where 767 effective samples remain after the filtering:
Long response with too many nonsense words
Specific domains involving chemistry/biology, where most LLM models do not possess the knowledge and always fail
We compare Robin-7B with Vicuna-13B on this test set. The case study is shown in Figure 2.
3 Alignment Tuning
We conduct an experiment on the HH-RLHF (Helpful and Harmless) datasethttps://huggingface.co/datasets/Dahoas/full-hh-rlhf , which is collected for model alignment according to human preferences. The dataset consists of 112K training samples and 12.5K test samples. Each sample of the HH-RLHF dataset consists of a prompt , which is a chat history between the “Human” and “Assistant”, and two responses and from the “Assistant” to the prompt where is the preferred compared to . Following , we first finetune the LLaMA-7B base model on the training set with the preferred responses to get the LLaMA-SFT model. To model human preference, we train a reward model based on GPT-Neo-2.7B. Then, we use RAFT to align the LLaMA-SFT model to get the aligned model LLaMA-RAFT.
For comparison, we use LLaMA-SFT and also LLaMA-PPO aligned by the PPO as two competitors. The evaluation metrics of these models are reported in Table 3. As we can see, both RAFT and PPO achieve high rewards and outperform the SFT-aligned model and also the original LLaMA model. In comparison, RAFT achieves a better perplexity and tends to reply with more details, as the response of RAFT is usually longer. We present representative examples with randomly sampled prompts in Figure 4.
It is worth noting that the RAFT training is very robust and the resulting models achieve stable performance across three independent experiments. In contrast, the PPO training requires a complicated hyper-parameter tuning process and the training can fail sometimes.
LMFlow Benchmark
Assessing the performance of chat-style large language models (LLMs) has been a significant challenge since the emergence of ChatGPT. Researchers and developers require a reliable method to compare two models and determine which one is better suited for a particular application scenario. Additionally, monitoring the model’s performance during training is essential to prevent issues such as forgetting. A recent study by Vicuna introduced human evaluation comparison methods, also known as Chatbot Arenahttps://chat.lmsys.org/?arena, and pioneered the use of GPT-4 to compare the outputs of two models. However, human evaluation is costly and not scalable for LLM development due to the expensive human labeling. Furthermore, taking GPT-4 as a referee suffers from a position bias , and simply changing the order of candidates could skew the evaluation result.
To address these issues, we present the LMFlow benchmark, a new benchmark that offers an affordable and user-friendly evaluation framework that can reflect various aspects of LLMs. We have open-sourced the dataset and codehttps://github.com/OptimalScale/LMFlow, enabling the LLM community to use these toolkits to evaluate and compare different LLMs.
In our evaluation framework, negative log likelihood (NLL) is used for evaluating LLM
The NLL metric measures the prediction probability of the LLM model over a corpus set based on its context. If the corpus set is indicative of a specific type of LLM capability, such as multi-round conversation, instruction following, math problem solving, or role-playing, then the NLL metric on those corpora can offer quantitative measures to assess those abilities.
Besides NLL, another similar and commonly used metric in NLP is perplexity (PPL):
However, perplexity is inherently biased toward the lengths of tokenized sequences, leading to unfair comparisons between models that use different tokenizers. For instance, a model with a smaller vocabulary size will result in longer tokenized sequences and lower token-level perplexity. Therefore, we used NLL instead of PPL in all our experiments. NLL evaluation has a significant advantage in that it does not require human involvement during the evaluation process. As long as the test reference corpus is provided, researchers can automatically evaluate various aspects of an LLM’s ability. This feature makes the evaluation of LLMs more accessible to researchers. Furthermore, NLL is an excellent metric in its own right. In our commonsense QA experiments, we discovered that NLL is correlated with QA accuracy when comparing different finetuned versions of a single model. In Figure 3, it is observed that the accuracy of QA is roughly correlated to NLL. Therefore, we claim that NLL is a good metric to reflect the magnitude of prediction level difference between models, where a huge gap in NLL normally entails a huge performance gap.
Conclusion
In conclusion, while large foundation model models have shown significant promise in general applications, further finetuning is often required for specialized domains or tasks. This is where the LMFlow toolkit comes in, offering an extensible, lightweight, and easy-to-use solution for developers and researchers to perform efficient finetuning and inference of large models with limited resources. With features such as continuous pretraining, instruction tuning, and RLHF, as well as simple and extensible APIs, LMFlow provides a complete finetuning workflow for large models. Moreover, with the ability to personalize training and achieve comparable or even better performance than ChatGPT, LMFlow represents a significant step forward in the development of large foundation models and their application to specialized tasks.