Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, Yongbin Li
Introduction
Human beings have harbored a longstanding desire to acquire additional abilities through various ways, as expressed in mediums like movies and games. For example, in X-Men’s Apocalypse, the character can absorb the powers of other mutants to strengthen himself. Likewise, the protagonist in the Super Mario games can gain superpowers like throwing fireballs by absorbing in-game items. In this paper, we astonishingly find that Language Models (LMs), similar to Apocalypse and Super Mario, can enhance their capabilities by absorbing other models without the need for retraining or even GPUs.
Formally, Supervised Fine-Tuning (SFT) is the most widely adopted strategy for unlocking task-specific abilities to LMs by optimizing their parameters Dodge et al. (2020); Zhao et al. (2023). The effectiveness of SFT is fully evident in the alteration of the model parameters before and after SFT, referred to as delta parameters Ding et al. (2023). We first show that SFT LM (either encoder- or decoder-based) always tends to acquire excessively redundant delta parameters. To be specific, we present DARE (Drop And REscale), which randomly sets certain delta parameters to zeros with a drop rate and subsequently rescales the remaining ones by a factor . Although conceptually simple, DARE can eliminate up to 99% delta parameters with minimal impact on the performance when the LM’s parameters reach 70 billion (see Figure 1(a)). Moreover, the more parameters the LM has, the larger it can tolerate. We attribute the effectiveness of DARE to its ability to approximate the original embeddings, which is verified both theoretically and empirically.
Furthermore, we can merge multiple homologous SFT LMs (fine-tuned from the same backbone) based on DARE without compromising their capabilities. As long as a small portion of the delta parameters remain unaffected during merging, the abilities of LMs unlocked by SFT can still be preserved. We first employ DARE to eliminate redundant delta parameters in each model before merging, which can potentially mitigate the interference of parameters among multiple models Yadav et al. (2023). Then, we apply established model merging techniques Wortsman et al. (2022); Ilharco et al. (2023); Matena & Raffel (2022); Jin et al. (2023); Yadav et al. (2023) to fuse the parameters with reduced redundancy for creating one model with diverse capabilities.
We conduct extensive experiments with encoder-based LMs on GLUE benchmark, and decoder-based LMs with three distinct abilities: instruction-following, mathematical reasoning, and code-generating. We observe that:
(1) SFT LMs exhibit a substantial number of redundant delta parameters regardless of their backbones (e.g., BERT, RoBERTa, LLaMA, Llama 2, or Code Llama). DARE can remove 90% or even 99% delta parameters without significantly affecting the model performance. DARE is able to approximate the original embeddings well and provide very similar embeddings for each layer of the LM. The rescale operation is crucial to guarantee the success of DARE, and dropping 30% or 40% delta parameters without rescaling would noticeably lead to worse results.
(2) DARE can often enhance the performance of various model merging methods on encoder-based LMs. For larger decoder-based LMs, simply averaging the parameters can already yield surprisingly good results. As shown in Figure 1(b), we merge WizardLM and WizardMath by combining DARE and parameter averaging, leading to a significant improvement of WizardLM’s mathematical reasoning ability with zero-shot accuracy from 2.2 to 66.3 on GSM8K, while also modestly enhancing its instruction-following ability with win rate from 67.2 to 67.5 on AlpacaEval. We also offer a merged LM with 7 billion parameters and it attains the top-ranking position on the Open LLM Leaderboard. It is fascinating that all these benefits are achieved by solely using CPUs without retraining.
(3) SFT delta parameters usually stay within 0.005, indicating minimal modifications to the pre-trained LM, and DARE works for delta parameters with relatively small value ranges. However, once models undergo continuous pre-training, the delta parameters can rapidly reach around 0.03, making DARE infeasible. Moreover, dropping only 10% fine-tuned parameters (i.e., the combination of pre-trained and delta parameters) would lead to a catastrophic decrease in performance, even approaching zero. This finding further confirms that SFT primarily unlocks the abilities of pre-trained LMs, rather than introducing new capabilities.
The used resources are publicly available at https://github.com/yule-BUAA/MergeLM, which integrates existing popular model merging methods and supports both encoder- and decoder-based LMs.
Related Work
Supervised Fine-tuning of Language Models. SFT of LMs aims to impart pre-trained LMs with particular abilities by optimizing them on task-specific data, which has become the de facto standard paradigm in natural language processing Dodge et al. (2020); Zhao et al. (2023). Generally, SFT can be divided into two categories: full fine-tuning Radford et al. (2018); Devlin et al. (2019) and parameter-efficient fine-tuning Houlsby et al. (2019); Liu et al. (2021); Li & Liang (2021); Lester et al. (2021); Hu et al. (2022). Indeed, the effects of SFT are reflected by the difference between parameters of LMs before and after SFT, i.e., delta parameters. In this paper, we reveal the extreme redundancy of various SFT LMs’ delta parameters by proposing an innovative approach DARE, achieving competitive performance with standard SFT LMs by removing 90% or even 99% delta parameters.
Network Pruning Technique. With the rapidly increasing size of neural networks, network pruning technique has been widely applied to reduce the computational costs Cheng et al. (2017); Liang et al. (2021). The objective of network pruning is to eliminate unnecessary parameters while maintaining the model performance Zhu & Gupta (2018); Liu et al. (2019b); Frankle & Carbin (2019); Gale et al. (2019); Xia et al. (2022). Magnitude-based pruning is one classical pruning method, which selects parameters according to their magnitudes (i.e., absolute parameter values) Han et al. (2015); Li et al. (2018); Lee et al. (2021). To be specific, parameters with magnitudes lower than a certain threshold are removed, and others are preserved. In fact, DARE is relevant to the concept of network pruning as it can also drop parameters. But DARE differs from existing pruning techniques in: (1) DARE focuses on delta parameters while most pruning methods deal with fine-tuned parameters; (2) DARE can work well without any retraining or extra data, which are often inevitably required by pruning methods.
Model Merging. Model merging has become a trending research direction in recent years, aiming to merge multiple task-specific models into a single model with diverse abilities Wortsman et al. (2022); Matena & Raffel (2022); Ilharco et al. (2023); Jin et al. (2023); Yadav et al. (2023); Zhang et al. (2023). The superiority of model merging over multi-task learning Crawshaw (2020); Zhang & Yang (2022) (which also intends to obtain one model with several abilities) is that model merging pays attention to the fusion of model parameters without accessing the original training data Matena & Raffel (2022); Jin et al. (2023). Average Merging Wortsman et al. (2022) is one common model merging approach, which utilizes averaged parameters to construct the merged model. Task Arithmetic Ilharco et al. (2023) employs a pre-defined scaling term to distinguish the importance of various models. Fisher Merging Matena & Raffel (2022) performs weighted fusions of parameters, where the weights are calculated by the Fisher information matrix Fisher (1922). RegMean Jin et al. (2023) masterly solves model merging by optimizing a linear regression problem with closed-form solutions. TIES-Merging Yadav et al. (2023) tackles the task conflicts in Ilharco et al. (2023) by trimming low-magnitude parameters, resolving sign disagreements, and disjointly merging parameters with consistent signs. In this paper, we use DARE as a versatile plug-in for existing model merging methods by first sparsifying delta parameters of several SFT homologous models and then merging them into a single model, which is equipped with the capabilities of all the SFT models.
Methodology
Model Merging Problem. Given a set of tasks and corresponding SFT models with parameters , model merging aims to fuse the parameters of models into a single model that can well handle tasks simultaneously. Following Matena & Raffel (2022); Jin et al. (2023); Yadav et al. (2023), we focus on merging fine-tuned models that are optimized from the same pre-trained backbone.
In this work, we reveal the extremely redundant properties of the delta parameters of SFT LMs and propose DARE to effectively reduce delta parameter redundancy (see Figure 2(a)). DARE is conceptually simple and consists of two steps: drop and rescale. Given delta parameters , DARE first performs random drop on based on a drop rate (setting their values to zeros) and then rescales the remaining ones by a factor as follows,
Finally, we combine and via addition to obtain the parameters for inference, i.e., . We prove that even after removing most delta parameters, DARE can well preserve the model performance by approximating the original embeddings.
Remark. We have given a rough proof of why DARE works. In practice, we find that DARE is applicable when the drop rate is properly set, and the tolerance of grows with LMs’ parameter sizes. Moreover, removing fine-tuned rather than delta parameters would cause a catastrophically decreased performance. A promising future direction is to explore DARE more deeply, such as inferring the upper bound of with respect to LM capacities and illustrating the intrinsic difference between fine-tuned and delta parameters.
Last, we highlight the connections and differences between DARE and Dropout Srivastava et al. (2014). Both methods involve random dropping and rescaling operations, but they differ in two key aspects: (1) DARE handles delta parameters while Dropout operates on model outputs; (2) DARE aims to reduce delta parameter redundancy without training, which permanently eliminates delta parameters and only retains others for inference. Dropout is used to prevent models from overfitting, which temporarily removes part of outputs during training but preserves all the outputs for inference.
2 Merging Models with DARE
As DARE effectively reduces the redundancy of delta parameters by setting most of them to zeros, we hypothesize that DARE can help address the interference of parameters when merging multiple models Yadav et al. (2023). Take Figure 2(b) as an example, when merging math- and code-related models, DARE can assist existing model merging methods to better absorb the abilities of two models with less or no parameter interference.
Formally, given models that are fine-tuned on corresponding tasks with parameters , we first apply DARE on each parameters (), and derive . Then, we adopt established model merging methods to fuse the derived parameters and obtain the merged single model. Let us take Task Arithmetic Ilharco et al. (2023) as an instance, whose official computation process is denoted by
where is the scaling term to determine the importance of the models to be merged. When equipped with DARE, the calculation process of Task Arithmetic is rewritten as
In Section 4.3, we find that DARE can effectively improve the performance of Task Arithmetic when merging multiple LMs. It is also worth noticing that DARE is a versatile plug-and-play module and can be applied to any model merging methods, such as Average Merging Wortsman et al. (2022), Fisher Merging Matena & Raffel (2022), RegMean Jin et al. (2023), and TIES-Merging Yadav et al. (2023).
Experiments
Datasets and Pre-Trained Backbones for Decoder-based LMs. We choose AlpacaEval Li et al. (2023) for evaluating instruction-following models (WizardLM Xu et al. (2023)). We use GSM8K Cobbe et al. (2021) and MATH Hendrycks et al. (2021b) for testing mathematical reasoning models (WizardMath Luo et al. (2023a)). HumanEval Chen et al. (2021) and MBPP Austin et al. (2021) are adopted for estimating code-generating models (WizardCoder-Python Luo et al. (2023b) and llama-2-13b-code-alpaca Chaudhary (2023)). These models are fine-tuned based on pre-trained backbones including LLaMA Touvron et al. (2023a), Llama 2 Touvron et al. (2023b), and Code Llama Rozière et al. (2023). Please see Table 3 in Section A.1 for their versions and correspondences with pre-trained backbones.
Datasets and Pre-Trained Backbones for Encoder-based LMs. For encoder-based LMs, the GLUE benchmark Wang et al. (2019) is used, containing one sentence acceptability dataset CoLA Warstadt et al. (2019), one sentiment detection dataset SST-2 Socher et al. (2013), two paraphrase datasets MRPC Dolan & Brockett (2005) and QQP Shankar et al. (2017), one sentence similarity dataset STS-B Cer et al. (2017), and three natural language inference datasets MNLI Bowman et al. (2015); Williams et al. (2018), QNLI Rajpurkar et al. (2016), and RTE Dagan et al. (2005); Haim et al. (2006); Giampiccolo et al. (2007); Bentivogli et al. (2009). As the test labels of GLUE are not publicly available, we split the original training data into training and validation sets with ratios of 90% and 10%. The original validation data is used as the test set. We choose bert-base-uncased Devlin et al. (2019) and roberta-base Liu et al. (2019a) as pre-trained backbones, and further fine-tune them to get SFT models on the eight datasets.
Evaluation Metrics. We calculate win rate for AlpacaEval, zero-shot accuracy for GSM8K and MATH, pass@1 for HumanEval and MBPP, Matthews correlation coefficient for CoLA, accuracy for SST-2, QNLI, and RTE, matched accuracy for MNLI, accuracy and F1 score for MRPC and QQP, and Pearson and Spearman correlation for STS-B.
Implementation Details. Following Xu et al. (2023); Luo et al. (2023a; b), the inference of decoder-based LMs is implemented by vLLM Kwon et al. (2023). Temperature is set to 0.0 for greedy decoding. The maximal number of generated tokens is 1,024 on GSM8K, and 2,048 on the other four datasets. For encoder-based LMs, We fine-tune bert-base-uncased and roberta-base for 10 epochs with a warmup strategy. The weight decay is 0.01. We use 1e-5 and 5e-5 as learning rates and list the optimal setting of each fine-tuned model in Table 4 in Section A.2. Experiments are conducted on NVIDIA Tesla V100 and A100 GPUs.
2 Extreme Redundancy in SFT Delta Parameters
We show the extremely redundant property of SFT delta parameters of both decoder- and encoder-based LMs. We vary drop rate in [0.0, 0.1, 0.2, , 0.9, 0.99] and apply DARE to get models after removing the corresponding ratio of delta parameters. When is equal to 0.0, we actually obtain the standard SFT LMs. We report the performance of decoder-based LMs on GSM8K and HumanEval as well as encoder-based LMs on eight GLUE datasets in Figure 3 and Figure 4. Please see results of decoder-based LMs on AlpacaEval, MATH, and MBPP in Figure 12 in Section B.1.
We conclude that: (1) the SFT delta parameters of both encoder- and decoder-based LMs are highly redundant. DARE can effectively remove 90% delta parameters without significantly decreasing the performance. In some cases, the drop rate can even reach 99%; (2) the tolerance of drop rate increases with the sizes of LMs, i.e., LMs with more parameters can withstand higher drop rate. For example, WizardMath-70B performs well when while WizardMath-7B and WizardMath-13B fail. This depicts some connections with the scaling laws of LMs Kaplan et al. (2020); Hoffmann et al. (2022), indicating that there may exist quantifiable correlations between model sizes and drop rates they can afford.
3 Merging Models with DARE on SFT LMs
We combine DARE with five model merging methods, including Average Merging Wortsman et al. (2022), Task Arithmetic Ilharco et al. (2023), Fisher Merging Matena & Raffel (2022), RegMean Jin et al. (2023), and TIES-Merging Yadav et al. (2023). Please see Section A.3 for more descriptions. For feasible computations, we evaluate decoder-based LMs with Task Arithmetic by choosing the scaling term in [0.5, 1.0] and report the best results. We merge WizardLM-13B, WizardMath-13B, and llama-2-13b-code-alpaca since all of them adopt Llama-2-13b as the pre-trained backbone. WizardCoder-Python-13B is not selected as it is fine-tuned from CodeLlama-13b-Python. We merge encoder-based LMs with all five methods and perform grid search on some hyperparameters (see Table 5 in Section A.4 for more details). Following Jin et al. (2023); Yadav et al. (2023), we also fine-tune the models under the multi-task learning setting and report the oracle results. We show partial results of merging decoder-based LMs and encoder-based LMs in Table 1 and Figure 5. Please refer to Table 6 and Figure 13 in Section B.2 for more results.
In Table 1, DARE often facilitates Task Arithmetic on merging decoder-based LMs, yielding better results than the single model in many cases. For instance, merging WizardLM-13B and WizardMath-13B effectively improves WizardLM-13B’s zero-shot accuracy from 2.2 to 66.3 on GSM8K (better than WizardMath-13B’s 64.2 performance), and also enhances the instruction-following ability with win rate from 67.2 to 67.5 on AlpacaEval. Moreover, since llama-2-13b-code-alpaca is not well fine-tuned for generating codes (it performs worse than WizardLM-13B), we hypothesize this will affect the model merging performance. Hence, we additionally evaluate the code-generating ability of the merger of WizardLM-13B and WizardMath-13B, which obtains better results than llama-2-13b-code-alpaca, explaining the suboptimal performance of the amalgamation of WizardMath-13B and llama-2-13b-code-alpaca. Therefore, a potential prerequisite for effective model merging is that each model to be merged should be well fine-tuned.
From Figure 5, we observe that DARE often yields modestly better results of various merging methods, achieving an average improvement of 0.58%, 0.36%, 0.37%, -0.03%, and 0.84% on Average Merging, Task Arithmetic, Fisher Merging, RegMean, and TIES Merging. However, the merged model still struggles to surpass the single model in some cases, which is in line with the conclusion in Matena & Raffel (2022); Jin et al. (2023); Yadav et al. (2023).
We further provide two merged LMs with 7 billion parameters (supermario_v1 and supermario_v2) and evaluate them on Open LLM Leaderboard Beeching et al. (2023). Please see Section A.5 for more details. From Table 2, we find that the merged LMs beat the individual models they are built upon, achieving considerable improvements. Notably, until January 28th, 2024, supermario_v2 achieves the first rank on the Open LLM Leaderboard. It is exciting that these benefits are cheaply obtained by only using CPUs.
4 Importance of the Rescale Operation
As analyzed in Section 3.1, the rescale operation in DARE is essential to approximate the original embeddings. To verify this, we introduce DropOnly which randomly drops delta parameters without rescaling. We calculate the similarities of embeddings between the original LM and LM with DARE or DropOnly. Specifically, we obtain the embeddings of each input token layer-by-layer and report the average cosine similarities. Results of WizardMath-7B on GSM8K and bert-base-uncased on CoLA are shown in Figure 6. We observe that DARE can perfectly maintain the original embeddings in each layer with similarities higher than 0.95 even when removing 90% delta parameters. However, DropOnly just preserves the original embeddings with and the similarities sharply decline when is higher. For example, the similarities on WizardMath-7B decrease to about 0.85/0.68 when is 0.5/0.9). We further show the distributions of embeddings’ cosine similarities in the last layer in Figure 7, demonstrating the ability of DARE in approximating original embeddings. Note that similar findings can be obtained on other LMs and datasets but they are not presented due to page limits.
We also report the performance of LMs with DARE and DropOnly in Figure 8. See Figure 14 and Figure 15 in Section B.3 for additional results. We observe that discarding the rescale operation usually leads to worse results, and the performance gaps between DARE and DropOnly become more significant with the increase of . This validates the effectiveness of the rescale operation in DARE once again.
5 Comparison with Magnitude-based Pruning
We compare DARE with the commonly used Magnitude-based Pruning (MP) Han et al. (2015); Li et al. (2018); Lee et al. (2021), which chooses parameters based on their magnitudes. For more fair and credible comparisons, we adapt MP to operate on delta parameters and discard the retraining process. We show partial results of LMs with DARE and MP in Figure 9. Please refer to Figure 16 and Figure 17 in Section B.4 for extra results. We find that DARE outperforms MP in most cases and the superiority of DARE is more obvious when the drop rate becomes higher, verifying the superiority of DARE in abandoning delta parameters. We have also tried to combine MP with the rescale operation but got worse results than using MP separately. This is because MP removes parameters with smaller magnitudes and retains certain parameters with the largest magnitudes. Simply rescaling the remaining ones would destroy the original embeddings and result in unpredictable performance.
6 When Can DARE Be Used?
We investigate the prerequisites that DARE can work. We choose Llama-2-13b instead of CodeLlama-13b-Python as the pre-trained backbone for WizardCoder-Python-13B and apply DARE to derive the model after dropping certain delta parameters for evaluation. We find that the pass@1 metric on HumanEval/MBPP drastically decreases from 63.41/55.4 to 0.0/0.0 when only 10% delta parameters are removed. We deduce this is because Code Llama models are additionally trained with 500B tokens of code-related data Rozière et al. (2023), resulting in more obvious changes in parameter values with respect to Llama 2 models. Since WizardCoder-Python-13B is fine-tuned based on CodeLlama-13b-Python, when it uses Llama-2-13b as the pre-trained backbone, the ranges of SFT delta parameters would become much larger, making DARE infeasible. To verify this, we depict the absolute values of SFT delta parameters of 13B decoder-based LMs vs. various pre-trained backbones in Figure 10. Please see Figure 18, Figure 19 and Figure 20 in Section B.5 for the SFT delta parameter ranges on decoder- and encoder-based LMs. Note that the results are plotted by randomly choosing 10% delta parameters due to the huge parameter numbers of LMs. We also present the statistics about the deciles of delta parameter ranges of both decoder- and encoder-based LMs in Table 7 in Section B.5.
From the results, we observe the absolute values of delta parameters of WizardCoder-Python-13B vs. Llama-2-13b (often greater than 0.01) are several orders of magnitude bigger than those of WizardCoder-Python-13B vs. CodeLlama-13b-Python (usually within 0.0002), causing the failure of DARE. For other 13B decoder-based LMs fine-tuned from Llama-2-13b, most of their absolute values of delta parameters are less than 0.005, making DARE a proper choice. To this end, we conclude that DARE can work well when the absolute values of SFT delta parameters are relatively small (e.g., less than 0.005). Otherwise, DARE may fail.
7 Can DARE Drop Fine-tuned Parameters?
As previous network pruning methods mainly operate on the fine-tuned instead of delta parameters, we also conduct experiments under this setting. For decoder-based LMs, we find they perform badly when removing fine-tuned parameters even with 0.1 as the drop rate. Quantitatively, the performance sharply drops from 67.20 to 8.56 on AlpacaEval for WizardLM-13B, from 64.22/14.02 to 0.38/0.16 on GSM8K/MATH for WizardMath-13B, from 63.41/55.40 to 0.0/0.20 on HumanEval/MBPP for WizardCoder-Python-13B. Similar observations can also be found on MP or decoder-based LMs with 7B, 34B, or 70B sizes. Partial results on encoder-based LMs are shown in Figure 11 and please see Figure 21 in Section B.6 for additional results. We observe that directly eliminating the fine-tuned parameters by either DARE or MP would lead to worse performance on encoder-based LMs. This confirms that the knowledge is inherent in pre-trained LMs, and SFT is responsible for unlocking instead of introducing new capabilities.
Conclusion
In this work, we first discussed the extremely redundant properties of SFT delta parameters in LMs and proposed a simple approach DARE to effectively reduce the number of delta parameters needed for SFT without any data, retraining, or even GPUs. DARE can impressively drop 90% or even 99% SFT delta parameters without sacrificing much performance compared with using all SFT delta parameters. We further employed DARE as a versatile plug-and-play approach for existing model merging methods to merge multiple task-specific fine-tuned models into a single model with diverse abilities. Extensive experimental results on both encoder- and decoder-based LMs demonstrated the effectiveness of DARE in reducing SFT delta parameter redundancy and facilitating the model merging performance. We also provided a deeper analysis of why DARE works as well as the prerequisites for using DARE.
Impact Statements
Recently, merging language models has become a promising research direction. Our work allows researchers to obtain a single model with diverse capabilities at a low cost. Thanks to our method, hundreds of models with different functionalities and domains have been created on the Hugging Face community. Several popular toolkits for model merging have also been established on the GitHub platform. Even though this work has no direct social impacts, the potentially harmful information generated by LLMs (e.g., gender bias, racial discrimination) may still exist when using our approach. It is necessary to advocate for careful regulation by communities and governments on this matter.
References
Appendix A Detailed Experimental Settings
Table 3 shows the versions and correspondences with pre-trained backbones of SFT decoder-based LMs.
A.2 Learning Rate Configurations of Encoder-based LMs on GLUE
The optimal settings of the learning rate of each fine-tuned encoder-based LM are presented in Table 4.
A.3 Descriptions of Existing Model Merging Methods
We experiment with five model merging methods:
Average Merging simply averages the parameters of multiple models to get the merged model Wortsman et al. (2022).
Task Arithmetic uses a scaling term to control the contributions between the pre-trained backbone and the models to be merged Ilharco et al. (2023).
Fisher Merging first estimates the importance of parameters by calculating the Fisher information matrix, and then fuses parameters based on their importance Matena & Raffel (2022).
RegMean recasts the model merging task as a linear regression problem and derives closed-form solutions to solve the problem Jin et al. (2023).
TIES-Merging aims to address parameter conflicts in model merging. It first trims parameters with lower magnitudes, and then resolves sign disagreements. Parameters with consistent signs are finally merged Yadav et al. (2023).
A.4 Details of Grid Search on Hyperparameters of Model Merging Methods for Encoder-based LMs
Table 5 shows the searched ranges of model merging methods’ hyperparameters for encoder-based LMs. For DARE, we search the drop rate in [0.1, 0.2, , 0.9] and select the optimal setting with the best performance.
A.5 Details of Our Merged 7B LMs and the Open LLM Leaderboard
We offer two merged LMs with 7 billion parameters, namely supermario_v1 and supermario_v2. Specifically, we choose NeuralBeagle14-7Bhttps://huggingface.co/mlabonne/NeuralBeagle14-7B and Turdushttps://huggingface.co/udkai/Turdus to build supermario_v1, where both of them all derived from Beagle14-7Bhttps://huggingface.co/mlabonne/Beagle14-7B. We set the drop rate in DARE to 0.3, and merge NeuralBeagle14-7B and Turdus by Task Arithmetic with 0.8 as the scaling term. We select WildMarcoroni-Variant1-7Bhttps://huggingface.co/BarryFutureman/WildMarcoroni-Variant1-7B and WestSeverus-7B-DPO-v2https://huggingface.co/FelixChao/WestSeverus-7B-DPO-v2 to obtain supermario_v2, where both of them adopt Mistral-7B-v0.1https://huggingface.co/mistralai/Mistral-7B-v0.1 Jiang et al. (2023) as the backbone. The drop rate in DARE is set to 0.5, and the scaling term in Task Arithmetic is also 0.5.
The Open LLM Leaderboardhttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard is established to evaluate open-sourced LLMs based on Eleuther AI Language Model Evaluation Harness Gao et al. (2023), which contains six benchmarks including AI2 Reasoning Challenge (ARC) Clark et al. (2018), HellaSwagZellers et al. (2019), MMLUHendrycks et al. (2021a), TruthfulQALin et al. (2022), WinograndeSakaguchi et al. (2020), and GSM8KCobbe et al. (2021). The average score on the six datasets is used for ranking models on the leaderboard. We refer interested readers to the original papers for detailed information on the datasets.
Note that due to space limits, in Table 2, we use Hella., TQA and Wino. as abbreviations for HellaSwag, TruthfulQA, and Winogrande. WildMarcoroni-7B and WestSeverus-7B are the abbreviations for WildMarcoroni-Variant1-7B and WestSeverus-7B-DPO-v2.
Appendix B Additional Experimental Results
Figure 12 shows results of decoder-based LMs on AlpacaEval, MATH, and MBPP with different drop rates. We notice that the performance of WizardLM-70B drastically declines on AlpacaEval when the drop rate is 0.9 (different from the observations of WizardMath-70B and WizardCoder-Python-34B). One possible reason is that the instruction-following task on AlpacaEval is harder and requires general abilities with more delta parameters via SFT, causing more obvious dependencies among parameters (especially on LMs with larger sizes). Therefore, when the ratio of dropped delta parameters reaches a relatively small value (e.g., 0.9 in this case), the dependent relationships among parameters are destroyed, leading to unsatisfactory performance.
B.2 Additional Results of Merging LMs with DARE
Table 6 presents the complete performance of merging decoder-based LMs. Figure 13 shows the performance of merging encoder-based LMs on GLUE.
B.3 Additional Results of Comparisons between DARE and DropOnly
The comparison results between DARE and DropOnly on AlpacaEval, MATH, HumanEval, and MBPP on decoder-based LMs and all results on GLUE on encoder-based LMs are shown in Figure 14 and Figure 15, respectively.
B.4 Additional Results of Comparisons between DARE and MP
Comparisons between DARE and magnitude-based pruning on AlpacaEval, MATH, HumanEval, and MBPP on decoder-based LMs and all results on GLUE on encoder-based LMs are shown in Figure 16 and Figure 17, respectively.
B.5 Ranges of SFT Delta Parameters of Decoder-based LMs and Encoder-based LMs
We show the SFT delta parameter ranges of decoder- and encoder-based LMs in Figure 18, Figure 19 and Figure 20. We give the statistics about the deciles of delta parameter ranges in Table 7, which are computed on 0.1%/10% randomly chosen parameters of decoder-based/encoder-based LMs for feasibility.
B.6 Additional Results of Dropping Fine-tuned Parameters on Encoder-based LMs
Figure 21 shows the results of removing fine-tuned parameters on GLUE on encoder-based LMs.