Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, Dacheng Tao
Introduction
Model merging, also known as model fusion, is an effective technique that merges the parameters of multiple separate models with different capabilities to build a universal model without needing access to the original training data or expensive computation. The concept most relevant to model merging is ensemble learning , as both facilitate knowledge fusion and transfer. As shown in Figure 1, the main difference between them is that ensemble learning must save all individual models and fuse the predictions (or outputs) of multiple models during the inference phase, whereas model merging performs merging directly at the parameter level and only has one final model during inference. This gives model merging more attractive properties.
Although model merging is a relatively young topic, it is evolving rapidly and has already found applications in several domains. For example, in foundation models, models fine-tuned by different downstream tasks are merged to enhance the capabilities of large language models, and image generative models with different styles are merged to create a new model with mixed-style capabilities. In particular, the number of pre-trained and fine-tuned checkpoints in the machine learning community has grown exponentially in recent years, including open-source repositories such as Huggingface , torchvision , and timm , making it easy for users to obtain well-trained expert models of varying abilities. These rich model repositories further promote the rapid development of model merging direction.
As model merging becomes increasingly popular in various areas of the machine learning community, it is crucial to have a comprehensive understanding of the advantages and limitations of existing model merging techniques and their applications across different domains. Although some efforts have been made by the community , there are still large gaps to be filled. More specifically, MergeKit and FusionBench are technical reports in which only seven representative methods are discussed in MergeKit, and eight merging methods are discussed in FusionBench. Additionally, Zheng et al. discuss the topic of “learning from models” and it only mentions model merging as a subsection (single page only) in the whole paper. The most related work to the “model merging” topic is , but in terms of application, it only discusses model merging in three scenarios: federated learning, fine-tuning, and distillation. It also ignores a lot of recently published articles due to the rapid evolution of the model merging direction. To address these gaps, this survey aims to elucidate the methods, theories, applications, and future trends in model merging direction, providing a comprehensive classification of relevant approaches. In particular, this paper enhances the comprehensive understanding of model merging by covering three main aspects:
First, how are existing model merging methods classified? We first propose a new taxonomy in Figure 2 (upper part) that divides existing model merging methods into two phases (§2): pre-merging and during-merging. (i) Pre-merging methods aim to create better conditions for merging. It is further divided into using linearized fine-tuning to achieve weight space and input space disentanglement, performing architectural transformations to convert heterogeneous models into homogeneous models, and aligning weights to place them in the same basin. (ii) During-merging methods focus on designing sophisticated techniques to merge multiple models into one. These methods address task conflict and interference problems when merging models. They can be further divided into basic merging methods that perform the simplest parameter merging strategy; weighted merging methods that merge multiple models according to the importance calculated by specific rules; subspace merging methods that project multiple models into sparse subspaces for merging; routing-based methods that dynamically merge models according to input samples during inference; and the post-calibration based method that corrects the merged model. In addition to these methods, we also discuss the theoretical or empirical analysis of model merging.
Second, which applications can benefit from model merging? We discuss in detail the various use cases of model merging in foundation models (§3) and over ten subfields of machine learning (§4). As shown in Figure 2 (lower part), model merging can be applied to a variety of foundation models, including large language models, multimodal large language models, and image generative models. For example, model merging in large language models can help mitigate untruthfulness and toxicity output, accomplish knowledge unlearning, and speed up training. Moreover, model merging also arises in different machine learning subfields, such as continual learning, multi-task/multi-domain learning, few-shot learning, and other subfields, to solve a variety of challenges. For instance, in continual learning, model merging can mitigate catastrophic forgetting of old tasks. In multi-task learning, multi-objective learning and multi-domain learning, it facilitates knowledge transfer. Additionally, in adversarial learning, model merging can be employed for both attack and defense strategies.
Third, what are the remaining challenges and future research opportunities for model merging? Despite the advancements in merging methods and their well-developed applications, there are still numerous open challenges and future research directions in the field (§5). For example, as the number of tasks increases, the performance gap between existing methods and independent expert models becomes significantly larger. Additionally, current model merging methods incur enormous memory costs during merging and lack trust guarantees as well as in-depth theoretical analysis. Addressing these gaps will require substantial efforts from researchers to further advance the flourishing development of this field.
To summarize, the main contributions of this paper include the following three aspects:
Methodology Overview: We provide a comprehensive summary of the technical aspects of model merging. Specifically, we propose a new taxonomy that divides existing model merging methods into two stages and further subdivides the methods in each stage according to key techniques. Additionally, we discuss theoretical analysis work related to model merging.
Application Overview: We offer a comprehensive summary of the application aspects of model merging. Specifically, we explore the application of model merging to foundation models and machine learning subfields, demonstrating how model merging can address existing challenges in these areas.
Future Directions: We outline several remaining challenges and future directions for model merging. We believe that model merging needs to be further explored in the future from the perspectives of performance gap, theoretical analysis, trustworthy guarantees, cross-disciplinary applications, etc.
The main structure of this paper is as follows: §1 is an introduction, and §2 offers a comprehensive discussion of advanced model merging methods from a technical perspective. In §3 and §4, we summarize the applications of model merging in various foundation models and different subfields within machine learning, respectively. The remaining challenges and future research directions are discussed in §5. Finally, we conclude this paper in §6.
Advanced Model Merging Methods
In this section, we first introduce the notation and problem definition of model merging in §2.1. We then elaborate on advanced model merging methods (Table 1 summarizes the primary purpose of each category of methods). Existing model merging techniques can be roughly divided into the following two categories: (i) Before Merging Methods in §2.2: it provides better prior knowledge for model merging. (ii) During Merging Methods in §2.3: it resolves task conflict/interference by various strategies, and then performs parameter merging operations. Finally, we conclude with theories or explanations for the effectiveness of model merging in §2.4.
Assume there are models () of the same architecture that need to be merged, and they train from scratch or fine-tune on the same pre-trained model respectively. The parameters (or weights) of the -th model are represented as , where denotes the -th layer of the model, and is the total number of layers.
In this survey, we focus on parameter-wise merging. In other words, the goal of model merging is to merge the parameters , and finally obtain the new parameters . One straightforward solution for merging models is weighted averaging , defined as . However, the performance of this approach is often unacceptably poor or infeasible due to several possible factors: (i) The lack of suitable merging conditions, such as multiple models not being in the same basin or having inconsistent architectures. (ii) There are conflicts and interference among multiple models. We illustrate how advanced methods address these issues in §2.2 and §2.3, respectively.
2 Pre-Merging Methods
To provide better preconditions for model merging, one class of work focuses on the fine-tuning step of independent models, such as fine-tuning the linearized model instead of the nonlinear model (in §2.2.1). Additionally, when multiple model architectures that need to be merged are inconsistent, they must be pre-transformed to the same architecture (in §2.2.2). Finally, another class of work attempts to align the weights/parameters before merging (in §2.2.3).
Ortiz-Jimenez et al. reveal that one necessary condition for effective model merging is ‘weight disentanglement’. This means that different directions of the weight space correspond to functional changes in disjoint regions of the input space. For example, if model in weight space corresponds to a function change on in input space and model 2 in weight space corresponds to a function change on in input space, then the merged model will not interfere with each other in terms of function change. In other words, two models satisfying this weight disentanglement property can coexist in one model without affecting their respective performance, which is a very attractive property.
To achieve weight disentanglement, Ortiz-Jimenez et al. propose fine-tuning the linearized model along the tangent space of the pre-trained model during the fine-tuning stage, rather than in the original space of the nonlinear model. However, linearized fine-tuning with all parameters is more expensive than nonlinear fine-tuning. To accelerate this process, some works suggest linearizing only part of the layers. For example, Tang et al. propose partially linearizing the Adapter modules and then merging Adapters. Jin et al. suggest linearly fine-tuning only the linear layers in the attention modules of the full model. Furthermore, TAFT develops an efficient linearization method for the Transformer architectures, which directly derives closed-form linearized solutions for transformer networks. In summary, fine-tuning in the tangent space makes it easier to disentanglement the input space and weight space, thereby reducing the interference in subsequent model merging.
2.2 Architecture Transformation
In some cases, models that need to be merged may have different architectures and cannot be merged directly. To solve this problem, some studies propose to perform architecture transformation before merging, that is, transform multiple models with different architectures into the same architecture as shown in Figure 3 (a). For example, GAN Cocktail attempts to merge multiple GAN models with different architectures. It transforms all GAN models () into a specified target model , that is, is used as the initialization to learn the output of , while adding implicit regularizations to ensure that does not forget knowledge of the task . Consequently, the transformed GAN models have the same structure and shared knowledge, facilitating further model merging. Similarly, FuseChat proposes to merge chat LLMs with diverse architectures and scales (e.g., NH2-Mixtral-8x7B , NH2-Solar-10.7B , OpenChat-3.5-7B in their practical applications). Specifically, FuseChat first uses knowledge distillation to transform all the architectures to match that of OpenChat-3.5-7B, and then performs the model merge operation. Unlike the above distillation-based approach, CLAFusion adds layers/blocks (with weights set to the identity matrix) to the smaller model to align its architecture with that of the larger model. In summary, merging models with different architectures requires first transforming all models into a common architecture to merge later.
2.3 Weight Alignment
The linear mode connectivity (LMC) property of deep neural networks demonstrates that there is a connected path between multiple local minima of deep neural networks along which the loss remains nearly constant . Numerous studies have shown that two independent models, starting from the same pre-trained model and fine-tuned with different hyper-parameter configurations, typically satisfy LMC. Further, Adilova et al. and Zhou et al. extended the study of LMC to the layer level. The LMC property implies that multiple local minima may be equivalent in the weight space, and different weight configurations of the same model may represent the same functionality. Inspired by this, many works proposed to permute the weights of one model (i.e., ) to align with the other model when merging/interpolating two separate models, as illustrated in Figure 3 (b). denotes a permutation function, and researchers have dedicated efforts to studying effective and efficient permutation strategies for model alignment.
OTFusion and Imfeld et al. adopt optimal transport to soft-align neurons across models. NeuronAlignment introduces an inexpensive heuristic algorithm to approximate the optimal neuron alignment. CCAMerge permutes by maximizing the correlation between linear combinations of neurons. Notably, Git re-basin proposes three methods –activation matching, weight matching, and straight-through estimation– to align (or permute) the weights of models trained on different tasks. Based on the Git re-basin, Peña et al. further incorporate a Sinkhorn-based projection to improve these alignment methods. In addition, MuDSC proposes simultaneously performing model alignment in weight and activation spaces. Unlike heuristic alignment strategies, Deep-Align proposes a learning-based approach to weight alignment, employing a novel learnable architecture that takes two sets of weights as input and outputs a permutation matrix for alignment.
Despite the significant improvement of these alignment algorithms, Jordan et al. argue that the success of these methods depends on the use of normalization layers (e.g., BatchNorm, LayerNorm, etc.) in the model; without these, the performance of the matching algorithms is greatly reduced. The authors call this the “variance collapse” problem and propose the REPAIR method to solve it. Additionally, Crisostomi et al. noted that previous pairwise permutations do not guarantee cycle consistency, making the alignment fragile. They further proposed to globally optimize the permutations of all layers simultaneously at each step. Overall, aligned models experience much less interference or conflict during merging compared to directly merging unaligned models.
3 During Merging Methods
In this section, we provide a detailed discussion on how to merge a set of well-trained models. The existing methods can be roughly divided into five categories: basic merging methods (§2.3.1), weighted-based merging methods (§2.3.2), subspace-based merging methods (§2.3.3), routing-based merging methods (§2.3.4), and post-calibration based methods (§2.3.5).
One of the most straightforward approaches to model merging is to directly weighted average the parameters of multiple models , i.e., . However, the performance of simple weight averaging is generally unsatisfactory. Recently, Task Arithmetic introduced the concept of “task vector” (in Figure 4(a)), which represents the model parameter fine-tuned on task subtract the pre-trained model parameter , i.e., . In other words, task vectors are thought to steer the behavior of a neural network meaningfully. For example, multitask learning (MTL) can be accomplished by adding task vectors, forgetting can be achieved by subtracting task vectors, and task analogies can be performed using analogous task vectors. Specifically, when we want the pretrained model to perform MTL, we can add multiple task vectors to the pretrained model, i.e., in Figure 4(b), where is a hyperparameter. Conversely, when we want the pretrained model to forget a function , we can subtract the corresponding task vector from pretrained model as Figure 4(c), i.e., . As shown in Figure 4(d), we can also implement task analogies by task vector analogies, thus enabling zero-shot learning of new tasks. Similarly, PEMs combines Adapters with different capabilities by extending task arithmetic to parameter-efficient fine-tuning settings. However, the performance of basic merging methods is not satisfactory most of the time, especially when the tasks interfere with each other.
3.2 Weighted-based Merging Methods
As we all know, different models (or task vectors) represent different functions, and intuitively, different functions have varying degrees of importance. Therefore, advanced weighted-based model merging methods design various clever rules to determine the merging coefficients, as shown in Figure 5(a). For instance, when merging two models and (or task vectors and ), the goal of the weighted merging method is to find the optimal coefficients and so that the merged model (or ) can retain the capabilities of the independent models as much as possible. However, when the number of models is large, it is impractical to use brute-force grid search to find the optimal merging coefficient because of the expensive search cost involved.
To determine the merging coefficient more effectively, Evolutionary-model-merge and Checkpoint Merging efficient searches for the merging coefficients using evolutionary algorithms and Bayesian optimization, respectively. AdaMerging uses gradient descent optimization to learn the merging coefficients by minimizing entropy as a surrogate loss in unlabeled test data. MetaGPT casts the model merging problem as an MTL formalism, where the goal is to minimize the average loss of the merged model and the independent model. It employs local linearization of the model and the orthogonality of the task vectors to derive the optimal merging coefficient for each model as follows: . SLERP performs spherical interpolation of the parameters of the two models. The interpolated coefficients of and are given by and , respectively, where denotes the angle between the two task vectors, and represents the merging coefficient of the initial setting.
The above-sophisticated weighting methods operate at the model (or task) level. It is well known that each layer and even each neuron in a deep neural network model play a significantly different role, and some research has developed more fine-grained weighted merging strategies. For example, Layer-wise AdaMerging and aTLAS adaptively learn different sets of merging coefficients for each layer or module of the model, respectively. RegMean indicates that closed-form solutions (relying on the data statistics provided by the training set) exist for linear layers in model merging, while nonlinear layers can simply perform weight averaging. Other works utilize the Fisher information matrix to assess the importance of parameters when merging. Fisher-Merging performs model merging based on the importance of the parameters in each independent model, that is, , where is the diagonal of the Fisher information matrix with respect to task . Fisher-nodes-merging also combines a set of Transformer models based on the Fisher information matrix. MaTS developed a block diagonal approximation for Fisher merging. Daheim et al. linked the inaccuracy of weighted average with gradient mismatch, and further proposed an uncertainty-based algorithm to reduce the matching error, ultimately merging the models based on a second-order Hessian estimation.
3.3 Subspace-based Merging Methods
Another class of advanced methods transforms models into sparse subspaces for merging, thereby mitigating task interference. The over-parameterized nature of neural networks and the success of model pruning show that removing most of the parameters from the model barely affects its accuracy . This insight opens up new opportunities for model merging, allowing us to remove insignificant neurons from a single model and merge multiple sparse models within the parameter subspace, as shown in Figure 5 (b).
TIES-Merging proposes to trim each individual model based on parameter magnitudes, retaining only the top 20% of parameters with the highest magnitudes. It further suggests eliminating parameter sign conflicts to reduce interference, and finally merging sparse models using Task Arithmetic . Similarly, Drop And REscale (DARE) also sparsifies by parameter magnitude, and highlights the importance of further performing rescaling on sparse models. In addition to removing the tail parameters with the smallest weight, the Model Breadcrumbs highlight the importance of removing the parameters (outliers) with the largest weights to further reduce noise in model merging and enhance generalization to hyperparameters. TALL-masks creates a mask matrix specific to each task based on a predefined threshold related to independent models. Further, Model Tailor masks unimportant parameters based on the sensitivity of fine-tuned parameters to loss changes and the significance of changes compared to pre-trained parameters. Unlike the standard practice of obtaining a single model through model merging, EMR-Merging proposes maintaining a shared model among multiple tasks alongside a sparse task-specific model. In this approach, the value of the shared model at each index is the largest parameter value among all models. In contrast to the mask construction rules of the aforementioned heuristics, Concrete frames mask construction and model merging as a learnable bi-level optimization problem. The outer-level optimizes the mask matrix, while the inner-level merges the model based on the mask matrix and optimizes it using the unlabeled test samples.
3.4 Routing-based Merging Methods
The basic, weighted-based, or subspace-based merging methods discussed in §2.3.1, §2.3.2 and §2.3.3 are static merging methods. This means that the merged model remains the same for all samples or tasks. Given that there are differences between input samples/tasks, the model’s ability may vary when processing different samples/tasks. As shown in Figure 5 (c), some works propose to dynamically merge models (or subsets of layers) based on the samples/tasks during the inference phase.
For a given input, SMEAR first computes a weighted average of the parameters of each expert by using the distribution of router inputs to the expert modules. The advantage of this approach is that it has a similar computational cost to that of a single expert. Twin-Merging also adaptively combines task-shared and task-private knowledge based on routing during the inference phase. Similarly, Weight-Ensembling MoE proposes a dynamic merging Transformer architecture. Specifically, they observed that the parameters of the linear layer in the fine-tuned model changed more dramatically than those of the nonlinear layer, which also significantly impacted the merging performance. Therefore, they use a standard weighted average to merge all modules except the linear layer. The linear layer is dynamically weighted and merged according to the routing network (sample features as the input of the routing, and merging coefficients as the output) during inference. PWE MoE further extends Weight-Ensembling MoE to a multi-objective optimization setting and uses the preference vector as input for routing.
3.5 Post-calibration based Methods
Recently, Yang et al. introduce a post-merging method to calibrate merged models. They observed that merged models (across multiple mainstream model merging methods) suffer from representation bias, meaning the representations extracted by the independent and merged models are very different, leading to performance degradation in the merged model. To alleviate this problem, they propose a module called ‘representation surgery’ to calibrate the representation bias. The core idea is to align the representation of the merged model after representation surgery with that of the independent model.
4 Theories and Analysis of Model Merging
In addition to designing various advanced methods in §2.2 and §2.3, the theoretical and effectiveness analysis of model merging is also crucial. Currently, there is limited work on the theoretical analysis of model merging. Based on the source of the models to be merged, the existing theoretical analysis can be roughly divided into three categories: (i) model merging of different checkpoints in the same training trajectory, (ii) model merging of different models fine-tuned on the same dataset, and (iii) model merging of different models fine-tuned on different datasets or tasks.
First, some analyses target model merging on the single-trajectory training, usually referring to stochastic weighted average (SWA) or exponential moving average (EMA). For example, Jain et al. theoretically proved that the excess risk of the EMA is an upper bound of a bias term and a variance term in the context of least squares regression. The bias term depends on the initialization state of the parameters and decreases exponentially with the number of iterations once the model starts averaging. The variance term depends on the noise covariance inherent in the data, which decays at a faster rate when model averaging is used . Similarly, Rame et al. applies bias-variance decomposition to the domain generalization setting to explain why model averaging improves out-of-distribution performance. In addition, Hardt et al. provide a stability bound for SWA under convex assumptions, while Wang et al. further establish generalization bounds analysis in both convex and nonconvex cases.
Second, some studies explain the merging of multiple models with different hyperparameter fine-tuning for the same dataset in terms of connectivity and flatness of the loss landscape. Specifically, some works apply the theory of linear mode connectivity (LMC) of neural networks to explain model merging. LMC reveals that neural network loss minima are not isolated points in the weight space. Recent studies have shown that two independent models, starting from the same pre-trained model and fine-tuned with different configurations, usually satisfy LMC. In other words, LMC is a general phenomenon that typically appears in fine-tuned models based on the ”pretraining-finetuning” paradigm, which is the current standard in the machine learning community. Therefore, performing weight alignment according to LMC provides a robust validity guarantee for model merging . On the other hand, other studies explain model merging from the perspective of a flatter loss landscape , arguing that merging multiple weights fine-tuned under different optimization configurations with the same data usually converges to a flat local minimum , thus revealing why model merging has better generalization .
Finally, an analysis by Ortiz-Jimenez et al. is based on multiple models fine-tuned on different datasets, identifying weight disentanglement as a necessary precondition for effective model merging. More specifically, Ortiz-Jimenez et al. provide theoretical and empirical analyses of the neural tangent kernel (NTK) and establish a compelling link between the task arithmetic and the spectral properties of NTK.
Application of Model Merging in Foundation Models
The emergence of foundation models, including large language models (LLMs), multimodal large language models (MLLMs), and image generative models, is a significant indicator of technological progress in the field of artificial intelligence in recent years. However, despite their advancements, these large models still face several challenges, such as generating harmful content in LLMs, MLLMs struggling with fusing information from different modalities, and the difficulty of producing mixed-style images in image generation models. Recent studies suggest that model merging techniques offer a promising solution to these inherent challenges in foundational models. Table 2 first briefly summarizes the application of model merging in foundational models. Then, §3.1, §3.2 and §3.3 provide a detailed discussion on how LLMs, MLLMs, and image generative models benefit from model merging, respectively.
In recent years, large language models (LLMs), such as GPT-4 , Gemini , PaLM and LLaMA , have made significant advancements and have been widely applied across various tasks. Despite their superhuman performance on most basic tasks, LLMs still face numerous challenges, including producing toxic content that violates laws or ethics, using unauthorized data during training, high training costs, and insufficient performance in specific domains. Model merging technology presents a promising opportunity to address these challenges.
Humans often hold diverse opinions about aesthetics, politics, or fairness. When LLMs serve humans, different people have different expectations of the model, e.g., some expect LLMs to generate harmless responses, while others seek engaging and enjoyable interactions . Consequently, the development of practical LLMs is generally divided into three stages, to generate responses that are more helpful, accurate, and safer : Pre-training on a large amount of unsupervised data, supervised fine-tuning (SFT) on a small dataset with high-quality annotation, and interaction with humans to further optimize LLM alignment (e.g., direct preference optimization (DPO) or reinforcement learning from human feedback (RLHF) ) with human preferences, rewards, or values.
Some works propose to achieve better, safer, or faster alignment of human preferences by model merging. For example, ExPO adds a task vector, constructed by a moderate model aligned using DPO or RLHF on a small amount of human preference data, to an unaligned SFT model. A more powerful aligned model can be directly obtained by setting a suitable merging coefficient. On the AlpacaEval 2.0 benchmark , fusing a model aligned on the 10%/20% preference data with an SFT model results in performance comparable to that of a model aligned on the full preference data. DogeRM proposed merging the reward model with LLMs fine-tuned on different downstream domains to create domain-private reward models directly. Additionally, Lu et al. propose an Online Merging Optimizer, that interpolates the gradient with the SFT model at each step of RLHF. This approach encourages RLHF to optimize toward reward maximization while preventing LLMs from forgetting general knowledge due to RLHF. Beyond preference alignment, several studies have examined the impact of model merging for secure alignment of LLMs . For example, Hammoud et al. find that merging two security-aligned models could compromise security. Thus, they proposed explicitly including secure alignment as an optimization objective when constructing synthetic data for model merging.
In practice, users often have various combinations of preferences rather than a single preference. Training a model separately for each combination of preferences is unrealistic due to the infinite combinations and the high training costs. Therefore, some studies suggest combining models with different reward alignments to create a series of integrated aligned LLMs. For example, Rame et al. and Jang et al. propose Reward Soups and Personalized Soups, respectively, as efficient and flexible solutions for diverse rewards. Specifically, Rewarded Soups first trains an expert model for each reward and then linearly interpolates the weights of the experts to approximate the set of Pareto optimal solutions for various reward combinations. This approach is cost-effective, as it only requires training separate models for each reward to combine any variety of rewards.
1.2 Detoxifcation of LLMs
LLMs have been widely noted for issues related to untruthfulness and toxicity in various applications , such as insults, threats, and profanity in responses to certain questions. To address the potential security risks in the application of LLMs, flexible techniques are needed to reduce the generation of toxic text, essentially detoxifying LLMs. A straightforward solution is to collect additional non-toxic data to fine-tune LLMs ; however, this approach requires significant computing resources and may interfere with the general capabilities of LLMs. Alternatively, directly reducing the probability of potentially toxic words during the decoding stage requires additional guidance information . Recent studies have shown that reducing the toxic data generation of LLMs through model merging is a simple and effective scheme .
Task Arithmetic negates the task vectors of GPT-2 model fine-tuned on toxic data (Civil Comments ) and shows that this operation effectively reduces the proportion of data classified as ”toxic”, with little change in the fluency of the language on the control task (WikiText-103). Additionally, some parameter-efficient models steer the toxic behavior of LLMs by manipulating a small number of parameters. PEM negates LoRA (and (IA)3 ) modules trained on poisoning data to maintain language proficiency while reducing toxicity of language model output. Ethos and Ext-Sub point out that while the task vector on toxic data is factually wrong, it also contains correct information about language modeling and logical narrative skills. Therefore, Ext-Sub decomposes the toxic task vector into two orthogonal subspaces that represent general capability and destructive capability, respectively. Toxic knowledge is then eliminated by removing only the component representing the destructive ability from the LLM.
1.3 Knowledge Unlearning of LLMs
LLMs may inadvertently learn copyrighted material, raising significant legal and ethical concerns , and broader questions about responsible AI use . In this context, the California Consumer Privacy Act and the General Data Protection Regulations of the European Union stipulate the right to data forgetting. The foundational model’s knowledge must be adapted to comply with these regulations. However, the cost of excluding copyrighted data for re-training from scratch is prohibitive. For instance, training a Llama-2-70B from scratch requires 1,720,320 GPU hours . Traditional methods often use gradient ascent (GA) to achieve forgetting by fine-tuning the model using the GA algorithm on the specific data to be forgotten . Unfortunately, this approach typically catastrophically destroys other parts of the model’s knowledge. That is, forgetting specific knowledge also erases other knowledge that should be retained. Recently, many studies based on model merging techniques have demonstrated the potential to forget LLM-specific knowledge without harming other knowledge .
Unlike the GA-based approach, the model merging approach does not require additional data for other tasks to maintain old knowledge. To achieve forgetting, model merging typically incorporates a negatively fine-tuned model into the target model (i.e., the task-specific fine-tuned knowledge is subtracted from the target model). For example, Task Arithmetic shows that negating task vectors degrade performance on specific tasks without substantial changes to the control tasks. Experiments demonstrate that model merging can forget the knowledge of the target task in a fine-tuned model without harming performance on control tasks. Similarly, Stable Sequential Unlearning (SSU) extends this forgetting to the setting of sequential unlearning on LLMs, where different copyrighted content must be unlearned at different time steps. Knowledge forgetting can also forget samples that represent bad behavior during pretraining. For instance, FuseToForget employs model merging as a debiasing tool to reduce privacy issues in language models. FLearning first subtracts the parameters related to the data to be forgotten and then fine-tunes the parameters with new data to achieve accurate knowledge updates. SKU explores the forgetting of harmful data in LLM, which is a two-stage scheme. Initially, harmful data (e.g., harmful question-answer pairs) is used to fine-tune the parameters corresponding to the location of harmful knowledge in the LLM (i.e., the task vector), and then the task vector is negated from the LLM to mitigate undesirable behavior in the LLM effectively. Generally, incorporating the opposite (anti-expert) task vectors into the pre-trained model can effectively accomplish the task of machine unlearning.
1.4 Faster Training of LLMs
Training LLMs requires numerous iterations on massive data, making the training process extremely expensive. For example, training LLAMA2-70B with 2T tokens required 1,720,320 GPU hours . Methods to accelerate LLM training include mixed-precision training, continual retraining, and pipeline parallelism. An orthogonal approach is checkpoint merging in training trajectories, which offers a simple and effective means to either speed up LLM training or enhance training performance at the same cost.
The first type of works incorporate checkpoints in a single training trajectory during LLM training to accelerate model training. For instance, LAWA demonstrated that merging checkpoints during intermediate stages of model training speeds up the process. For example, training a ResNet50 model on the ImageNet dataset reduced the training time by 68 GPU hours, and training a RoBERTa-Base model on the WikiText-103 dataset saved 30 GPU hours. Sanyal et al. further showed that the combination of checkpoint averaging in pre-trained trajectories and a high learning rate contributes to faster convergence. Checkpoint Merging comprehensively evaluates the effectiveness of model merging at different stages of the Baichuan2 LLM model pre-training process. The second type of work involves combining existing models to create a more powerful initial model, thereby accelerating learning speed and improving accuracy on downstream tasks. For example, Fusing and ColD Fusion mixture of multiple existing fine-tuned models as base models and used for downstream task fine-tuning shows that this merged model outperforms the naive pre-trained model.
1.5 Combine the Capabilities of Expert LLMs
LLMs exhibit strong generalizability in general tasks, but often lack knowledge in specific vertical domains. Pretrained LLMs typically require fine-tuning within different corporations to become expert LLMs in various fields. Integrating the expertise of multiple specialists is particularly critical for solving more complex tasks. Research on model merging techniques indicates that a composite LLM can be created by combining the parameters of different expert LLMs . For example, Dekoninck et al. demonstrate the ability to flexibly control text generation by merging multiple LLMs with different styles and applying personalized weighting. Robust Weight Signatures proposes a robustness “patching” framework via model merging to enhance the overall robustness of the model against various naturally corrupted versions of clean data. In summary, model merging offers a straightforward and effective strategy for enhancing LLM’s capabilities.
2 Model Merging in Multimodal Large Language Models (MLLMs)
Foundation models often involve processing and interacting with data from different modalities, such as video, images, speech, and text. In order to build a generally large model, a key obstacle is the diversity and heterogeneity of tasks and modalities. Traditionally, most existing approaches train a modality-specific model for each modality. However, these methods face limitations: on the one hand, they require separate models for each modality; on the other hand, jointly training a large multimodal model necessitates the expensive collection of paired training data (image, text, video, speech) and the retraining of the entire model when a new modality is added.
An interesting question is whether we can merge multiple modality-specific models to obtain a single, effective, and parameter-efficient modality-agnostic model. We aim for the merged unified model to encode inputs from different modalities, learn cross-modal interactions, and maintain performance comparable to that of well-trained independent modality-specific models. Compared to traditional multimodal learning, model merging techniques offer new opportunities. This model-merging approach offers several benefits: (1) it eliminates the costly and labor-intensive process of collecting labeled paired multimodal training examples, which is required for jointly training multimodal models; (2) it enhances the adaptability of multimodal models, allowing for the seamless integration of new modalities; and (3) it fully leverages knowledge collaboration across multiple modalities, thereby benefiting from cross-modal knowledge transfer.
Recently, many studies have focused on merging models from different modalities into a single model, thereby enhancing the diversity of knowledge across modalities. For instance, JAM proposes to merge two specialized (one for text-to-image and one text-only) autoregressive, decoder-only, large transformer models to seamlessly generate multimodal outputs. Similarly, DAMC introduces a method for fusing multimodal LLMs across image, audio, video, and point cloud modalities, further reducing cross-modal interference through parameter decoupling and adjusting modality fusion coefficients.
To evaluate the impact of various factors on model merging, VL-Merging performs a comprehensive empirical analysis of multimodal model merging. The overall framework consists of three steps: independent modality fine-tuning, multimodal merging, and downstream task fine-tuning. Through experiments involving different initializations, merging methods, and architectures in multimodal model merging, the authors propose the following guidelines: (1) Models across multiple modalities should be based on the same pretraining starting point to ensure they are in the same basin and share more information. (2) Simple model averaging achieves better performance, and if more computation and storage resources are available, more fine-grained merges can be conducted. (3) Merging the entire model rather than just a subset of layers generally yields more satisfactory results, as fine-tuning only a subset of layers may restrict the capabilities of single-modality models. Unlike the above model merging approaches that are developed based on specific architectures, UnIVAL is the first to design a unified architecture for four modalities: image, video, audio, and language. It transforms the tasks across all modalities into a “sequence-to-sequence” format, with the training objectives of all modalities converted into a “next token prediction” format. This allows for a uniform feature extraction and classifier to be applied across all modalities. Additionally, UnIVAL provides favorable architectural conditions for model merging and demonstrates that linear interpolation of models fine-tuned across multiple modalities in weight space results in a general single model that performs well on both seen and unseen tasks.
2.2 Model Merging for Cross-Modal Knowledge Transfer
Some works attempt to transfer knowledge from one modality to another through a model merging approach. For instance, MAM investigates whether the attention layers of Transformers generalize across different modalities. Specifically, it examines if the knowledge acquired by Transformer models trained on high-resource modalities (e.g., data-rich images and text) can be transferred to Transformer models trained on low-resource modalities (e.g., data-sparse speech and audio). This paper demonstrates attention merging for models across various tasks, modalities, and initializations. The final results show that MAM achieves an 18.42% reduction in classification error on the audio classification task (using the ESC-50 dataset ) compared to the standard fine-tuning paradigm.
3 Model Merging in Image Generative Models
The goal of image generative models, such as generative adversarial networks (GANs), variational autoencoders (VAEs), normalizing flows (Flows), and denoising diffusion probabilistic models (Diffusions), is to approximate the underlying data distribution behind a given dataset so as to generate more new samples with the same distribution. However, image generative models still face the following challenges: the inability to flexibly generate samples with multiple style combinations, the high cost of generative model training, and the inability to generate all the details specified in the instructions. This dilemma has led to an interest in expert models, which train a set of experts with specific abilities on different data shards or distributions, allowing for the flexible addition or removal of certain styles of experts at inference time. Considering the difficulty in deploying and the cost of resources for ensemble learning, model merging offers a new perspective on combining skill-specific experts of different styles without additional memory and inference costs.
Existing generative models typically generate distributions based only on the training data. However, in real deployments, different users or artists often want to generate artwork with different combinations of styles. Collecting additional data for these mixed distributions is expensive, and fine-tuning the model can result in the forgetting of other capabilities. Model merging offers the potential to flexibly combine multiple styles.
Earl GAN Cocktail attempted to merge several pre-trained GAN models. Recently, diffusion-based image generative models have gained more attention than GAN-based models due to their superior generative capabilities. Consequently, most research focuses on fusing different diffusion models. Specifically, Diffusion Soup demonstrates the ability to linearly merge diffusion models fine-tuned on data shards of different styles (e.g., data provided by different domains/categories or different users), resulting in hybrid style zero-shot generation. In addition, Diffusion Soup empirically verifies that model merging has an anti-memorization effect, meaning the generated images are less likely to replicate training data, which is beneficial for generating diverse images. Unlike Diffusion Soup, which directly merges model parameters, MaxFusion is inspired by ZipIt and proposes merging intermediate features of multiple diffusion models based on the same input noise to generate images that satisfy multiple conditions. However, merging multiple diffusion models based on full parameter fine-tuning can be costly when the number of tasks is large. To address this issue, ZipLoRA and MoLE aim to seamlessly merge parameter-efficient LoRA modules. For example, ZipLoRA proposes merging independently trained content/subject (e.g., a specific object or person) LoRAs with artistic style (e.g., drawing or painting, etc.) LoRAs, allowing the diffusion model to generate any user-provided combination of subjects and styles . This approach enables users and artists to easily combine publicly available subjects and styles LoRAs of their choice.
3.2 Reducing Training Cost of Generative Models
In real-world scenarios, large-scale training data typically originates from different domains or is provided by various users. Given the need to add new data or remove outdated data, retraining a single model with updated data every time is often impractical . For instance, training a CM model using 8 A100 GPUs takes about one week. This is because existing methods only apply the final convergence weights during generative model training and ignore the intermediate training trajectories. LCSC demonstrates that a simple combination of training trajectories in the middle of the diffusion model by the evolutionary algorithm can significantly reduce the training cost. Specifically, only a few iterations or a small batch size is required to train the diffusion model to achieve image quality comparable to that of a fully trained diffusion model. For example, on the CIFAR-10 dataset, LCSC improves the training process for consistency distillation and consistency training by factors of and , respectively. The underlying reason is that each local checkpoint of the optimized trajectory has many high-quality basins (i.e., areas of better generation quality) nearby that cannot be reached by stochastic gradient descent due to substantial variance in gradient estimations. However, checkpoint interpolation provides an opportunity to reach these basins.
3.3 Enhancing the Faithfulness of Generative Models
Some studies on Text-to-Image (T2I) show that although the existing T2I generation model can generate high-quality images according to text prompts, the images often fail to fully capture and reflect the semantic details in the text, such as generating multiple subjects or correctly depicting the spatial relationships between objects . To enhance the fidelity of the T2I generation models, SELMA designed a novel four-stage paradigm. In the first and second stages, a series of input texts (corresponding to different skills) are collected through diversified prompts from existing LLMs, and the corresponding image data are generated using the T2I model. The third stage involves fine-tuning skill-specific experts (i.e., LoRA) separately on images of different skills. In the fourth stage, expert models with different skills are merged to obtain the final model during inference. Compared to the paradigm of joint learning of multiple skills, this method of merging expert skills after independent learning may help alleviate knowledge/skill conflicts while also being more efficient.
Application of Model Merging in Different Machine Learning Subfields
Model merging is a simple and effective technique widely used in various subfields of machine learning, such as continual learning, multi-task learning, domain generalization, federated learning, few-shot learning, and adversarial defense, etc. In this section, we comprehensively discuss the application of model merging in the different machine learning subfields. Table 3 provides a brief summary, and in §4.1 to §4.6, we introduce each application case in detail.
Continual Learning (CL) involves training a model using a streaming, non-stationary data stream. The primary challenge in CL is the ‘catastrophic forgetting’ problem; that is, the CL model’s prediction accuracy for old tasks drops dramatically after training on new tasks. The mainstream CL methods are mainly divided into memory replay-based methods, architecture expansion-based methods, regularization-based methods, and subspace projection-based methods . Recently, there has been a growing interest in using model merging to address the catastrophic forgetting problem. This novel approach offers several benefits, such as avoiding additional parameters and inference costs associated with network expansion-based methods and eliminating the need to cache old data as required by memory-based methods.
Tangent Model Composition proposes fine-tuning each task independently in the tangent space of the pre-trained model and then linearly fine-tuning these models to perform CL. This approach does not depend on the specific settings of CL and can be easily applied to task, class, and domain-incremental learning scenarios. In addition, ITA emphasizes the necessity for the fine-tuned model to be in the same basin as the pre-trained model to ensure the composability of nonlinear models. It introduces a regularization term similar to EWC in traditional CL to constrain the distance between the fine-tuned weights and the pre-trained weights when training the independent model. WARP suggests linearly interpolating the pre-trained LLM’s weights with its aligned weights via RLHF on a preference dataset, thus mitigating the forgetting of knowledge from the pre-trained LLM. BAM continuously adapts LLMs to new languages by merging models while preserving general capabilities. Model Tailor explores the problem of catastrophic forgetting during fine-tuning of MLLMs, and proposes to merge only the most important subset of parameters in the fine-tuned MLLM model into the pre-trained MLLM model, so as to retain the generalization ability of the pre-trained model as much as possible, while compensating the selected weights to reduce the performance of the fine-tuning task. MagMax merges pruned task vectors to further alleviate parameter sign conflicts and old knowledge forgetting. Equifinality, PAINT and LM-Cocktail interpolate the weights of the fine-tuned model and the zero-shot model to improve accuracy on downstream tasks without degrading accuracy on supported/general tasks.
In contrast to merging full models, some research focuses on merging parameter-efficient modules. Chitale et al. propose a CL method based on task arithmetic . This method first fine-tunes a task-specific LoRA for each task, then constructs a task vector based on the difference between fine-tuned and pre-trained models. Multiple task vectors are then merged, and a small amount of data (10 samples per class) is used to fine-tune the merged model. Compared to traditional CL methods, particularly those based on replay, this approach eliminates the need to replay data from old tasks at each iteration, thereby accelerating model training. Additionally, fine-tuning the merged model with a class-balanced subset helps mitigate CL model bias. Similarly, DynaMMo applies lightweight model merging (i.e., Adapter) in a CL setting for medical images. In contrast to architecture expansion-based CL methods, this approach does not result in a linear increase in the number of parameters with the number of tasks. Unlike the static aggregated parameter-efficient fine-tuning (PEFT) modules of DynaMMo, DAM introduces dynamic aggregated PEFT modules during inference to perform CL. AMM proposes merging convolutional layers to facilitate incremental new class discovery and prevent forgetting fundamental knowledge. Disperse-Then-Merge suggests merging submodels trained on different data partitions during the supervised fine-tuning of LLMs to reduce data bias and mitigate the forgetting of generic pre-trained knowledge.
2 Model Merging in Multi-Task/Multi-Objective/Multi-Domain/Auxiliary Learning
In machine learning, to optimize resource efficiency, we typically use a single model to handle multiple tasks, objectives, or data domains with varying distributions. The traditional multi-task learning (MTL), multi-objective learning (MOO), or multi-domain learning (MDL) paradigm requires gathering data from all tasks, objectives, or domains to collaboratively train a model, leading to high data management and model training costs. This approach becomes particularly costly when new tasks, goals, or domains are introduced, as retraining a comprehensive model from scratch using all available data is resource-intensive. Numerous recent studies have proposed efficient methods for integrating knowledge across tasks, goals, or domains by merging models directly.
The goal of multi-task learning (MTL) is to enable a single model to perform multiple tasks simultaneously, thereby facilitating knowledge transfer between these tasks . As shown in Figure 1(c), to avoid the high cost of joint training, a straightforward approach is to merge multiple independently trained models on different tasks to accomplish MTL. Almost all of the model merging methods discussed in §2.3 can be used to merge multiple models trained on different tasks to perform MTL. In this section, we take some representative tasks as examples. For MTL tasks in computer vision, Task Arithmetic , Ties-Merging , AdaMerging and other studies proposed to combine ViT models trained on different visual classification tasks, and the obtained model can complete the object classification of multiple tasks. The results of Task Arithmetic demonstrate that merging independently trained models from any two datasets yields a merged model whose performance is comparable to that of a single-task model. Similarly, ZipIt , which merges ResNet architectures trained on different tasks, achieves comparable results. For MTL tasks in natural language processing, DARE introduces a method to assimilate homologous models, augmenting LLMs as a ”free lunch”. For instance, merging WizardLM with WizardMath significantly boosts WizardLM’s performance on GSM8K (a benchmark for evaluating the mathematical reasoning ability of LLMs) from 2.2 to 66.3. Akiba et al. suggest that directly merging an LLM with mathematical capabilities and an LLM with Japanese language proficiency results in a model capable of solving Japanese mathematical problems. Furthermore, numerous studies have demonstrated that combining PEFT modules (such as Adapter or LoRA) trained on different tasks can also achieve MTL .
2.2 Knowledge Transfer in Multi-Objective Optimization
Multi-objective optimization (MOO) aims to optimize multiple objective functions simultaneously. These objective functions may conflict with one another, so the MOO problem typically does not have a single optimal solution. Instead, it involves finding a trade-off among the multiple objectives, which corresponds to identifying a set of Pareto optimal solutions. Tang et al. propose approximating the entire Pareto set using a mixture of experts (MoE) based model merging approach. Specifically, their method trains an independent model for each objective and learns a routing network to balance the trade-offs between the multiple objectives (models). The input of the routing network is the task preference vector, and its output consists of the merging coefficients for the independent models. Considering that directly evaluating Pareto solutions based on the original evaluation metric is time-consuming, MAP proposes a second-order Taylor expansion model as a surrogate model for the true evaluation metric, and further uses an evolutionary algorithm to calculate the Pareto front based on the surrogate model.
2.3 Knowledge Transfer in Multi-Domain Learning
Unlike existing model-merging-based MTL approaches that focus on datasets with different object categories, Ye et al. explore model merging across multiple domains, where datasets share the same categories but differ in environmental contexts. To mitigate conflicts between multi-domain models, this paper introduces a weight similarity criterion to assess the correlation between different model layers. For layers with high correlation, a simple weight averaging or RegMean strategy is employed to merge models that have been fine-tuned in different domains of the same task. For layers with low correlation, the weights are flexibly combined using a gating mechanism during the inference phase. Branch-Train-Merge demonstrates the effectiveness of training expert language models on 64 different domains and subsequently merging them.
2.4 Knowledge Transfer in Auxiliary Task Learning
The goal of auxiliary task learning (ATL) is to enhance the performance of the target task by leveraging knowledge obtained from related auxiliary tasks. Unlike MTL, which aims to optimize the average performance across all tasks, ATL focuses solely on improving the performance of the main task. However, ATL often encounters the issue of gradient conflict, leading to negative transfer, where the inclusion of auxiliary tasks interferes with the main task’s performance. To mitigate negative transfer, Jiang et al. propose ForkMerge, a method that periodically performs ‘fork’ and ‘merge’ operations. The model is first periodically duplicated into multiple branches: the first branch is trained exclusively on the main task, while the remaining branches are trained jointly on both the main and auxiliary tasks. An optimal merging coefficient is then determined using the validation set to merge the models updated by the various branches. Empirical results show that ForkMerge achieves positive transfer gains across several auxiliary task learning benchmarks.
3 Model Merging in Out-of-Distribution/Domain Generalization
The common goal of out-of-distribution generalization (OODG) and domain generalization (DG) is to improve a model’s performance on unseen data. The key difference between them is that OODG focuses on enhancing a model’s generalization ability on unknown data with significantly different distributions from the training data. In contrast, DG emphasizes improving a model’s generalization ability on unseen domains. Numerous recent studies have demonstrated that model merging contributes to enhanced training stability and overall performance in both OODG and DG.
In real-world scenarios, a trained model may be deployed in environments with changing distributions. For example, autonomous driving models are trained on a clean dataset, but in practice, they are vulnerable to unforeseen distributions such as natural corruptions (e.g., camera noise, motion blur) and more significant distribution shifts (e.g., summer to winter) . The goal of OODG is to enhance the model’s ability to generalize to unknown data that significantly differs from the training distribution.
Stochastic weight averaging (SWA) is a straightforward and widely used technique to improve machine learning models’ training stability and OOD performance. From a statistical perspective, weight averaging helps reduce variance during model training. Many works merge intermediate weight states (i.e., checkpoints) from training trajectories while training models . For example, WiSE fine-tuning demonstrates that linearly combining the weights of a pre-trained model and a fine-tuned model can significantly improve accuracy in the case of distribution shifts, while maintaining high accuracy on the original distribution. SWA simply averages all checkpoints from the beginning of a particular epoch to the end of training. This approach is explained to help the model converge to flat rather than sharp local optima, thereby improving generalization . Adaptive SWA highlights that executing SWA too early may lead to underfitting, while executing it too late may result in overfitting. It proposes averaging only when generalization on the validation set improves, effectively combining SWA with an early stopping mechanism. However, simple average weights are often suboptimal. In particular, TWA addresses this by showing that the averaging coefficients of the weights can be determined in a training manner. Consequently, TWA, unlike simple SWA, can perform averaging from the initial epoch of training, eliminating the need to define an additional hyperparameter for the epoch at which weight averaging should start.
In contrast to previous works that average weights obtained along one training trajectory, methods such as Model Soups , AdapterSoup , Model-Ratatouille , WARM , WARP , PAPA , WASH , DART , and DiWA propose merging multiple independently fine-tuned or trained models. These models are usually more diverse, which improves OOD performance. Independently trained models differ in hyperparameters (e.g., learning rate, weight decay, dropout), batch order, data augmentation techniques (e.g., random crops, horizontal flips), and the number of training steps, among other factors. Specifically, Model-Ratatouille , starts from the same initial model, fine-tunes multiple models on an auxiliary task, then continues to fine-tune these models on the target task, and finally merges the diverse models to improve OOD performance. WARM further increases the diversity of fine-tuned models by sampling different checkpoints from the trajectories of the pre-trained model as the initial weights for the downstream preference fine-tuning task. To reduce the additional cost of training multiple models, Model Stock proposes that we can exploit the geometric properties of the weight space and the anchoring effect of pretrained models to approximate the merged weights using only a few fine-tuned models. MEHL-Soup develops a scalable and efficient method to learn model merging coefficients for model soup. It only loads a subset of models for each iteration, significantly reducing the computation and memory requirements of naive model soup for learning merging coefficients.
The above analysis reveals that the SWA lacks diversity due to its reliance on a single trajectory. In contrast, Model Soups and DiWA train independently, which can lead to multiple models with significant differences, resulting in weight averaging failure. To balance these two approaches, Lookaround introduces a gradient descent optimizer based on the weight averaging. This optimizer iteratively performs ‘around’ and ‘average’ steps throughout the optimization process. In the ‘around’ step, multiple independent models are trained from the same starting point, each using different data augmentations. In the ‘average’ step, the diverse models are averaged, and the result is used as the starting point for the next iteration.
3.2 Model Merging for Better Domain Generalization
Domain generalization methods aim to generalize to an unknown target domain using only training data from source domains. For instance, in the context of traffic sign recognition, the training data for a machine learning (ML) model tasked with identifying traffic signs in various urban environments come from multiple cities (i.e., source domains). However, when deployed, the model must recognize traffic signs in new urban environments (i.e., target domains) that it has never encountered before. Existing DG methods can be classified into domain alignment, data augmentation, regularization, and meta-learning frameworks . Complementary to these approaches, model merging techniques can be seamlessly integrated to further improve out-of-domain performance without modification. Specifically, model merging in DG mainly occurs during the training process of the source domain model. Merging the intermediate weight states from different training stages helps improve the stability and generalization of the final model.
SWAD demonstrates that flatter minima generalize better to unseen domains. Inspired by SWA , SWAD proposes a dense and overfit-aware stochastic weight sampling strategy to identify these flatter minima. More specifically, unlike SWA, it starts from a predefined epoch until the final epoch, and collects a random weight every epochs for averaging. SWAD collects weights densely, that is, one is collected every iteration/step, and the start and end of random weight collection are determined by the performance changes on the validation set. EoA also shows that model averaging can improve out-of-domain performance stability, and that ensembling multiple moving average models can further enhance performance compared to ensembling models without weight averaging.
4 Model Merging in Federated Learning
Federated Learning (FL) is a distributed learning approach that allows multiple clients to collaboratively train a model without sharing data. FL primarily includes two settings: centralized (with a central server) and decentralized (without a central server). Each client updates the model or calculates the gradient based on local data and sends the updated information to the central server (in centralized FL) or other clients (in decentralized FL) for aggregation to update the global model, thus ensuring data privacy protection.
Model merging is a routine and crucial operation in FL. Taking centralized FL as an example, it typically involves clients and a central server . Each client has a private set of training data. Specifically, the training process in the centralized FL paradigm consists of five steps: (1) Model initialization: The central server initializes the global model parameter; (2) Model distribution: The latest model on the server is sent to the local client in the -th round of communication. (3) Update of the local model: The -th client updates the model by calculating the gradient based on the local data. (4) Model upload: The updated models of all local clients are sent to the server aggregation. (5) Model aggregation: The multiple local models on the server are aggregated. These five steps are repeated until the model converges or the maximum number of training rounds is reached. Since this paper is not a survey of FL, we focus on implementing the ‘model aggregation’ step. In FL, model merging refers to summarizing model parameters from various clients during each communication round, thereby forming an updated global model.
4.2 Model Merging for Local Knowledge Aggregation
Most FL methods adopt a simple coordinate-wise average to aggregate the local models. For example, they calculate local model merging coefficients according to some heuristic rules. FedAvg , the most classic FL method, proposes to merge local models on the server weighted by the amount of training data from each client. FedNova normalizes and scales model updates on the client side based on the number of update steps, efficiently aggregating local models to obtain a high-performance global model. FedAtt calculates layer-wise attention coefficients based on the similarity of client and server parameters, fusing local models based on these coefficients. FedFisher computes the Fisher information matrix of the parameters in each client to merge the local models. In more challenging FL tasks, the above direct coordinate-wise merging methods may result in suboptimal global model performance. Inspired by the property of permutation invariance of neural networks, PFNM , OTFusion and FedMA propose to permute neurons of local models before merging them. Similarly, GAMF transforms the model merging problem into a multi-graph matching problem based on graph matching and then merges the aligned local models.
5 Model Merging in Zero-shot/Few-shot Learning
In practical applications of machine learning models, collecting a large amount of labeled data can be expensive or infeasible in specific scenarios (e.g., medical diagnosis, real-time monitoring). Users often want deep models to effectively perform new tasks that have not been encountered before, that is, an ability commonly referred to as cross-task generalization . Zero-shot and few-shot learning can reduce the dependence on large amounts of data and allow the model to better deal with unseen categories or small numbers of samples, improving the cross-task generalization ability of the model. In few-shot learning, a common approach is to fine-tune the model using the limited examples available. However, because of the minimal data, this fine-tuning process is often unstable and yields only modest performance improvements. Recently, some studies have explored merging pre-trained models (from some publicly accessible resources) to enhance cross-task generalization under zero-shot and few-shot conditions.
Model merging has demonstrated the effectiveness of zero-shot learning across several applications. Some examples of practical applications include cross-lingual transfer , hybrid style image generation , and multi-modal processing .
Some works achieve cross-lingual transfer through model merging, such as chat , text summarization , or reasoning . A well-performing language-specific LLM needs to be fully trained, and with 7,000 languages in the world, not all of them have enough labeled data to support model fine-tuning. Therefore, cross-lingual knowledge transfer is particularly important. For example, Huang et al. build a Chat vector based on fine-tuned LLAMA2-chat and pre-trained LLAMA2 on chat data in the English language, and assembles it with the continuously pre-trained LLAMA2 model on other non-English languages. This allows the new model to chat in non-English languages. Chronopoulou et al. develop a zero-shot multilingual summarization framework. It uses a merged model (a supervised summarization model and an unsupervised pre-trained model for a high-resource language, along with an unsupervised pre-trained model for a low-resource language) to perform text summarization tasks in low-resource languages. Similarly, AdaMergeX demonstrates the effectiveness of model merging for cross-language transfer across three tasks: reasoning, natural language understanding, and natural language generation. In the hybrid style image generation task, Diffusion Soup and MaxFusion show that the zero-shot generation ability can be enhanced by merging multiple diffusion models. In the multi-modality task, DAMC experiments prove that zero-shot multi-modal extension can be achieved by merging multi-modal models, provided they are initialized from the same LLM. For example, by merging a visual LLM and an audio LLM, the combined model can not only perform image or audio tasks independently but also acquire the zero-shot ability to process inputs containing both visual and auditory information simultaneously.
5.2 Model Merging for Cross-task Generalization in Few-shot Learning
Parameter-efficient fine-tuning (PEFT), such as LoRA or Adapter, facilitates the creation and sharing of thousands of custom PEFT modules, each trained on different data for various downstream tasks. A natural question is whether combining PEFT modules pre-trained on different upstream tasks can improve the transfer accuracy for unseen downstream tasks with limited samples.
Recent work on model merging suggests a positive answer, showing that merged models can enhance generalization in few-shot settings . For example, LoraHub proposes to merge LoRA modules available on HuggingFace to achieve adaptive performance for unseen tasks, where the merging coefficients of different LoRA are searched in a black-box gradient-free manner with few-shot samples. As expected, few-shot LoraHub performs better than few-shot in-context learning and reduces inference costs by eliminating the need for examples as input to LLMs. LoraRetriever further proposes dynamically retrieving the most relevant LoRAs based on the input and merging them. Similarly, MerA proposes merging pretrained adapters into a single adapter for few-shot NLP scenarios. In general, well-trained LoRAs or adapters can serve as valuable resources that users can easily share, access, and apply to a variety of downstream tasks. In the real world, upstream and downstream tasks can be entirely disparate, originating from different datasets, domains, or even different parts of the same dataset. Asadi et al. comprehensively evaluates model merging in the few-shot learning setting. Specifically, this study examines three cases of label, domain, and task drift between upstream and downstream tasks. The results demonstrate that model merging enhances the model’s generalization ability in few-shot learning scenarios across different contexts.
6 Model Merging in Adversarial Learning
In the machine learning community, the open-source availability of pre-trained models has accelerated technological advancements. In this context, developers often download unvalidated checkpoints to fine-tune their models or even outsource the training process to third-party platforms . Consequently, open-source models are also vulnerable to malicious attacks, such as poisoning attacks, where hidden malicious behaviors can be triggered by specific inputs. This raises several intriguing questions: Can model merging lead to attacks, and can it be used to develop defense mechanisms? Additionally, how can intellectual property protection be enhanced in the context of model merging?
Parameter-Efficient Fine-Tuning (PEFT) methods , such as LoRA , exhibit functional transferability. This means that a LoRA model fine-tuned for a specific task based on a pretrained model can be successfully transferred to another pretrained model . In practice, developers often download LoRA models from open-source platforms to address their specific downstream tasks . If a poisoned LoRA, which could be seen as a Trojan horse, is inadvertently downloaded and integrated into a model, it may introduce security vulnerabilities. Research by LoRA-as-an-Attack demonstrates that merging a poisoned LoRA—trained on compromised data—with a benign LoRA, trained on clean data, can result in a backdoor injection. This finding also holds when multiple LoRAs are merged. In addition, BadMerging has developed a two-stage backdoor attack framework specifically for model merging, and through a large number of experiments, it has shown that the success rate of on-task and off-task attacks on merged models exceeds 90%, and existing defense measures cannot defend against BadMerging.
6.2 Model Merging as a Defense Strategy
Contrary to the attacks described in §4.6.1, the transferability of LoRA also offers an opportunity for model merging as a defense strategy. Specifically, if we know that a model may be susceptible to certain attacks, can we train some LoRAs to enhance the model’s defense (i.e., reduce the attacker’s success rate)? For example, Liu et al. demonstrate that GPT-3.5 was used to generate a benign dataset containing backdoor triggers. A dedicated defense LoRA was then trained on this benign data and merged into the poisoned pre-trained model. This defensive model merging ultimately led to a reduction in the backdoor effect. Furthermore, research has shown that in the context of full parameter fine-tuning, model merging can serve as a ”free lunch” for model defense. Experiments involving four model architectures and four datasets revealed that merging multiple poisoned models without additional effort mitigated these poisoning attacks, with the accuracy on the benign dataset remaining nearly unaffected. Rebuffi et al. and Croce et al. merge a set of (for various ) robust fine-tuned models to easily control the robustness level of each threat model against boundary adversarial attacks. Similarly, the experimental analysis by indicates that model merging offers an effective defense mechanism against jailbreak attacks .
In another practical scenario, merging unauthorized models may infringe on the intellectual property rights of the model owner. Malicious users might merge several high-quality open-source models (e.g., those authorized for research use only) to create a new model, then claim that this new model was entirely developed and trained from scratch by themselves, subsequently offering model services for commercial gain. In such cases, it becomes particularly crucial for model owners to detect whether others have merged their models. MergeGuard performs a preliminary analysis of the effectiveness of two existing defense methods—Quantization Watermarking and Instructional Fingerprint —in the context of model merging. The study observed that while the watermarking method cannot be detected in the merged model, the fingerprint method remains detectable.
Remaining Challenges and Future Directions
Although §2, §3 and §4 present various advanced model merging methods and applications, challenges remain in the technology and application of existing model merging approaches. Additionally, there are numerous areas that warrant further research in the future.
(1) Closing the Performance Gap Between the Merged and Independent Models. In practical settings, guaranteeing the performance of model merging remains challenging. The effectiveness of current model merging techniques heavily relies on the ”pretraining-finetuning” paradigm. Specifically, successful model merging requires that multiple models be fine-tuned based on the same pre-trained model, with careful control over the number of epochs and learning rate during fine-tuning. If these hyper-parameters are not set properly, the models may not converge in the same or close basin. Even based on the pre-trained fine-tuning paradigm, there is still a significant gap between the merged and independent models, especially when the number of models/tasks is large. Therefore, a promising direction for future research is to explore how to ensure the effectiveness of model merging under more relaxed conditions. For example, investigating how to merge multiple models that are trained independently from scratch for different tasks without compromising performance could be valuable.
(2) In-depth Theoretical Analysis for Model Merging. The validity and explanation of existing model merging techniques are largely empirical and lack sufficient theoretical guarantees. As discussed in §2.4, there is currently a limited amount of work on the theoretical aspects or explanations of model merging. The few existing studies mainly focus on merging multiple models trained on the same trajectory or on the same dataset with different fine-tuning settings. There is almost no theoretical research or explanation concerning merging multiple models fine-tuned on different datasets or merging multiple models trained from scratch on different datasets. Therefore, future research should aim for a more comprehensive and in-depth theoretical analysis to enhance the success and reliability of model merging. Furthermore, a deeper understanding of the effectiveness of model merging can, in turn, facilitate the discovery of prior conditions that are more conducive to model merging.
(3) Trustworthy Model Merging. Model merging is prone to intellectual property disputes and poisoning attacks, making the development of a reliable and trustworthy merging scheme an urgent research priority. Research on the reliability of model merging can be categorized based on two key roles: the model owner and the model combiner. On one hand, for model owners, protecting the intellectual property of their models is a primary concern. This protection involves both active and passive defense strategies: (1) Active Defense: Model owners may want to ensure that their published models are used independently and not merged by other users. The ideal outcome of an active defense strategy is that the model performs stably when used as intended, but breaks down completely if merged with other models. (2) Passive Defense: When model owners suspect that their models have been merged, there needs to be a robust method to verify whether the merged models contain their original models. On the other hand, for the model combiner, a key research direction is how to effectively prevent the inclusion of malicious injections, such as backdoors or poisoning attacks, when merging a set of authorized models.
(4) Effective and Efficient Model Merging. Existing high-performance model merging methods often come with significant costs in terms of efficiency and memory. First, most of these methods require all models to be loaded into memory during execution. For instance, merging 72 fine-tuned ViT-B/32 models necessitates more than 200GB of memory . Additionally, heuristics for determining model merging coefficients involve repeated evaluations of the combined model, while learnable methods depend on additional data and training. In the future, it would be beneficial to develop more efficient model merging methods that do not require training, additional data, GPUs, or large amounts of memory.
(5) Merge Heterogeneous Expert Models. Existing methods primarily focus on merging homogeneous models. However, in practice, numerous heterogeneous models excel in various tasks. A limited number of existing methods for merging heterogeneous models involve transforming multiple heterogeneous models into homogeneous ones using knowledge distillation techniques, followed by the merging process . The distillation process relies on the data from the original tasks and involves costly training. Therefore, it is also worth exploring approaches to merge these heterogeneous models without incurring the high costs associated with architectural transformations.
(6) Interdisciplinary Application of Model Merging. As discussed in §3 and §4, model merging has been adeptly applied across various foundation models and machine learning subfields to address different challenges and achieve interesting tasks. The question of how to adapt model merging strategies from one subfield to another presents an exciting avenue for exploration.
Conclusions
Model merging is a straightforward and effective technique for model enhancement that combines multiple models to achieve diverse capabilities. In this survey, we first provide a comprehensive overview of the advanced methods and theories currently available in the field of model merging. Next, we discuss the application of model merging techniques across various foundation models (i.e., LLMs, MLLMs) and more than ten subfields of machine learning, highlighting their use in addressing various challenges and difficulties. Finally, we identify ongoing issues within the model merging and propose six research directions that are worthy of further exploration. We believe that model merging technology, as an efficient and modular model empowerment solution, will play a significant role in more practical scenarios in the future.