Exploring Mode Connectivity for Pre-trained Language Models

Yujia Qin, Cheng Qian, Jing Yi, Weize Chen, Yankai Lin, Xu Han, Zhiyuan Liu, Maosong Sun, Jie Zhou

Introduction

Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP (Han et al., 2021), with the state-of-the-art across various NLP tasks consistently being pushed (Devlin et al., 2019; Liu et al., 2019b; Raffel et al., 2020). Through large-scale self-supervised training, PLMs acquire versatile semantic (Liu et al., 2019a) and syntactic (Tenney et al., 2019) knowledge, which could be utilized when conducting transfer learning on downstream tasks.

From the perspective of parameter space, PLMs provide generic initialization for downstream adaptation. Starting from the initialization, many high-performance minima can be found through gradient-based optimization. Up to now, plenty of works have studied how to effectively and efficiently adapt PLMs to high-performance minima, including adjusting hyperparameters (Liu and Wang, 2021), conducting transfer learning using auxiliary training data (Pruksachatkun et al., 2020), tuning PLMs in a parameter-efficient way (Ding et al., 2022), etc. Under different adaptation configurations, PLMs may finally reach local minima distributed in highly distinct regions. Although these minima all correspond to excellent performance (low loss), little has been known about their geometric connection in the parameter space.

A straightforward way to explore such geometric connection is to look into the loss landscape around different minima, which is inherently intractable due to the high dimensionality brought by the tremendous parameter size of PLMs. Instead of probing the full landscape, we propose to investigate the relation of different minima through the lens of mode connectivity (Garipov et al., 2018), which measures whether two different minima can be connected via a parametric path, along which the loss of the downstream task remains low. Exploring the mode connectivity of PLMs contributes to understanding the geometric connection among different minima. Such connection reflects the inherent relation of various adaptation configurations and may help us fathom the inner workings of PLM downstream adaptation under different settings.

To the best of our knowledge, systematic studies for the mode connectivity of PLMs are still lacking. In this paper, we first investigate what factors may affect PLM’s mode connectivity by answering the following research questions:

(Q1) How could different adaptation configurations (hyperparameters, tuning methods, and training data) affect PLM’s mode connectivity?

(Q2) How does mode connectivity change during pre-training?

We first consider the mode connectivity when different minima are trained on the same dataset. We investigate the effects of several hyperparameters (e.g., training data order, initialization of the tunable parameters, training steps, etc.) on PLM’s mode connectivity, and find that among these factors, initialization has the greatest impact. In addition, we show that fine-tuning leads to better mode connectivity than parameter-efficient delta tuning (e.g., adapter (Houlsby et al., 2019)).

Then we extend the connectivity analysis to minima trained on different datasets. We demonstrate that: (1) the mode connectivity is good for two minima trained on data belonging to the same distribution, but without overlap of specific instances. This means instead of memorizing training data, PLMs learn advanced task-level knowledge during training, and mode connectivity originates from the high overlap of task knowledge of two minima. (2) Although two minima trained on different tasks are inherently disconnected, pre-training gradually pulls the optimal regions of different tasks closer in an implicit way. This phenomenon may help explain PLM’s excellent cross-task transferability.

Beyond exploring the effects that could affect PLM’s mode connectivity, we also study the intrinsic properties of model solutions between two minima, which leads to the third question:

(Q3) How does the PLM’s task knowledge change along the path connecting two minima?

In the experiments, we observe that for two minima obtained independently on two tasks, when traversing from the minima trained on a source task to that of a target task, a PLM suffers from catastrophic forgetting (McCloskey and Cohen, 1989) on the source task, and gradually absorbs the knowledge from the target task. Besides, PLMs prioritize forgetting those elusive source knowledge and acquiring easy-to-grasp target knowledge.

In general, to fathom the connection of minima reached under different settings, we conduct empirical studies on the mode connectivity of PLMs. We also show that our findings may have potential significance in broad research areas, such as designing better ensemble methods for PLMs, understanding the task-level transferability of PLMs and revealing the mechanism of PLM downstream adaptation. We expect our evaluation setup and findings could inspire more future works in this field.

Related Work

To effectively and efficiently utilize the knowledge learned during pre-training, many strategies have been developed to better tune a PLM, including: (1) hyperparameter search, which aims to find an optimal hyperparameter configuration through traditional grid search or modern automated search (Liu and Wang, 2021); (2) pre-finetuning, which trains PLMs on intermediate auxiliary tasks before fine-tuning on a target task (Pruksachatkun et al., 2020; Aghajanyan et al., 2021). In this way, PLMs achieve better downstream performance by taking advantage of cross-dataset knowledge transfer; (3) prompt learning, which casts downstream tasks into the form of the pre-training objective by leveraging natural language prompts (Brown et al., 2020; Schick and Schütze, 2021a, b). Prompt learning exhibits superior performance especially under the few-shot and zero-shot scenarios; (4) delta tuning (also known as parameter-efficient tuning). Optimizing all the parameters of a PLM is computationally cumbersome. As a lightweight alternative, delta tuning optimizes only a few tunable parameters and keeps other parameters frozen, achieving comparable performance to fine-tuning (Ding et al., 2022).

Although plenty of adaptation strategies have been developed to better tune a PLM, little is understood about the connection of local minima reached under different training configurations. In this work, we take the first step by utilizing mode connectivity as the analysis tool.

Mode Connectivity for Neural Networks.

Despite being extremely high-dimensional, the loss landscape of neural networks exhibits a simple geometric pattern of mode connectivity (Garipov et al., 2018; Freeman and Bruna, 2017; Draxler et al., 2018). It is shown that starting from different initialization, the local minima obtained by gradient-based optimizations are often connected by low-loss paths, along which high-performance solutions could be easily found and ensembled to achieve better performance (Garipov et al., 2018). These paths are typically non-linear curves, which require a process of curve finding with task supervision. Excellent mode connectivity indicates that different minima are not isolated points in the parameter space, but essentially form a connected manifold (Draxler et al., 2018).

Frankle et al. (2020) further contend that from the same initialization, local minima obtained with different training data order can be connected by a linear low-loss path, reducing the burden of curve finding. Such a phenomenon is dubbed as linear mode connectivity, which is closely related to lottery ticket hypothesis (Frankle and Carbin, 2019), and has direct implications for continual learning (Mirzadeh et al., 2020). Compared with the non-linear counterpart, linear mode connectivity is a stronger constraint, requiring that the convex combination of two minima stay in the same loss basin.

Previous works typically study mode connectivity using non-pretrained models in the field of computer vision. Until recently, Neyshabur et al. (2020) observe linear mode connectivity on pre-trained vision models. Despite the great efforts spent, a systematic understanding of the mode connectivity of PLMs is still lacking. In this paper, we focus on investigating the effects that would influence PLM’s mode connectivity and analyze the knowledge variation along the connecting path. Different from existing works that study mode connectivity for minima trained on the same dataset, we additionally extend the analysis to different datasets.

Mode Connectivity Evaluation

Connecting Path.

Connectivity Criterion.

After defining the curve connecting both minima, we traverse along the curve and evaluate the loss and performance of the interpolations. We deem two minima θC1\theta_{C_{1}} and θC2\theta_{C_{2}} mode connected if there does not exist a significant loss barrier or performance drop along the defined curve between θC1\theta_{C_{1}} and θC2\theta_{C_{2}}. In the experiments, we evaluate evenly distributed interpolations on ϕ(α)\phi(\alpha).

Empirical Analysis

In this section, we conduct experiments to investigate the aforementioned research questions.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q1. (a) How could different hyperparameters and the specific tuning method affect the mode connectivity of PLMs?]

We first investigate the effects of several hyperparameters that could affect PLM’s mode connectivity, including (1) training data order, initialization of tunable parameters, training step (main paper), (2) learning rate and batch size (§ B.1). To explore the effects of the specific tuning method, we experiment with both fine-tuning and a representative delta tuning method, i.e., adapter (Houlsby et al., 2019). Adapter inserts tunable modules into a PLM and keeps other parameters fixed during adaptation. Unless otherwise specified, we mainly conduct the experiments using T5BASE\text{T5}_{\texttt{BASE}} (Raffel et al., 2020), and choose two representative NLP tasks (MNLI (Williams et al., 2018) and ReCoRD (Zhang et al., 2018)).

In each experiment, the training configurations of the two endpoints only differ in one hyperparameter, while other settings are kept the same for a fair comparison. To explore the effects of training steps, we evaluate the performance when both endpoints are trained for {1010k, 3030k, 5050k} steps, respectively. We evaluate 2424 evenly distributed interpolations and 22 endpoints along a linear path, i.e., we evaluate a series of ϕ(α)\phi(\alpha), where α∈{025,125,...,2525}\alpha\in\{\frac{0}{25},\frac{1}{25},...,\frac{25}{25}\}. Since we find that the trends of loss and performance are generally highly correlated (i.e., a performance drop corresponds to a loss barrier), we report the performance in the main paper and leave the results of loss in appendix D. All experiments are conducted 33 times with random seeds and we report the average results on test sets. For more training details, please refer to appendix C.

PLM’s downstream adaptation generally involves mini-batch gradient-based optimization, where training samples are learned in a random order. To explore its effect, we adapt two copies of a PLM with two different random data order. Then we visualize the performance of linear interpolations in Figure 1, from which we observe that for fine-tuning, both endpoints are well connected by a linear path; while for adapter tuning, there exists a slight but negligible performance drop near the midpoint. In general, we conclude that local minima are well connected under different random training data order.

Effects of Initialization.

Before downstream adaptation, additional parameters (e.g., extra modules defined by delta tuning, the classification head, etc.) may be introduced; in addition, Wu et al. (2022) recently show that adding noise to the pre-trained weights improves the fine-tuning performance on downstream tasks. Thus, both fine-tuning and delta tuning require proper initialization for the tunable parameters. Since different initialization could lead to distinct optimization trajectories, we explore the effects of initialization on PLM’s mode connectivity.

Specifically, for those newly introduced modules, we randomly initialize them with a Gaussian distribution; for those pre-trained weights that require tuning, we add a random Gaussian noise. Two endpoints are initialized with the same configuration (e.g., mean and standard deviation of the Gaussian distribution), but different random seeds. The linear interpolation results are depicted in Figure 2, from which we observe that the mode connectivity of fine-tuning is generally good; while for adapter tuning, there exists a significant performance drop between two differently initialized minima. This means starting from different initialization, PLMs tend to reach non-connected local minima in the parameter space, especially for delta tuning. In short, initialization of tunable parameters has a great impact on mode connectivity.

Effects of Training Step.

As mentioned before, the experiments in Figure 1 and 2 are conducted when both minima are trained for {1010k, 3030k, 5050k} steps. Comparing the results at different training steps, we observe that (1) longer training leads to poorer connectivity for adapter tuning under certain cases; while (2) the mode connectivity of fine-tuning is good at different steps. In § B.2, we further show that (1) the mode connectivity becomes poorer when one endpoint is trained with more steps while the other is trained with fixed steps, and (2) with the training step increasing, the Euclidean distance between two minima is also prolonged, which may partially explain the poorer mode connectivity.

Effects of Tuning Method.

Comparing the results of fine-tuning and adapter tuning in Figure 1 and 2, we observe that in general, the linear mode connectivity of fine-tuning is better than adapter tuning. In other words, when using fine-tuning, PLMs are more likely to be optimized to linearly-connected minima. A similar phenomenon also occurs for minima trained with different learning rates or batch sizes (see § B.1). Considering that adapter optimizes only 2.382.38% parameters than fine-tuning, we hypothesize that more tunable parameters may yield better mode connectivity and leave more explorations as future work.

Different Minima are Generally Connected by a Non-linear Path.

Considering that linearity is a strong constraint for mode connectivity, even if a direct linear path connecting two minima incurs a high loss, both minima may still be connected by a low-loss non-linear path. To explore whether this holds for PLMs, we follow the setting of tuning adapters with different initialization, which has been shown in Figure 2 to have poor linear mode connectivity. We try to use the supervision from the downstream task to find a low-loss parametric path connecting two endpoints θC1\theta_{C_{1}} and θC2\theta_{C_{2}}. Following Garipov et al. (2018), we consider a quadratic Bezier curve defined as follows:

We visualize the performance of the interpolation on the found Bezier curve in Figure 3. We observe that the two minima are well-connected by the found Bezier curve, without a significant performance drop. In fact, such a low-loss Bezier curve exists for minima reached under various different settings (see more experiments in § B.3). Given the above results, we conjecture that there may exist multiple loss basins which are connected via a low-loss non-linear path, instead of a linear path. For most of the minima within the same loss basin, their convex combination also lies in this basin. In this sense, if two minima are connected linearly, then both of them probably lie in the same basin; otherwise in different basins (e.g., the case of adapter tuning with different initialization).

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q1. (b) What are the effects of training data?]

In previous experiments, we focus on the connectivity of two minima trained with the same dataset. From now on, we extend the mode connectivity to two minima trained on different datasets, focusing on two facets: data overlap and data domain.

Effects of Data Overlap.

PLMs have been demonstrated to be adept at memorizing the training data (Carlini et al., 2021, 2022). To show that the connectivity of both minima does not originate from PLM’s memorization, we explore whether such mode connectivity still exists when two minima are obtained on data belonging to the same distribution, but without overlap of specific training samples. Specifically, we partition the original training data of MNLI into two equal splits. Then we adapt two copies of T5BASE\text{T5}_{\texttt{BASE}} on either split using the same training configurations. The experiments are conducted using both fine-tuning and adapter tuning.

The performance of linear interpolations is recorded in Figure 4. The results show that two minima are well connected for both tuning methods, demonstrating that mode connectivity does not originate from PLM’s memorization of specific training data; instead, during training, PLMs learn advanced task-level knowledge, and the connectivity reflects the high overlap of task knowledge of two local minima.

Effects of Data Domain.

PLMs are shown to generalize well on out-of-distribution data (Hendrycks et al., 2020), implying the connection of minima trained with different data distributions. To gain a deeper understanding, we choose two natural language inference datasets (MNLI and ANLI (Nie et al., 2020)), and two sentiment analysis datasets (Rotten Tomatoes (Pang and Lee, 2005) and Yelp Polarity (Zhang et al., 2015)) sourced from different domains. Then we fine-tune two copies of T5BASE\text{T5}_{\texttt{BASE}} on two datasets of the same task, and evaluate the linear mode connectivity between two minima. Note previous works typically study mode connectivity on the same dataset; while in our setting, we extend the analysis by evaluating the interpolations on two datasets.

The results are shown in Figure 5. Take the NLI task as an example, starting from one endpoint (α ⁣= ⁣0\alpha\!=\!0) of a source task (MNLI), with the interpolation approaching the other endpoint (α ⁣= ⁣1\alpha\!=\!1) of the target task (ANLI), the performance on MNLI / ANLI exhibits almost a monotonous drop / rise. Besides, there does not exist a performance valley where the performance is significantly lower than both endpoints. Intuitively, the performance change reflects the variation of the interpolation’s task knowledge along the connecting path. Due to the difference in data domain, the task knowledge of two endpoints only partially overlap with each other. When traversing from the source minimum to the target minimum, PLM suffers from catastrophic forgetting on the source knowledge, but gradually absorbs target knowledge, leading to the performance drop / rise on the source / target task. We defer more in-depth analyses to Q3.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q2. How does mode connectivity change during pre-training?] Previous works demonstrate that compared with random initialization, the initialization obtained by pre-training leads to a wider loss basin after downstream adaptation (Hao et al., 2019; Neyshabur et al., 2020). Intuitively, if a local minimum lies in a more flat basin, it should be easier to connect with other minima. In this sense, pre-training may be closely related to mode connectivity. To investigate this, we re-train a RoBERTaBASE\text{RoBERTa}_{\texttt{BASE}} model from scratch and explore how mode connectivity changes at different pre-training steps. We follow the pre-training setting of Liu et al. (2019b), with more details described in § C.4.

Pre-training Facilitates Mode Connectivity.

We select a series of checkpoints at different pre-training steps. For each checkpoint, we adapt two copies on MNLI using different initialization, and evaluate the performance of their linear interpolations. From Figure 6, we observe that for both fine-tuning and adapter tuning, with the pre-training step becoming larger, the mode connectivity of the PLM becomes better. This implies that pre-training implicitly facilitates the mode connectivity. Specifically, when using fine-tuning, there does not exist a performance drop for checkpoints pre-trained with more than 6.256.25k steps. Considering that pre-training with a batch size of 20482048 for 6.256.25k steps corresponds to almost 55% the computational cost of BERT (a batch size of 256256 for 10001000k steps), we conclude that PLMs acquire good mode connectivity at an early stage of pre-training.

Pre-training Pulls Task Boundary Closer.

Further, we look into the performance variation along a linear path between two minima trained on two different tasks (MNLI and SST-2 (Socher et al., 2013)). Similarly, we choose a series of checkpoints at different pre-training steps. Then for each checkpoint, we adapt two copies on MNLI and SST-2 under the same setting, and conduct linear interpolation between two adapted weights. We also conduct experiments on MNLI and QQP (link) in § B.6. We evaluate the performance of each interpolation on both tasks and illustrate the results in Figure 7. It can be derived that (1) due to the inherent difference of both tasks, the minimum obtained on one task achieves the performance near random guess (≈50%\approx 50\% for SST-2 and ≈33.3%\approx 33.3\% for MNLI) on another task. This indicates that minima of different tasks are disconnected. (2) In addition, there is a strong general trend that for a checkpoint pre-trained longer, the intersection of both tasks’ high-performance regions becomes wider. In other words, the boundaries of both tasks’ optimal regions are gradually pulled closer by pre-training. (3) For the checkpoint pre-trained with 6060k steps, we do not observe a region where the performance on both tasks is poor. This means starting from the initialization of a PLM pre-trained with enough steps, the optimal regions of various downstream tasks are closely packed. This finding may help explain PLM’s cross-task transferability, and we leave more discussions in § 5.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q3. How does the task knowledge change along the path connecting two minima?]

Having shown that mode connectivity reflects the high overlap of task knowledge of different minima, we further investigate the knowledge variation along the path connecting two minima. To quantify a model’s task knowledge, we resort to the memorization of the training data as a rough estimate. In experiments, we evaluate two minima obtained on data of different distributions as mentioned in Q1. (b)We choose this setting because (1) there does not exist a performance valley between two minima, which means the knowledge is properly combined, and (2) the knowledge of both tasks is diverse enough..

Specifically, we adapt two copies of T5BASE\text{T5}_{\texttt{BASE}} on MNLI (source task) and ANLI (target task), respectively. Denote θs\theta_{s} and θt\theta_{t} as two minima trained on the source dataset Ds={xi,yi}i=1∣Ds∣\mathcal{D}_{s}=\{x_{i},y_{i}\}_{i=1}^{|\mathcal{D}_{s}|} and the target dataset Dt={xi,yi}i=1∣Dt∣\mathcal{D}_{t}=\{x_{i},y_{i}\}_{i=1}^{|\mathcal{D}_{t}|}. We investigate the knowledge variation from θs\theta_{s} to θt\theta_{t} by choosing 44 evenly distributed linear interpolations (ϕ1\phi_{1}, ϕ2\phi_{2}, ϕ3\phi_{3}, ϕ4\phi_{4}) and 22 endpoints (ϕ0\phi_{0}, ϕ5\phi_{5}), i.e., ϕj=θs+j5⋅(θt−θs)\phi_{j}=\theta_{s}+\frac{j}{5}\cdot(\theta_{t}-\theta_{s}), j ⁣∈ ⁣{0,1,...,5}j\!\in\!\{0,1,...,5\}, where ϕ0 ⁣= ⁣θs\phi_{0}\!=\!\theta_{s}, ϕ5 ⁣= ⁣θt\phi_{5}\!=\!\theta_{t}. Then we measure whether each source training sample xi ⁣∈ ⁣Dsx_{i}\!\in\!\mathcal{D}_{s} is memorized (correctly classified) by ϕj\phi_{j}. We find empirically that with ϕj\phi_{j} approaching θt\theta_{t}, training samples of Ds\mathcal{D}_{s} are gradually forgotten, but seldom re-memorized under this setting. Therefore, we only record those newly forgotten samples for ϕj\phi_{j} (i.e., those classified correctly by ϕj−1\phi_{j-1} but wrongly by ϕj\phi_{j}) and denote them as Fj\mathcal{F}_{j}. Similarly, we denote those newly memorized samples of Dt\mathcal{D}_{t} as Mj\mathcal{M}_{j} (i.e., those classified wrongly by ϕj−1\phi_{j-1} but correctly by ϕj\phi_{j}).

After that, we characterize the role of each sample using dataset cartography (Swayamdipta et al., 2020). For a brief introduction, each sample of Ds\mathcal{D}_{s} (Dt\mathcal{D}_{t}) is characterized by the training dynamics of θs\theta_{s} (θt\theta_{t}). Take Ds\mathcal{D}_{s} as an example, assume we train the PLM for EE epochs on Ds\mathcal{D}_{s}, and the weights of the PLM are adapted to θs(e)\theta_{s}(e) after the ee-th epoch, where 1 ⁣≤ ⁣e ⁣≤ ⁣E1\!\leq\!e\!\leq\!E. For each training instance (xi,yi)∈Ds(x_{i},y_{i})\in\mathcal{D}_{s}, denote Pθs(e)(yi∣xi)\mathcal{P}_{\theta_{s}(e)}(y_{i}|x_{i}) as the probability θs(e)\theta_{s}(e) assigns to the true label, we record the PLM’s prediction after each epoch and calculate two statistics:

∙\bullet confidence measures how confidently the PLM assigns the true label to a given input, it is defined as the mean probability of the true label:

∙\bullet variability captures how consistently the PLM judges each training instance, it is defined using the standard deviation of the true label’s probability:

After obtaining both statistics for each training sample of Ds\mathcal{D}_{s} and Dt\mathcal{D}_{t}, we illustrate the average statistics for newly forgotten / memorized samples (Fj\mathcal{F}_{j} / Mj\mathcal{M}_{j}) in Figure 8. We observe that with ϕj\phi_{j} approaching θt\theta_{t}, the average confidence of the newly forgotten data gradually increases, while the variability gradually drops; symmetrically, the average statistics of the newly learned data exhibit an opposite trend. According to Swayamdipta et al. (2020), instances with high confidence but low variability are generally easy-to-learn ones; while those with low confidence are generally ambiguous or hard-to-learn data. In this regard, when gradually leaving the source minimum, the PLM prioritizes forgetting the source knowledge of those difficult instances, and then forgets the source knowledge of the easy-to-learn data. On the contrary, the easy-to-learn target knowledge is learned before the elusive and obscure target knowledge.

Discussion

The property of linear mode connectivity is related to recent explorations of weight averaging (Wortsman et al., 2021, 2022; Matena and Raffel, 2021), which combines independently fine-tuned models in the parameter space. In this way, the knowledge from multiple models can be merged. Our findings have direct implications for designing better weight averaging methods: (1) for two minima, weight averaging can be seen as choosing the midpoint on the linear path. We have shown that a non-linear curve may have better mode connectivity under certain cases. This implicates that linear interpolation may not find the optimal combination despite its simplicity; instead, there may exist better methods to ensemble weights (see experiments in § B.8); (2) our findings on the effects of different training configurations can also inspire choosing more appropriate models (with better mode connectivity) to ensemble.

Task-level Transferability.

Although PLMs are demonstrated to have excellent cross-task transferability (Vu et al., 2020; Pruksachatkun et al., 2020; Poth et al., 2021; Su et al., 2022), it is still under-explored why PLMs have such an ability. Our findings that pre-training implicitly pulls the task boundary closer may help explain this phenomenon. Since the optimal regions of various tasks are packed closely, PLMs are easier to traverse across the task boundary, without getting blocked by a loss barrier.

Knowledge Quantification.

Investigating the knowledge variation along the connecting path helps better understand how different model knowledge is merged. Quantifying the task knowledge of various models may also provide insights for research topics like knowledge distillation (Hinton et al., 2015) and knowledge transfer (Weiss et al., 2016). While we use training data memorization as a rough estimate for task knowledge, it would be interesting to explore whether there exist more granular methods, such as knowledge probing (Petroni et al., 2019; Liu et al., 2019a).

Conclusions

In this paper, we conduct empirical analyses on the mode connectivity of PLMs, aiming to fathom the connection of minima reached under different settings. We investigate how different downstream adaptation configurations and pre-training affect PLM’s mode connectivity. In addition, we explore the knowledge variation along the path connecting different minima. In general, exploring the mode connectivity of PLMs contributes to understanding the inner workings of PLM downstream adaptation. We expect our evaluation setup and analyses could inspire more future explorations in this field.

Acknowledgments

This work is supported by the National Key R&D Program of China (No. 2020AAA0106502) and Institute Guo Qiang at Tsinghua University.

Yujia Qin designed the experiments and wrote the paper. Cheng Qian and Jing Yi conducted the experiments. Yankai Lin, Zhiyuan Liu, Maosong Sun, and Jie Zhou advised the project. All authors participated in the discussion.

The authors would like to thank anonymous reviewers for their valuable feedback.

Limitations

There are some limitations not well addressed in this paper:

The goal of this paper is to investigate the connection among different minima. However, since we use mode connectivity as the analysis tool, we only investigate the connection between two minima at a time. In this regard, it would be interesting to develop more advanced tools to explore the connection among multiple minima simultaneously, which is left as future work.

We do not give a strict mathematical definition for “good mode connectivity”. For instance, we do not set a specific threshold of performance drop (e.g., >5%>5\% for the case of “not well-connected minima”). We argue that this is because, all of our experimental results are significant enough, thus there is no need to follow a strict definition.

References

Appendices

Appendix A Details for Finding a Bezier Curve

We follow Garipov et al. (2018) to find a quadratic Bezier curve connecting two endpoints θC1\theta_{C_{1}} and θC2\theta_{C_{2}}. The Bezier curve is defined as follows:

where qθ(α)=∣∣ϕθ′(α)∣∣dα∫01∣∣ϕθ′(α)∣∣dαq_{\theta}(\alpha)=\frac{||\phi_{\theta}^{\prime}(\alpha)||d\alpha}{\int_{0}^{1}||\phi_{\theta}^{\prime}(\alpha)||d\alpha}. Since qθ(α)q_{\theta}(\alpha) is dependent on θ\theta, it is generally intractable to compute the original loss Lcurve(θ)\mathcal{L}_{\text{curve}}(\theta). To this end, Garipov et al. (2018) suggest optimizing a more computationally tractable loss as follows:

where α\alpha is sampled from a uniform distribution U(0,1)\text{U}(0,1) instead of qθ(α)q_{\theta}(\alpha). In experiments, we initialize θ\theta with 12θC1+12θC2\frac{1}{2}\theta_{C_{1}}+\frac{1}{2}\theta_{C_{2}} (i.e., starting from a linear curve), which makes training more stable than using randomly initialized weights or the pre-trained weights.

Appendix B Additional Experiments

We perform experiments to explore the effects of both learning rates and batch sizes. For the former, we evaluate when both endpoints are trained with a learning rate of {1×10−4,5×10−4}\{1\times 10^{-4},5\times 10^{-4}\} and {1×10−4,5×10−5}\{1\times 10^{-4},5\times 10^{-5}\}, and the batch size is set to 1616; for the latter, we chose a batch size of {16,8}\{16,8\} and {16,32}\{16,32\}, and the learning rate is set to 1×10−41\times 10^{-4}. The experiments are conducted using both full-parameter fine-tuning and adapter tuning on MNLI and SST-2 with T5BASE\text{T5}_{\texttt{BASE}}. For MNLI, we experiment when both endpoints are trained for {10k, 30k, 50k} steps; for SST-2, both endpoints are trained for {3k, 9k, 15k} steps. We illustrate the results of linear interpolation in Figure 9 and Figure 10. We could conclude from both figures that, the minima obtained by fine-tuning are always well-connected by the linear path; however, the connectivity of adapter is poor under certain cases. This is aligned with the finding in the main paper that the mode connectivity of fine-tuning is generally better than delta tuning. We also observe that with the training steps becoming larger, the connectivity of adapter tuning sometimes becomes poorer.

B.2 Additional Experiments for the Effects of Training Steps

In the main paper, when exploring the effects of training steps, we experiment with the setting where both endpoints are trained with the same number of steps. We show that mode connectivity could be poorer when both endpoints are trained for longer steps. To more rigorously investigate the effects of training steps, we experiment when both minima are obtained when using different training steps. Specifically, we adapt T5BASE\text{T5}_{\texttt{BASE}} model on MNLI and ReCoRD using both fine-tuning and adapter tuning. We train two endpoints with different initialization, which is implemented by utilizing different random seeds, but keeping the configuration of the initialization (mean and standard deviation of the normal distribution) the same. After that, one endpoint (denoted as θC1\theta_{C_{1}}) is adapted for 5050k steps, while the other endpoint is adapted for {1010k, 2020k, 3030k, 4040k} steps, and denoted as {θC210k\theta_{C_{2}}^{10\text{k}}, θC220k\theta_{C_{2}}^{20\text{k}}, θC230k\theta_{C_{2}}^{30\text{k}}, θC240k\theta_{C_{2}}^{40\text{k}}}, respectively. Then we evaluate the linear interpolations between {(θC1\theta_{C_{1}} and θC210k\theta_{C_{2}}^{10\text{k}}), (θC1\theta_{C_{1}} and θC220k\theta_{C_{2}}^{20\text{k}}), (θC1\theta_{C_{1}} and θC230k\theta_{C_{2}}^{30\text{k}}), (θC1\theta_{C_{1}} and θC240k\theta_{C_{2}}^{40\text{k}})}, respectively. The results are shown in Figure 11, from which we observe that, the mode connectivity of fine-tuning is generally good, while two minima of adapter are not well-connected. In addition, with the gap of the training steps between two endpoints becoming larger, the mode connectivity of adapter becomes poorer. These results suggest that the number of training steps can affect PLM’s mode connectivity under certain cases.

To better understand the reason why training steps could affect mode connectivity, we record the Euclidean distance variation of two endpoints during downstream adaptation. Both endpoints start from different initialization. We adapt the PLM on MNLI using both fine-tuning and adapter tuning for {1010k, 2020k, 3030k, 4040k, 5050k} steps, and obtain a series of checkpoints: {(θC110k\theta_{C_{1}}^{10\text{k}} and θC210k\theta_{C_{2}}^{10\text{k}}), (θC120k\theta_{C_{1}}^{20\text{k}} and θC220k\theta_{C_{2}}^{20\text{k}}), (θC130k\theta_{C_{1}}^{30\text{k}} and θC230k\theta_{C_{2}}^{30\text{k}}), (θC140k\theta_{C_{1}}^{40\text{k}} and θC240k\theta_{C_{2}}^{40\text{k}}), (θC150k\theta_{C_{1}}^{50\text{k}} and θC250k\theta_{C_{2}}^{50\text{k}})}. Then we calculate the Euclidean distance of two endpoints as: ∣∣θC1∗−θC2∗∣∣2||\theta_{C_{1}}^{*}-\theta_{C_{2}}^{*}||^{2}. The change of Euclidean distance is visualized in Figure 12, from which we observe that, with the training steps becoming larger, the distance between two endpoints is also prolonged. This may partially explain the poorer mode connectivity with the increasing of training steps. We have shown in the main paper that PLMs have multiple loss basins connected by a non-linear path, instead of a linear path. Within the same loss basins, most of the solutions have good linear mode connectivity. However, since the loss basin has a boundary, when the distance between two minima becomes large enough, they may finally cross the border of the loss basin. Under this scenario, the linear path connecting both endpoints would incur a high loss.

B.3 More Experiments for Mode Connectivity along a Non-linear Path

In the main paper, to explore the connectivity of two minima along a non-linear path, we experiment on the setting of tuning adapters with different initialization. This setting has been shown to have poor linear mode connectivity but good non-linear mode connectivity. In fact, in our pilot experiments, we find that such a low-loss non-linear curve exists for minima reached under various different settings. In this section, we provide some of the experiments to demonstrate the above finding using T5BASE\text{T5}_{\texttt{BASE}}.

Specifically, we experiment with three tasks: MNLI, ReCoRD, and SST-2. For adapter tuning, we choose the setting where two minima are trained with (1) different training data order and (2) data from the same distribution but without specific overlap of training instances. For (2), same as before, we randomly partition the original training dataset into two equal splits, and adapt two copies of PLM on each split. For fine-tuning, we choose the setting where two minima are trained with (1) different training steps (the setting is the same as § B.2), and (2) different initialization.

The performance of the interpolations are visualized in Figure 13 for adapter tuning and Figure 14 for fine-tuning. We observe that under all the settings, we do not observe a significant performance drop along the non-linear curve, showing that the connectivity is good. The above results demonstrate that PLMs may have multiple loss basins which are connected via a low-loss non-linear path. In this paper, following Neyshabur et al. (2020), we spend most of the efforts on convex hull and linear interpolation to avoid possibly trivial connectivity results.

B.4 Additional Experiments for the Effects of Data Overlap

In the main paper, when evaluating the effects of data overlap, we present the results on MNLI. In this section, we visualize the results when using ReCoRD and SST-2 in Figure 15. Other settings are kept the same as the main paper.

B.5 Additional Experiments for the Effects of Data Domain

In the main paper, when evaluating the effects of data distributions (data domain), we present the results when using fine-tuning. In this section, we visualize the results when using adapter tuning in Figure 16. Other settings are kept the same as the main paper.

B.6 Additional Experiments for the Change of Mode Connectivity during Pre-training

In the main paper, when evaluating the performance variation between two minima trained on two different tasks, we report the results of MNLI and SST-2. In this section, we present the results of MNLI and QQP in Figure 17. In fact, in our pilot studies, we find that the conclusions in diverse tasks are very consistent. Due to the concern about the energy cost, we only report the performance of two pairs of tasks.

B.7 Performance along the Connecting Path

We show that better performance could be achieved by interpolating two independently trained weights in the parameter space. Specifically, we choose the scenario where two copies of PLMs are trained with different training data order. As mentioned in Q1. (a) in the main paper, PLMs have excellent mode connectivity under this setting. We experiment with T5BASE\text{T5}_{\texttt{BASE}} using fine-tuning and adapter tuning on MNLI, and conduct both linear interpolation and curved interpolation. We evaluate the performance of 2424 evenly distributed points on the curve on a development set, select the best-performing one and evaluate its performance on the test set. We also compare the interpolation with the endpoints (we report the best performance of the two endpoints). All experiments are conducted 33 times and we report the average test results in Table 1. We observe that by traversing along the connecting curve between two minima, we could find a solution that performs better than both endpoints. In addition, traversing along a linear path finds an interpolation with higher performance than traversing along a curved pathAlthough we have shown that the non-linear mode connectivity is generally good for different minima, it does not mean that the best performance on a non-linear curve is always better than that on a linear curve.. In general, this finding demonstrates that it is promising to combine the knowledge of multiple models through weight averaging.

B.8 Other Strategies for Weight Ensemble

where σ\sigma denotes a sigmoid function. During training, both θC1\theta_{C_{1}} and θC2\theta_{C_{2}} are kept frozen, and only α\bm{\alpha} is tuned. We design three intuitive strategies for parameter division: (1) layer-wise division, where the parameters within the same layer share the same combination ratio; (2) module-wise division, where we discriminate the combination ratio of the feed-forward network module and multi-head attention module in each layer. This means each module in each layer is assigned with an individual combination ratio; (3) matrix-wise division, where each weight matrix in each module is combined individually. Matrix-wise division is the most fine-grained one among the above three strategies. Since we use the T5BASE\text{T5}_{\texttt{BASE}} model, which consists of 1212 encoders and 1212 decoders, the number of components MM for block-wise, layer-wise, and matrix-wise divisions are 2424, 6060, and 120120 for adapter; and 2828, 6464, and 282282 for fine-tuning. During training, we perform grid search on a series of learning rates {0.1,0.05,0.01}\{0.1,0.05,0.01\} and set the batch size to 88, and max training steps to 100100k.

Two endpoints are obtained by fine-tuning T5BASE\text{T5}_{\texttt{BASE}} on MNLI using different training data order. The results are shown in Table 2, from which we find that, among the proposed three combination strategies, matrix-wise division achieves the best performance. The performance is also better than using linear interpolation. This phenomenon demonstrates that there exist better ways for combining two minima’s knowledge than linear interpolation. We hope our findings on mode connectivity in this paper could inspire future works to design better weight ensemble methods.

Appendix C Training Details

For the T5BASE\text{T5}_{\texttt{BASE}} model, we use the checkpoint provided by Lester et al. (2021), who conducted additional 100100k steps of language modeling adaption on the official checkpoints released by Raffel et al. (2020). Such adaptation is demonstrated to help stabilize downstream adaptation and improve the performance, especially for delta tuning methods (Lester et al., 2021). We use AdamW (Loshchilov and Hutter, 2019) as the optimizer for all the experimented PLMs. All the implementation codes, trained checkpoints and used datasets would be released after publication.

We download all the experimented datasets from Huggingface Datasets (Lhoest et al., 2021). Since some datasets do not contain a test set, we first merge all the data points, and then split them into the new training split, development split, and test split with an approximate ratio of 8:1:18:1:1. The above procedure is conducted on all the experimented datasets.

For different tasks fine-tuned on T5BASE\text{T5}_{\texttt{BASE}}, we first conduct grid search to find an optimal hyperparameter combination. Specifically, the chosen hyperparameter of different tasks for fine-tuning is shown in Table 3; for adapter tuning, in our prior experiments, we find that a learning rate of 5×10−45\times 10^{-4} and a batch size of 1616 performs good on all tasks, thus we set them as the default configuration. For both tuning methods, we save 55 checkpoints during training, with different saving intervals for different tasks as shown in Table 3.

For adapter tuning, all the modules newly introduced are initialized using a Gaussian distribution. As for fine-tuning, we add Gaussian noise to all the tunable parameters. The mean and standard deviation of the Gaussian distribution are set to and 0.00020.0002, respectively. We use different random seeds to generate different initialization.

C.2 Additional Details for Curve Finding

When optimizing the Bezier curve, we set the learning rate to 1×10−41\times 10^{-4}, batch size to 88, max training steps to 55k for fine-tuning; and set the learning rate to 1×10−41\times 10^{-4}, batch size to 1616, max training steps to 1010k for adapter tuning. During curve finding, we evaluate the development performance of the current curve for every 100100 steps, using a series interpolations with α∈{0.25,0.5,0.75}\alpha\in\{0.25,0.5,0.75\}.

C.3 Additional Details for Calculating Confidence and Variability

As mentioned in the main paper, we use the training dynamics to characterize each training sample. For MNLI / ANLI, we adapt the model for 88 epochs / 2020 epochs to calculate both confidence and variability. We tune more epochs for ANLI because the size of its training dataset is far smaller than that of MNLI.

We closely follow the pre-training setting of Liu et al. (2019b), except that for pre-training data, we use the concatenation of Wikipedia and BookCorpus Zhu et al. (2015) same as BERT (Devlin et al., 2019), and we pre-train our model with a batch size of 20482048. The pre-training implementations for RoBERTaBASE\text{RoBERTa}_{\texttt{BASE}} are based on those of Qin et al. (2022a, b). Adam (Loshchilov and Hutter, 2019) is chosen as the optimizer. The hyperparameters for the optimizer is set to 1×10−61\times 10^{-6}, 0.90.9, 0.980.98 for ϵ\epsilon, β1\beta_{1}, β2\beta_{2}, respectively. We set the dropout rate to 0.10.1, weight decay to 0.010.01 and use linear learning rate decay. The model architecture is the same as the official RoBERTaBASE\text{RoBERTa}_{\texttt{BASE}} model (Liu et al., 2019b). The pre-training is conducted using 88 NVIDIA V100 GPUs.

Appendix D The Visualization of Loss for Interpolations

As mentioned before, we record both loss and performance for each interpolation. Since we find that the trends of loss and performance are generally highly correlated, due to the length limit, we only report the performance in the main paper. In this section, we visualize the loss for most of the experiments conducted in this paper, see Figure 18, Figure 19, Figure 20, Figure 21, Figure 22, Figure 23, Figure 24, and Figure 25.