Improving Sharpness-Aware Minimization with Fisher Mask for Better Generalization on Language Models
Qihuang Zhong, Liang Ding, Li Shen, Peng Mi, Juhua Liu, Bo Du, Dacheng Tao
Introduction
The “pretraining-finetuning” paradigm has become the de facto standard for the community of natural language processing (NLP) Devlin et al. (2019); Liu et al. (2019); Clark et al. (2019b); Raffel et al. (2020); Brown et al. (2020); Lewis et al. (2020). Given a pretrained language model (PLM), the dominant fine-tuning manner is tuning the entire pretrained parameters for each downstream task Radford et al. (2018); Devlin et al. (2019). While fine-tuning the entire PLM can improve performance on a wide range of NLP tasks, it usually suffers from over-fitting and poorer generalization ability Xu et al. (2021); Bahri et al. (2022), especially in the large-scale PLMs and limited training data scenarios.
Hence, some existing efforts attempt to provide more regularization in the fine-tuning stage Zhang et al. (2018); Müller et al. (2019); Xu et al. (2021), among which the optimization of the training loss is an intuitive and effective method. Specifically, motivated by the finding Keskar et al. (2016); Neyshabur et al. (2017) that the smoother loss landscape refers to the better model generalization, Foret et al. (2020) propose the “sharpness-aware minimization” (SAM) to simultaneously minimize loss value and loss sharpness, where the sharpness can be quantified as the maximized difference of loss when a perturbation is added to the current weights. In practice, SAM performs two forward-backward computations for each optimization step, where the first forward-backward is to obtain the perturbation for each model parameter and the second one is to update the parameters. Many prior works Wu et al. (2020); Zheng et al. (2021) show the effectiveness of SAM in the vision domain, motivated by this, Bahri et al. (2022) first apply the SAM to the language domain, more recently.
Although Bahri et al. (2022) empirically show the remarkable performance of SAM on several language understanding tasks, SAM calculates perturbations indiscriminately for all parameters, which is time-consuming and hinders the application of SAM. Furthermore, inspired by the finding Keskar et al. (2016) that only about 5% of parameters are sharp and rise steeply during optimization, we notice that not all parameters contribute equally to the optimization of training. Hence, this raises a question that whether we can calculate perturbations for only some individual parameters, and thus make the optimizer focus on these important parameters.
To this end, we propose a novel optimization approach, Fisher SAM (FSAM), which introduces a Fisher mask to improve the efficiency and effectiveness of SAM. In short, FSAM first uses the Fisher information Fisher (1922) as the metric to identify the sharper parameters We refer to these parameters as the important ones, because they will rise steeply during optimization and affect the model generalization significantly. and formulates a binary Fisher mask correspondingly. Then, the Fisher mask is multiplied with the perturbations to obtain the sparse perturbations, which are lastly used to perform regularization in the parameter update. In this way, only parts of sharper parameters will be added into the perturbations, and the optimizer can thus focus more on these important parameters. Also, the sparse perturbations could ensure the training acceleration via sparse back-propagation Since the fine-grained sparse training is limited to the hardware, we do not achieve actual sparse speedup in this work. Despite it, we still believe that FSAM has great potential to achieve true training acceleration in the future, with the development of hardware for fine-grained sparse operation.. Moreover, one may concern that the sparse Fisher mask would affect the convergence rate of FSAM Lin et al. (2019). Hence, we theoretically provide the convergence analysis of FSAM, ensuring that the convergence of FSAM is irrelevant to the Fisher mask.
We conduct a large-scale and systematic study to evaluate the performance and effectiveness of FSAM. Firstly, we apply SAM and FSAM to fine-tune various PLMs on parts of GLUE and SuperGLUE benchmarks, where the results show that FSAM consistently outperforms the vanilla SAM by 0.671.98 average score among these PLMs, and surpasses the Adam Kingma and Ba (2015) optimizer by 1.411.91 points. Secondly, we conduct experiments on two popular generation tasks (i.e., XSUM and CoNLL2014) and prove that FSAM can deliver promising results against SAM. Lastly, quantitative analysis and in-depth discussion demonstrate the universality and effectiveness of FSAM in various complex scenarios, and prove that FSAM indeed brings better model generalization. Specifically, we show that our Fisher mask strategy not only works well in the SAM, but also can be applied to other SAM variants.
To summarize, our contributions are two-fold: (1) We propose a novel optimization approach (namely FSAM) with theoretical convergence guarantee for PLMs. Specifically, FSAM improves the performance and efficiency of recently-proposed SAM via a Fisher mask strategy, which can also be applied to more SAM variants. (2) Extensive experiments show that FSAM consistently outperforms the SAM by a large margin on both language understanding and generation tasks. The systematic study demonstrates the effectiveness and universality of FSAM on improving model generalization.
Related Work
Hochreiter and Schmidhuber (1994) first show the strong correlation between the flat minima and the generalization of a model, inspired by this, Foret et al. (2020) propose the SAM to find a flat minimum and thus improve model generalization. While many existing works prove the effectiveness of SAM on various computer vision tasks Wu et al. (2020); Chen et al. (2021); Zheng et al. (2021), the double forward-propagation process of SAM brings more computational cost. To this end, Du et al. (2021) propose an Efficient SAM (ESAM) for reducing the computational cost of SAM. Additionally, there are also some efforts that focus on more efficient and effective SAM optimization Zhuang et al. (2021); Kwon et al. (2021); Mi et al. (2022).
Improving Generalization.
Recently, we have witnessed numerous PLMs that achieved tremendous success in the community of NLP Yang et al. (2019); Devlin et al. (2019); Brown et al. (2020); Lewis et al. (2020); Raffel et al. (2020); Joshi et al. (2020); He et al. (2020); Qi et al. (2021); Zhong et al. (2022). The current dominant fine-tuning approach needs to tune all pretrained parameters for each downstream task, which makes the PLM easily memorize the training data and thus leads to overfitting. To tackle this issue, some works attempt to provide implicit and explicit regularization into the training of models, such as dropout Srivastava et al. (2014), label smoothing Müller et al. (2019), mixup Zhang et al. (2018) and other data-augmentation methods Sennrich et al. (2016); Wang et al. (2018b); Zhong et al. (2021); Wang et al. (2022); Ding et al. (2022). On the other hand, motivated by the successful applications of SAM in the vision domain, Bahri et al. (2022) involve applying SAM to optimize the T5 Raffel et al. (2020) model on multiple language tasks and show that SAM can improve the generalization of PLMs effectively.
We depart from the prior work Bahri et al. (2022) and ours as follows: 1) different motivations: instead of verifying the effect of vanilla SAM on several language understanding tasks, we aim to improve the efficiency and effectiveness of SAM. 2) different contributions: our main contribution is to propose a fisher mask strategy, which can be applied to both SAM and its variants. 3) more analysis: we provide more experimental results and analysis towards the effectiveness of our method in more complex scenarios.
Methodology
In this section, we first review the Sharpness-Aware Minimization, and then propose our Sharpness-Aware Minimization with Fisher mask, coined as FSAM. Finally, we theoretically analyze the convergence of FSAM with adaptive learning rate.
Sharpness-Aware Minimization.
Foret et al. (2020) propose the Sharpness-Aware Minimization (SAM) to improve the generalization, which is achieved by the following min-max problem:
where is a predefined value to control the neighborhood size, and the is the perturbation vector on model weight. The optimization is expected that the model loss will not significantly rise with a certain amount of weight change controlled by , which is intuitively consistent with the generalization capacity of model.
With the Taylor expansion, the perturbation vector could be achieved approximately:
and the object function could be simplified as
The solution of the above function could be obtained by a two-step gradient descent. In the first gradient descent step, the perturbation vector is calculated by Equation 2. The second gradient descent step is the actual weight update.
However, despite the improvement of SAM on many tasks, SAM requires a two-step gradient calculation which leads to the double overhead compared to the conventional optimizer, e.g., Stochastic Gradient Descent (SGD) and Adam.
2 Sharpness-Aware Minimization with Fisher Mask
In this subsection, we propose the Sharpness-Aware Minimization with Fisher Mask (FSAM) in detail, which reduces the computation of SAM by sparse calculation.
To be specific, we compute only a fraction of the elements in the perturbation vector , which would be multiplied by a sparse binary mask . To control the amount of perturbation, the sparse mask satisfies , where the is the predefined sparse ratio and empirically set to 0.9. The objective function of FSAM is denoted as
where is the Hadamard product, i.e., the element-wise multiplication. For the stability of optimization, we update the mask with a fixed interval (denoted as Fi) during training. The algorithm of FSAM is shown in Algorithm 1.
To find the optimal mask during training, we apply the Fisher information to achieve sparse perturbation. The Fisher information is proposed by Fisher (1922) to measures the information carried by an observable random variable about the unknown parameters of the distribution. The Fisher information is defined by
The second expectation is over , which can be achieved by the label for data in supervised learning. Finally, we calculate the Fisher information as "Empirical Fisher":
where is the set whose elements in the mask are 1, i.e., , and returns the top largest values among . On the other hand, the other weights with small Fisher values will not be perturbed, i.e., the corresponding element in mask will be set to 0:
3 Theoretical Analysis
In this subsection, we theoretically analyze the convergence and generalization of FSAM. Due to the space limitation, we only show the convergence analysis here, and the generalization analysis and whole proof are presented in Appendix A.1.
(-smooth.) Consider is differentiable with gradient Lipschitz property: It exists s.t.
(Bounded stochastic gradients.) The variance of stochastic gradient is bounded:
(Bounded gradient.) The stochastic gradient is bounded: It exists s.t.
Consider the function under the assumption 1,2,3, and a fixed base learning rate satisfies that , we have
The Theorem 1 shows that when is large, FSAM could achieve the linear speedup convergence rate with respect to mini-batch size under the setting of and , i.e.,
Experimental Setup
To investigate the effectiveness and universality of our FSAM method, we conduct extensive experiments on various NLP tasks. Specifically, different from Bahri et al. (2022) that only verify the method on several language understanding tasks, we evaluate our method on both language understanding and generation tasks.
Following many previous works Vu et al. (2022); Bahri et al. (2022); Zhong et al. (2022), we conduct experiments on a combination of tasks from GLUE Wang et al. (2018a) and SuperGLUE Wang et al. (2019) benchmarks, including linguistic acceptability (CoLA), natural language inference (RTE, CB), paraphrase and similarity (MRPC and STS-B), question answering (BoolQ), word sense disambiguation (WiC) and coreference resolution (WSC). In practice, we evaluate the performance with Accuracy (“Acc.”) metric for most tasks, except the additional F1 score for MRPC, the Pearson-Spearman correlations (“Pear./Spea.”) for STS-B and the Matthew correlation (“Mcc.”) for CoLA.
Language Generation Tasks.
We also use two popular generation tasks following Liu et al. (2021); Zhang et al. (2022) as the benchmarks, i.e., abstractive summarization (XSUM) and grammatical error correction (CoNLL2014). For the XSUM, we report results in terms of standard ROUGE metrics Lin (2004), i.e., Rouge-1, Rouge-2 and Rouge-L, respectively. For the CoNLL2014, MaxMatch scores Dahlmeier and Ng (2012) are used for evaluation with Precision, Recall, and values Due to the space limitation, we present the details of all used tasks and datasets in Appendix A.2.
2 Implementations
In practice, we use the pretrained models and code in HuggingFace https://github.com/huggingface/transformers Wolf et al. (2019). Specifically, for the understanding tasks, we employ 4 widely used PLMs in our study, i.e., BERT Devlin et al. (2019), ELECTRA Clark et al. (2019b), ALBERT Lan et al. (2019) and RoBERTa Liu et al. (2019). Furthermore, an representative sequence-to-sequence model, BART Lewis et al. (2020), is used for the generation tasks.
We compare our proposed FSAM method with the base optimizer (without using any SAM approach) and vanilla SAM method. Specifically, the Adam Kingma and Ba (2015) is used as the base optimizer to tune our models. The and weight decay of Adam are set as 0.999 and 0.01. SAM and FSAM use the same settings as above. More specially, we grid search for the neighborhood size of SAM and FSAM on {1e-2, 5e-3, 1e-3}. Additionally, for each downstream task, we follow the same hyper-parameter settings from the prior works Lewis et al. (2020); Xu et al. (2021). The detailed hyper-parameters of fine-tuning on these downstream tasks can be seen in Appendix A.3. We report the averaged results over 5 random seeds for NLU tasks, while for NLG tasks, we follow existing works Collins et al. (2005); Ding et al. (2021) and use the Bootstrap test Berg-Kirkpatrick et al. (2012) to calculate the statistical significance.
Main Results
Table 1 shows the results of all understanding tasks. We can observe that SAM achieves better average scores than the base Adam in most scenarios, confirming the effectiveness of SAM in improving generalization Bahri et al. (2022). Moreover, with the help of our Fisher mask strategy, FSAM consistently improves the vanilla SAM by a large margin across all PLMs. Specifically, FSAM yields an improvement of up to 1.98 average score on ELECTRA, 1.01 average score on ALBERT and 1.34 average score on BERT. The average improvement on RoBERTa is slight but also higher than 0.67.
FSAM also works well on the generation tasks.
Prior works Kwon et al. (2021); Bahri et al. (2022), which involve the study of SAM or its variants, usually conduct experiments on the image or text classification tasks, e.g., CIFAR-10 Krizhevsky et al. (2009) and ImageNet Krizhevsky et al. (2012). The effectiveness of optimizer on other types of tasks, e.g, generation tasks in NLP, has not been explored well. Thus far, we evaluate our FSAM on the generation tasks and present the results in Table 2. It can be seen that FSAM can deliver promising results against the vanilla SAM as well. Note that both FSAM and SAM outperform the base Adam optimizer, indicating the applicability of SAM and its variants on generation tasks.
FSAM improves performance on various model sizes.
To investigate whether our FSAM is helpful for various scales of PLMs, we evaluate the performance on smaller PLMs, i.e., BERT-base, RoBERTa-base and BART-base. The results are showed in Table 3 and Table 2, respectively. We can see that FSAM consistently outperforms the vanilla SAM on multiple smaller PLMs, to be specific, the relative improvements of BERT-base and BART-base are up to 0.92 and 0.64 average scores. These results prove that FSAM works well on various model sizes.
Analysis and Discussion
In this section, we examine whether our approach works in more complicated scenarios, and provide a more intuitive comparison between different optimizers towards the generalization. More analysis and results can be found in Appendix.
There are two important hyper-parameters (i.e., and Fi) in our FSAM, where the refers to the sparse ratio and Fi is used to control the update frequency of Fisher mask. Here, we evaluate the performance of FSAM with different and Fi on several downstream tasks to analyze their effects.
Firstly, Figure 1 shows the results based on different . We can observe FSAM outperforms the vanilla SAM and base Adam in most settings, indicating the robustness of FSAM. Specifically, when the sparse ratio is 0.9, FSAM consistently achieves the best performance on both tasks. Secondly, for Fi, we show the performance of FSAM on different Fi in Table 4. Too small Fi (e.g., 10) may lead to the Fisher mask updating too fast, thus affecting the stability of model optimization. Recall that we set and as the default setting.
2 Complementarity with Other Optimizers
As aforementioned, we show the effectiveness of our Fisher mask strategy on SAM optimization. To further prove the universality of our proposed strategy, we examine whether the strategy is complementary with i) more base optimizers and ii) other efficient SAM variants.
To verify i), we use the additional AMSGrad and Adagrad as the base optimizers and evaluate the performance with different strategies, respectively. Table 5 lists the results of RoBERTa-large. It can be seen that FSAM consistently achieves the best performance upon these base optimizers, showing our strategy is not sensitive to the base optimizers.
For ii), we apply our strategy to another two cutting-edge SAM-variant optimizers, i.e., ESAM Du et al. (2021) and GSAM Zhuang et al. (2021). Table 6 shows the results, where F_ESAM and F_GSAM refer to the optimizations using our strategy. When evaluating RoBERTa-large on these tasks, compared to the vanilla ESAM and GSAM, our method can bring a 0.70 average score improvement. This indicates that our Fisher mask strategy is not only beneficial to the vanilla SAM, but also can be applied to other efficient SAM variants.
3 Results in Low-resource Scenarios
Prior works Chen et al. (2021); Bahri et al. (2022) show that SAM helps more when there is less training data. Here, we verify how our Fisher mask strategy affects the effectiveness of SAM in low-resource scenarios. In practice, we follow Bahri et al. (2022) and sub-sample the training splits for several GLUE datasets at rates ranging from 10% to 90%. Notably, due to the space limitation, we only report parts of results on BERT-large and RoBERTa-large in Figure 2.
We can observe consistent gains from both SAM and our FSAM across all sizes of sub-sampled training sets, which confirms the statement in prior work Bahri et al. (2022). Moreover, it can also be seen that our FSAM improves the vanilla SAM by a large margin in low-resource scenarios, especially when there is only 20% training data. More specifically, when fine-tuning the RoBERTa-large on the STS-B dataset, the relative improvements of FSAM are up to 15.0 and 15.1 in terms of accuracy and F1 score, respectively. These results show that our method is more helpful in low-resource scenarios.
4 Does FSAM Bring Better Generalization?
We prove the effectiveness of our FSAM by large-scale experiments as above. Here, to examine whether FSAM indeed brings better generalization, we i) measure the generalization properties (i.e., task generalization) of different optimizations, and ii) visualize the generalization of models via the training loss landscapes.
The common wisdom is that models with better generalization would perform better on out-of-domain data Xu et al. (2021). Thus, to measure the generalization ability of the model quantitatively, we follow the experiments from Xu et al. (2021) and evaluate the performance of various fine-tuned models on out-of-domain data. In practice, we first fine-tune RoBERTa-large on the QNLI task (one of GLUE tasks) and then transfer it to other tasks, i.e., CoLA, MRPC, STS-B and RTE. The results of different optimization strategies are illustrated in Figure 3.
We can observe that FSAM consistently outperforms the base Adam and vanilla SAM on different transferred tasks. To be more specific, compared with vanilla SAM, our FSAM brings a 2.37 relative average improvement score on these tasks, indicating that our method helps more in improving the generalization of model.
Visualization of Landscape.
Here, we visualize the loss landscapes of RoBERTa-base model fine-tuned on CoLA with different optimizers. In practice, we follow Li et al. (2018); Zan et al. (2022) and show the 3D loss surface results in Figure 4 by sampling 2525 points in the range of from random “filter normalized” directions Li et al. (2018). Additionally, following Hao et al. (2019); He et al. (2021), we also plot the 1D loss curve in Figure 5 by linear interpolation between the pretrained model weights before (denoted as ) and after (denoted as ) fine-tuning, i.e., “”, where is a scalar parameter that is ranged from -1 to 1. We can find that the landscape of FSAM is much flatter than both base Adam and SAM, especially in the area of low loss. These results prove that FSAM can smooth the loss landscape and improve the generalization of PLMs effectively.
Conclusion
In this paper, we improve the recently-proposed SAM optimization method with a novel Fisher mask strategy, and propose a new approach FSAM. Different from the vanilla SAM that adds a constant perturbation to all parameters, FSAM uses the Fisher information to calculate the Fisher mask and further obtains the sparse perturbation. Such a method can not only reduce the computation cost of optimization potentially, but also make the optimizer focus on the optimization of the important sharper parameters. Extensive experiments on five PLMs and various language understanding and generation tasks show that our FSAM consistently improves the performance of SAM by a large margin across all PLMs and tasks. Additionally, in-depth analysis and discussion demonstrate the robustness and universality of FSAM on improving the generalization of language models.
Limitations
Indeed, our work has some potential limitations, and we will discuss them in this section. Firstly, we only evaluate the BART on two generation tasks with different optimizers, and prove the effectiveness of our FSAM optimization method. It would be more valuable to consider other sequence-to-sequence PLMs and more generation tasks, e.g., fine-tuning T5 Raffel et al. (2020) on CNN-DM Hermann et al. (2015).
Additionally, as aforementioned in Section 1, we do not achieve the actual sparse training in this work, due to the limitation of the hardware. Specifically, to actually accelerate the unstructured sparsity (fine-grained sparsity), we need to implement the relevant sparse matrix calculation using the CUDA API on the recent NVIDIA Ampere A100 GPUs Choquette et al. (2021) equipped with Sparse Tensor Cores Pool (2020) (Notably, although there is python API (i.e., ASP) provided by NVIDIA for accelerating the unstructured sparsity, it is only applicable to accelerate the model parameter sparsity, but not to the gradient-level sparse acceleration in our FSAM scenario). Unfortunately, it is relatively impracticable for us to do that. However, we still believe that FSAM has great potential to achieve true training acceleration in the future, with the development of hardware for fine-grained sparse operation.
Acknowledgements
We are grateful to the anonymous reviewers and the area chair for their insightful comments and suggestions. This work was supported in part by the National Natural Science Foundation of China under Grants 62141112, 62076186 and 62225113, and in part by the Science and Technology Major Project of Hubei Province (Next-Generation AI Technologies) under Grant 2019AEA170. The numerical calculations in this paper have been done on the supercomputing system in the Supercomputing Center of Wuhan University.
References
Appendix A Appendix
With the term defined before, we have the following inequality
where the is a undetermined scalar.
The is the term to be determined.
where the is a undetermined scalar.
By using the definition of L-smooth and the Lemma 2, 3, 4 and 5, we have
From the Lemma 2, we have the following inequality. Specifically,
From the Lemma 3, we have the following inequality. Specifically,
By taking the expectation, we have the following inequality. Specifically,
From the Lemma 4, we have the following inequality. Specifically,
Set the and set the , we can simplify the inequality
Note that the is bounded, we re-arrange the inequality and achieve
With probability over the choice of training set , we have the following inequality.
The proof is mainly based on PAC-Bayesian Generalization Bound theorem. We start by any prior over parameters with probability for any posterior distribution , we have
Assume , the follows the Chi-square distribution. From the Lemma 1 in Laurent and Massart (2000), we have the following inequality for any positive :
Combine the above function and subtract the same constant on both side, following the assumption in Foret et al. (2020), we finish the proof.
A.2 Details of Tasks and Datasets
As mention in Section 4, we conduct extensive experiments on parts of tasks from GLUE and SuperGLUE. In addition, two widely-used generation tasks are also used in this work. Here, we introduce the descriptions of the used tasks and datasets in detail. Firstly, we present the statistics of all datasets in Table 7. Then, each task is described as:
CoLA. Corpus of Linguistic Acceptability Warstadt et al. (2019) is a binary single-sentence classification task to determine whether a given sentence is linguistically “acceptable”.
MRPC. Microsoft Research Paraphrase Corpus Dolan and Brockett (2005) is a task to predict whether two sentences are semantically equivalent.
STS-B. Semantic Textual Similarity Cer et al. (2017) is a task to predict how similar two sentences are on a 1-5 scale in terms of semantic meaning.
RTE. Recognizing Textual Entailment Giampiccolo et al. (2007), given a premise and a hypothesis, is a task to predict whether the premise entails the hypothesis.
QNLI. Question Natural Language Inference is a binary classification task constructed from SQuAD Rajpurkar et al. (2016), which aims to predict whether a context sentence contains the answer to a question sentence.
CB. CommitmentBank De Marneffe et al. (2019) is a task that can be framed as three-class textual entailment on a corpus of 1,200 naturally occurring discourses.
BoolQ. Boolean Question Clark et al. (2019a) is a question answering task where each sample consists of a short passage and a yes/no question about the passage.
WiC. Word-in-Context Pilehvar and Camacho-Collados (2019) is a word sense disambiguation task that aims to predict whether the word is used with the same sense in sentence pairs.
WSC. Winograd Schema Challenge Levesque et al. (2012) is a co-reference resolution task which aims to determine the correct refer-rent of the pronoun from among the provided choices.
XSUM. The Extreme Summarization dataset Narayan et al. (2018) is one of abstractive Summarization task that aims to convert the given document into a short and adequate summary in the same language.
CoNLL2014. CoNLL2014 Ng et al. (2014) is a popular grammatical error correction task that aims to rewrite the input sentence with grammatical errors into the corresponding correct sentence, where the original and target sentences have the similar sentence lengths.
A.3 Hyper-parameters of Fine-tuning
In this paper, we fine-tune four different large-scale PLMs with our FSAM on the tasks of GLUE and SuperGLUE, including BERT-large (340M) https://huggingface.co/bert-large-cased, ELECTRA-large (340M) https://huggingface.co/google/electra-large-discriminator, ALBERT-xxlarge-v2 (223M) https://huggingface.co/albert-xxlarge-v2 and RoBERTa-large (355M) https://huggingface.co/roberta-large. Additionally, the BART-large (406M) https://huggingface.co/facebook/bart-large is used for the generation tasks. The training epochs/steps, batch size, learning rate and warmup steps are listed in Table 8 and Tabel 10. Notably, the maximum sequence length of language understanding tasks is set as 128/256. For two generation tasks, we empirically set the minimum and maximum length of the XSUM dataset as 10 and 60, and closely follow Chollampatt and Ng (2018) to preprocess the data of CoNLL2014.
A.4 Training Curves
In this sub-section, we visualize the training curves of Adam, SAM and FSAM in detail. Specifically, Figure 6 shows the evaluation metrics v.s. training epochs. The metric curves prove that FSAM boosts the performance effectively during the training, which shows the effectiveness of FSAM.
A.5 More Results
In addition to the results in Table 1 with the momentum as 0.9, we also conduct the experiments with the momentum as 0 (i.e., in Algorithm 1) to evaluate the influence of the adaptive learning rate. In practice, we evaluate the performance on several downstream tasks upon two base optimizers (Adam and AMSGrad) and two PLMs (RoBERTa-large and RoBERTa-base). Table 9 shows the results. When there is no momentum term, FSAM also achieves better performance against the vanilla SAM and base optimizers. These results prove the universality of FSAM in various scenarios.