Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs

Yuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun, Yiwu Yao, Xingjia Han, Jared Tanner, Shiwei Liu, Rongrong Ji

Introduction

Large language models (LLMs) (Zhang et al., 2022a; Touvron et al., 2023a; Brown et al., 2020) have recently emerged as the new favorite in various domains of natural language processing (NLP) (Wei et al., 2022b; a; Bubeck et al., 2023). Nevertheless, LLMs face a significant constraint: their extensive parameterization and computational demands present substantial challenges in terms of storage and deployment. For example, the GPT-175B model (Brown et al., 2020) eats up 320G of memory to load its parameters in FP16 precision, requiring at least five A100-80G GPUs for inference (Frantar & Alistarh, 2023). In response to this issue, there has been a surge of interest in compressing LLMs, as it holds the promise of LLMs while remarkably reducing memory usage and computational costs. To date, the majority of current effort for LLM compression falls into quantization (Yao et al., 2022; Lin et al., 2023; Frantar et al., 2022; Dettmers et al., 2023; 2022; Xiao et al., 2023), which compresses LLMs by diminishing the number of bits employed to represent weights or hidden states.

On the other hand, network pruning (LeCun et al., 1989; Han et al., 2015; Mocanu et al., 2018), a technique that removes superfluous weights to create a sparse and lightweight model, has received relatively little attention (Frantar & Alistarh, 2023; Sun et al., 2023). The plausible reason is that, network pruning usually appreciates at least one, usually many, iterations of fine-tuning or re-training to guarantee top performance (Frankle & Carbin, 2019; Yin et al., 2023). This fine-tuning step would cause a significant amount of compute and memory footprints due to the colossal model size and massive training data of modern LLMs, which even unnerves large corporations, let alone individual researchers.

Two previous arts have explored the possibility to scale pruning to billion-level LLMs without any fine-tuning. SparseGPT (Frantar & Alistarh, 2023) formulates LLM pruning as a layer-wise weight reconstruction problem, where the target falls into mitigating the output discrepancy, w.r.t., reconstruction error, between dense and sparse LLMs. To solve the row-Hessian challenge, i.e., the need for calculating the expensive inversion of a huge matrix for each row individually, SparseGPT iteratively applies OBS (Hassibi et al., 1993) to individually prune and updates weights in a column-wise manner, ultimately reaching the same optimal solution as applying the closed-form regression reconstruction. Wanda (Sun et al., 2023) proposes a new pruning metric that takes both weight magnitude and their corresponding input activations into consideration, performing on part with SparseGPT without the need for the expensive second-order information. The intuition behind Wanda lies in the existence of emergent outlier feature dimensions in large-scale LLMs which are significantly larger than typical features and meanwhile are essential for the optimal performance of LLMs (Dettmers et al., 2022). While these two approaches enable LLM pruning without performing fine-tuning, their performance is still far from satisfactory, e.g., starting to lose performance at 20% sparsity with LLaMA-30B. Therefore, it is imperative to enable fine-tuning for sparse LLMs to fully unlock the potential of sparsity to escalate the affordability of LLMs.

In a parallel vein, Dynamic Sparse Training (DST), as outlined in previous research (Mocanu et al., 2018; Liu et al., 2019; Evci et al., 2020), has garnered considerable attention recently due to its significant saving potentials in the context of neural network training. Instead of training an entire network, DST selectively updates and maintains a subset of the network throughout the training process, while allowing the sparse network topology to dynamically evolve via a weight operation (Mocanu et al., 2018). Given its demonstrated efficacy in achieving efficient training, DST seems to be a promising candidate for efficient LLMs fine-tuning. However, it is essential to note that DST intrinsically requires the training of subnetworks via backpropagation, and the effectiveness of mask adaptation highly relies on a sufficient number of weight updates (Liu et al., 2021). Moreover, prior studies have indicated its failure when employed for fine-tuning small-scale BERT-level language models (Liu et al., 2023).

Fortunately, it is noteworthy that the pruning-and-growing step employed in DST solely stands as a training-free methodology, enabling sparse mask adaptation based on certain weight status, e.g., magnitude (Mocanu et al., 2018). This offers an alternative perspective for addressing the aforementioned challenge: While fine-tuning sparse LLMs through backpropagation can result in substantial computational overhead, we can explore the possibility of iteratively updating sparse mask in a training-free fashion as a viable alternative. Based on this intuition, we introduce a training-free fine-tuning approach – Dynamic Sparse No Training (DS\faBanT). This approach empowers the further refinement of sparse LLMs without any weight updates. To facilitate mask adaptation in favor of the sparse reconstruction problem, we propose new criteria for mask pruning and growing, by considering both the expectation and variance of the reconstruction error reduction when recovering a specific weight. It is worth emphasizing that the DS\faBanT functions independently of the need for computationally intensive operations, such as gradient or Hessian matrices. Instead, it exclusively relies on a singular matrix multiplication operation to assess the reconstruction error.

We conduct comprehensive experiments to evaluate the effectiveness of DS\faBanT with a variety of LLMs, including LLaMa-V1 (Touvron et al., 2023a) and LLaMa-V2 (Zhang et al., 2022a), Vicuna (Chiang et al., 2023), and OPT families (Zhang et al., 2022a), from 7 billion to 70 billion parameters. Our results demonstrate that DS\faBanT consistently improves the performance of sparse LLMs by a good margin, especially at high sparsity levels >> 50%. For instance, DS\faBanT is able to improve the performance over Magnitude pruning, SparseGPT, and Wanda by 1.1e6, 4.31, and 1.87 perplexity with OPT-13B on WikiText-2 at 60% sparsity only using 7.3s on a single NVIDIA A100 GPU. Our work provides fresh insights in efficient sparse LLM fine-tune without weight updates and we hope to encourage more research in exploring benefits of sparsity in LLMs.

Related Work

Network Sparsification. The process of eliminating redundant weights, known as network sparsification or network pruning, has served as a practical strategy to diminish the complexity of deep neural networks over the past decades (LeCun et al., 1989; Han et al., 2015). Despite the substantial body of literature, network pruning can be roughly classified based on the granularity of sparsity and the dependency of the pre-trained dense models. I. Granularity of Sparsity: The granularity of sparsity varies from coarse grains to fine grains. The coarse-grained granularity can be a group of weights (Gray et al., 2017; Ding et al., 2017), a complete neuron (Jiang et al., 2018); a filters/channels (Li et al., 2017), or an attention head (Voita et al., 2019), etc. On the other hand, fine-grained granularity eliminates the least important weights based on the selected criteria, regardless of where they are (Gale et al., 2019). The advantage of coarse-grained sparsity is its pronounced acceleration effect, which yet typically suffers from larger performance loss. Fine-grained sparsity enjoys performance superiority compared to other more structured forms of sparsity but receives limited support in common hardware. Nonetheless, recent advancements of dedicated fine-grained sparse patterns, such as N:M sparsity (Zhou et al., 2021; Zhang et al., 2022b), can be effectively accelerated. As such, this paper focuses on fine-grained network pruning. II. Dependency of Pre-trained Networks: In parallel, sparsification techniques can be grouped into dense-to-sparse, and sparse-to-sparse methods based on the necessity of an over-parameterized dense network. The former entails embarking from a pre-trained dense model and discovering a sparse network (Han et al., 2015; Wen et al., 2016; Molchanov et al., 2017; Gale et al., 2019; Kurtic et al., 2022), usually followed by a retraining process to recover the optimal accuracy. On the other hand, sparse-to-sparse methods aim to train sparse neural networks from scratch, omitting any preliminary steps involving dense pre-training (Mocanu et al., 2018; Lee et al., 2019; Evci et al., 2020; Wang et al., 2020; Liu et al., 2021). Among them, Dynamic Sparse Training (DST) (Mocanu et al., 2018; Evci et al., 2020; Liu et al., 2021) stands out and receives upsurging interest due to its promise in saving both training and inference phases. In contrast to the conventional practices of pre-training followed by pruning, DST distinguishes itself by commencing with a randomly initialized sparse neural network. During a single training run, it dynamically adjusts the sparse network topology by such as pruning-and-growing, without the need for pre-training, while maintaining moderate training costs by, for example, keeping the similar sparsity ratios across all varying masks (Mostafa & Wang, 2019; Dettmers & Zettlemoyer, 2019; Yuan et al., 2021; Jayakumar et al., 2020).

While the crux of this paper focuses on the first category, i.e., pruning a pre-trained LLM model, our proposed method is mainly inspired by the pruning-and-growing utilized in DST to iteratively refine the binary masks in a training-free manner, even though we do not conduct weight training as such. Another line of research, akin to our approach, demonstrates the existence of “supermasks” within randomly initialized network (Zhou et al., 2019; Ramanujan et al., 2020; Huang et al., 2022) or pre-trained networks (Mallya et al., 2018; Wortsman et al., 2020; Zhang et al., 2023), exhibiting the capacity to achieve commendable performance solely by seeking binary masks. However, it is imperative to note that these methods heavily rely on backpropagation, which is ill-suited for LLMs.

Pruning of LLMs. Compared to the well-established promise of pruning in pre-LLM small-scale models, the advancement of pruning in the context of LLMs appears to exhibit relatively modest progress. Firstly, traditional pruning generally requires at least one iteration of re-training to recover performance. Considering the substantial model size and massive datasets associated with LLMs, the prospect of conducting such resource-intensive re-training becomes a formidable challenge. To mitigate the above challenge, researchers have introduced pruning algorithms specifically devised for LLMs compression. Ma et al. (2023) explored structured sparse LLM by applying Taylor pruning (Molchanov et al., 2017) to remove entire weight rows, followed by the parameter efficient fine-tuning (PEFT) technique (Hu et al., 2021) fine-tuning. However, the fine-tuning phase still demands a considerable amount of data while the performance suffers a significant degradation, attributed primarily to the coarse-grained level of sparsity. Recent research endeavours have evolved towards the direction of unstructured pruning in one-shot without fine-tuning, demonstrating significant progresses. SparseGPT (Frantar & Alistarh, 2023) incorporates the Hessian inverse for pruning and subsequent residual weight updates, whereas Wanda (Sun et al., 2023) directly arrives at a sparse LLM model by a criterion depicted by the multiplication of the absolute values of weights and their activations with the aim to preserve outliers (Dettmers et al., 2022) emerged in LLMs. DS\faBanT serves as an orthogonal perspective and can be organically integrated on top of them.

Dynamic Sparse No Training – DS\faBanT

Dynamic Sparse No Training. The problem defined in Eq. (1) can be addressed from two complementary perspectives. Firstly, it can be resolved through the initialization of sparse networks i.e., devising criteria to prune weights that exhibit minimal impact on model output. For instance, SparseGPT (Frantar & Alistarh, 2023) employs second-order Hessian inverses, while Wanda (Sun et al., 2023) considers products of weight and activation norm as the guide for weight removal. Secondly, for the obtained sparse networks, the remaining weights can be naturally fine-tuned to further compensate for the reconstruction error (Han et al., 2015). Unfortunately, this requires substantial training resources, which is not practical given the large volumes of LLMs. Therefore, SparseGPT adjusts the remaining weights via an iterative OBS update (Hassibi & Stork, 1992), which as a consequence remarkably reduces the computing demands.

In this work, our focus is on the second part, i.e., how to efficiently reduce the reconstruction error of a given pruned sparse network to its dense counterpart? Instead of fully fine-tuning (Han et al., 2015) or partially updating the pruned LLMs (Frantar & Alistarh, 2023) to recover performance, we introduce an ultra-efficient yet effective alternative to refine the sparse mask after pruning based on their contribution to the reconstruction error. Our approach is inspired by the pruning-and-growing operation used in Dynamic Sparse Training (Mocanu et al., 2018; Evci et al., 2020). DST incorporates the processes of weight pruning and weight growing within the framework of sparse network training, contributing to the discovery of improved sparse topologies. Note that this pruning-and-growing operation solely serves as a training-free approach that is able to adapt sparse masks towards a desirable perspective, e.g., loss minimization. Based on this insight, we propose DS\faBanT, a training-free fine-tuning method for sparse LLMs that strips weights updating in DST and keeps the pruning-and-growing by converting the optimization objective to the reconstruction error of each weight row. We isolate pruning-and-growing from network training, and formulate it as an iterative approach to progressively optimize sparse masks towards the desirable ones achieving minimal reconstruction error represented by Eq. (1).

Specifically, DS\faBanT starts with a sparse LLM which can be pruned by any existing criteria (Jaiswal et al., 2023; Sun et al., 2023; Frantar & Alistarh, 2023). Then, it performs iterative weight growing and pruning by looking at the reconstruction error as defined in Eq. (1), with especially-designed criteria to decrease the output discrepancy between sparse LLMs and their dense counterparts. The framework of DS\faBanT is illustrated in Figure 2 and its main parts are detailedly described below.

Growing Criterion. As each output neuron is computed independently, we use one weight row Wr\mathbf{W}_{r} and the corresponding mask Mr\mathbf{M}_{r} for illustration. Given sparse weight row Mr⊙Wr\mathbf{M}_{r}\odot\mathbf{W}_{r}, we attempt to revive pruned weight that leads to the most decrease on Δr\Delta_{r} across different input activations. Therefore, our growing criterion considers both the expectation and variance of the reconstruction error change when recovering a weight back. In particular, the index ii of the revived weights is derived as follows:

Pruning Criterion. After choosing revived weights, we need to select another weight for pruning in order to maintain a fixed sparsity rate. However, the circumstances here are distinct: if we prune weights based on the impact of reconstruction error change as per Eq. (2), there is a risk of removing weights that significantly influence the output. This concern becomes especially critical when pruning LLMs due to the presence of emergent large magnitude features within them (Dettmers et al., 2022; Wei et al., 2022a; Schaeffer et al., 2023). To alleviate this, we utilize a transformed version of the Wanda metric (Sun et al., 2023). In addition to its standard criterion for pruning weights, we mandate that the selected weights should also contribute positively towards the reduction of reconstruction error when being pruned. This helps in preserving critical weights from removal without compromising the stable decrease of reconstruction error during the training-free fine-tuning process. Therefore, the pruning index jj is obtained as follows:

Workflow. Given the criteria depicted above, the workflow of DS\faBanT is outlined in Algorithm 1. In particular, it iteratively performs weight growing and pruning with respect to Eq. (2) and Eq. (3), with the reconstruction error updated until it reaches a pre-defined threshold. Meanwhile, we set a maximum pruning-and-growing cycle TT to prevent certain rows from being unable to reach the settled threshold ϵ\epsilon.

Remark. It’s noteworthy that Algorithm,1 outlines the processing of each row in a sequential manner, primarily for the sake of simplicity. However, it’s imperative to acknowledge that each row can, in fact, undergo parallel processing by employing a binary indicator to assess whether a particular row has satisfied the termination condition. Furthermore, the DS\faBanT process eliminates the necessity for resource-intensive procedures such as backpropagation or the computation of gradient and Hessian matrices. Instead, it relies solely on several matrix multiplications to calculate the reconstruction error, a task that can be executed efficiently on GPUs. Subsequently, during each iteration of the DS\faBanT process, the only operation is to update the reconstruction error through straightforward addition and subtraction operations during the pruning-and-growing process. This approach effectively circumvents the introduction of additional algorithmic complexity. In summary, DS\faBanT preserves the simplicity associated with pruning LLMs, akin to the approaches employed in Wanda and Magnitude pruning. It’s important to note that while we share the same optimization objective with SparseGPT, our approach adopts a significantly more efficient pruning-and-growing operation to minimize the reconstruction error. As demonstrated in the next section, this efficiency operation can further improve the performance of SparseGPT.

Experimental Results

Implementation details. The implementation details of our proposed DS\faBanT are presented as follows, mostly conforming to the existing setups (Frantar & Alistarh, 2023; Sun et al., 2023). In context to pruning configuration, we adhere to SparseGPT (Frantar & Alistarh, 2023), where a uniform sparsity is imposed for all layers with the first embedding layer and the final classification head skipped. Meanwhile, the calibration data consists of 128 segments, each with 2048 tokens. These segments are randomly selected from the first shard of the C4 dataset (Raffel et al., 2020). For the hyper-parameter settings, we set the maximum cycle T=50T=50 and the update threshold ϵ=0.1\epsilon=0.1 in all experiments. We implement DS\faBanT in PyTorch (Paszke et al., 2019) and use the HuggingFace Transformers library (Wolf et al., 2019) for handling models and datasets. All pruning experiments are conducted on NVIDIA A100 GPUs with 80GB of memory.

Baselines. We principally work with the LLaMA-V1 (Touvron et al., 2023a), LLaMA-V2 (Touvron et al., 2023b), Vicuna (Chiang et al., 2023), and OPT families (Zhang et al., 2022a), from 7 billion to 70 billion parameters, which are among the most powerful and open-source Large Language Models (LLMs) in the field today. We run DS\faBanT on sparse LLMs pruned by various methods including (1) Magnitude-based pruning (Han et al., 2015) that discards weights based on their magnitudes. (2) SparseGPT (Frantar & Alistarh, 2023) that utilizes second-order Hessian inverses to ascertain unimportant weights. (3) Wanda (Sun et al., 2023) that removes weights with the smallest magnitudes multiplied by the corresponding input activation norms.

Evaluation. In accordance with prior studies (Frantar et al., 2022; Dettmers et al., 2023; Yao et al., 2022; Frantar & Alistarh, 2023), we assess the performance of pruned models by calculating the perplexity of language generation experiments on separate validation sets derived from WikiText2 (Merity et al., 2016). While perplexity has served as a stable and robust indicator of the generative performance of models (Dettmers & Zettlemoyer, 2023), we also examined the zero-shot capabilities of pruned models. In detail, we report the accuracy in six zero-shot tasks including PIQA (Bisk et al., 2020), StoryCloze (Mostafazadeh et al., 2017), ARC Easy and Challenge (Clark et al., 2018), HellaSwag (Zellers et al., 2019) and OpenBookQA (Mihaylov et al., 2018). We implement the lm-eval-harness (Gao et al., 2021) for the execution of all zero-shot tasks, with the report including both the accuracy results on each benchmark and overall average accuracy.

2 Language Modeling

Quantitative results. The results for fine-tuning sparse LLM models at a uniform sparsity rate of 60% are presented in Table 1. Irrespective of the datasets used for evaluation, DS\faBanT consistently delivers performance improvement for sparse LLMs with their original sizes varying from 7B to 70B. For instance, when pruning LLaMA-V1 with 7B parameters, DS\faBanT is able to enhance the performance of Magnitude (Jaiswal et al., 2023), SparseGPT (Frantar & Alistarh, 2023), and Wanda (Sun et al., 2023) by 4.94e2, 0.76, and 0.47 perplexity on the Wikitext-2 validation sets, respectively. It is worth noting that, without any weight updating, DS\faBanT consistently demonstrates better performance than SparseGPT, which requires expensive second-order Hessian inverses to update the sparse model. For larger models, the efficacy of DS\faBanT is still hold with performance gain from 13.35 to 6.77 perplexity when fine-tuning sparse LLaMA-V2-70B obtained by magnitude pruning (Han et al., 2015). These findings suggest DS\faBanT’s versatility, being adaptable to boost the performance of sparse LLMs with different parameter budgets.

Varying Sparsity Rates. We further investigate the efficacy of DS\faBanT when fine-tuning sparse LLMs with varying pruning rates. Table 2 shows that DS\faBanT offers effective performance enhancement across various pruning methods at different sparsity levels. Particularly, this improvement becomes increasingly evident as the sparsity level grows.

Computing efficiency. We further demonstrate the computing efficiency of DS\faBanT. Following Wanda (Sun et al., 2023), we only report the total pruning time and exclude the forward pass process shared by all methods. Table 5 compares the quantitative wall-clock overhead evaluated on NVIDIA A100 GPUs. It is indeed encouraging to observe that, as a fine-tuning approach, DS\faBanT maintains a comparable computing time to Wanda, while demonstrating significantly higher efficiency compared to SparseGPT.

Comparison with LoRA Fine-tuning. To further demonstrate the ultra efficiency of DS\faBanT in terms of fine-tuning, we also compare DS\faBanT with parameter efficient fine-tuning (PEFT) method LoRA (Hu et al., 2021). Table 5 presents a comparison of the time and performance of both methods in fine-tuning sparse LLaMA-7B. LoRA leverages the complete C4 dataset for a 5-hour fine-tuning and achieved a perplexity of 6.84. In stark contrast, DS\faBanT only requires a brief duration of 4.3 s and 128 samples to deliver a comparable performance, 7.12 perplexity. Taking into consideration the additional network parameter burden incorporated by LoRA, the efficiency and practicality of DS\faBanT is hold.

N:M Fine-grained Sparsity. Compared with unstructured sparsity, N:M fine-grained sparsity offers more practical speedup on the NVIDIA Ampere sparse tensor core (Nvidia, 2020). Thus, we also evaluate the effectiveness of DS\faBanT on N:M fine-grained sparsity. Given the unique pattern of N:M sparsity that stipulates N non-zero components within M consecutive weight block, our implementation of DS\faBanT involves a restriction on the position of pruning-and-growing weights. In particular, we select the pruned weight within the same block as the revived weight, thus the N:M characteristic is still maintained after fine-tuning. Table 5 lists the results for pruning LLaMA-V1 model family at 2:4 and 4:8 sparse patterns. Interestingly, even with the aforementioned extra restriction, DS\faBanT can achieve more significant performance improvement compared to previous methods. For instance, when pruning LLaMA-V1 with 7B parameters, DS\faBanT archives a perplexity of 10.89, enhancing Wanda (11.53) by a noticeable 0.64 ppl. Similar findings can be concluded when it comes to other models and sparse patterns. These results highlight the effectiveness of DS\faBanT in boosting the performance of sparse LLMs, even with more complex sparsity constraints.

3 Zero-shot Tasks

Following (Frantar & Alistarh, 2023; Sun et al., 2023), we also provided the accuracy performance of the LLaMA-V1 model family pruned at 50% sparsity rate on seven downstream zero-shot tasks. Averaging the accuracy over all tasks suggests DS\faBanT’s efficacy for enhancing sparse LLMs of any size. Particularly, DS\faBanT improves the average accuracy of SparseGPT by 1.6% when pruning LLaMA-V1-7B (52.7% for DS\faBanT and 51.1% for SparseGPT). For task-wise performance, DS\faBanT is beneficial on all tasks, while there is not a fixed superiority for fine-tuning models obtained by different pruning methods. This phenomenon may evidence the reported relatively noisy evaluation results from these zero-shot experiments (Dettmers et al., 2022). However, the advantages of consistent performance improvement and efficiency of DS\faBanT for zero-shot tasks are obvious.

4 Performance Analysis

Next, we investigate the influence of the components within DS\faBanT, unfolds as its update schedule, pruning-and-growing criteria, and robustness to calibration samples. All experimental setups are based on the LLaMA-7B model pruned by the Wanda metric (Sun et al., 2023) with 60% sparsity.

Update schedule. In Figure 3 (left), we examine the performance of DS\faBanT under different hyper-parameter setting for the update schedule, including the maximum cycle CC and stop threshold ϵ\epsilon. The best performance is obtained with 50 cycles and 0.1 updating threshold. To analyze, smaller CC and larger ϵ\epsilon both lead to an insufficient procedure for the decrease in reconstruction error. In contrast, running DS\faBanT without termination conditions also resulted in poor performance, most likely due to over-fitting of calibration data.

Robustness to calibration samples. In Figure 3 (right), we show the performance of pruning methods with varying numbers of sampled sequences for calibration. As can be observed, SparseGPT suffers serious performance degradation when calibration samples are limited, mostly due to the difficulty in estimating Hessian inverses in such cases. Fortunately, DS\faBanT consistently the performance of SparseGPT, even if only very few samples are given. These results further highlight the robustness of DS\faBanT for mitigating the reconstruction error.

Pruning-and-growing criteria. We further investigate the influence on criteria for prune and grow in Table 7. Note that when we transfer Eq. (2) to the prune criteria, the election of extreme values is also correspondingly reversed. As for the prune criterion, it can be seen that pruning weights that could bring the most reduction in reconstruction error actually led to a significant performance decrease. This indicates that while pursuing the reduction of reconstruction error, it is also essential to keep weights that exhibit an extremely large influence on the output, e.g., weights within outlier channel. On the other hand, our proposed criteria based on the expectation and variance of the reconstruction error reduction achieved the best results among all growing criteria.

Conclusion

In this work, we introduce DS\faBanT, a training-free fine-tuning approach that enhances the performance of sparse LLMs without the expensive backpropagation or any weight updates. Taking inspiration from the success of sparse training in the pre-LLM pruning age, DS\faBanT adapts iterative weights growing and pruning in a sparse LLM, with a transferred target for minimizing the reconstruction error between dense and sparse LLMs outputs. To furnish guidance in the selection of weights to be pruned and grown, we introduce novel criteria that take into account the expectation and variance of the reconstruction error reduction by growing each weight concerning different inputs. Extensive experiments on pruning representative LLMs across various language benchmarks demonstrate the efficiency and effectiveness of DS\faBanT in boosting the performance of sparse LLMs.

Acknowledgement

This work was supported by National Key R&D Program of China (No.2022ZD0118202), the National Science Fund for Distinguished Young Scholars (No.62025603), the National Natural Science Foundation of China (No. U21B2037, No. U22B2051, No. 62176222, No. 62176223, No. 62176226, No. 62072386, No. 62072387, No. 62072389, No. 62002305 and No. 62272401), and the Natural Science Foundation of Fujian Province of China (No.2021J01002, No.2022J06001).

References