LLM-Pruner: On the Structural Pruning of Large Language Models
Xinyin Ma, Gongfan Fang, Xinchao Wang
Introduction
Recently, Large Language Models (LLMs) OpenAI (2023); Touvron et al. (2023); Thoppilan et al. (2022); Scao et al. (2022); Xue et al. (2020); Chiang et al. (2023); Zeng et al. (2022) have demonstrated remarkable proficiency in language understanding and generation. With the increase in model size, they are better equipped to handle complex tasks Brown et al. (2020); Chowdhery et al. (2022); Wei et al. (2022b); Wu et al. (2020) and even exhibit emergent abilities Wei et al. (2022a). However, notwithstanding their impressive performance, LLMs pose challenges in deployment and inference. Their extensive scale engenders substantial computational demands, and the multitude of parameters involved can induce long latencies and other related issues. Several techniques are proposed to solve these problems, like model pruning Wang et al. (2019b); Xia et al. (2022); Zafrir et al. (2021); Kurtic et al. (2022), knowledge distillation Sun et al. (2019); Pan et al. (2020b); Sun et al. (2020a),quantization Bai et al. (2020); Frantar et al. (2022) within the context of pre-trained language model (PLM).
While previous methods have effectively maintained model performance amidst parameter reduction, they primarily target compression within specialized domains or for designated tasks in the context of task-specific compression. For instance, a PLM is fine-tuned on a particular dataset, such as one of the classification tasks in the GLUE benchmark Wang et al. (2018), after which these models are distilled into a smaller classification model Sun et al. (2019); Hou et al. (2020). Although this paradigm could potentially be employed for LLM compression, it compromises the LLM’s capacity as a versatile task solver, rendering it suited to a single task exclusively.
Thus, we strive to compress the LLM in a new setting: to reduce the LLM’s size while preserving its diverse capabilities as general-purpose task solvers, as depicted in Figure 1. This introduces the task-agnostic compression of LLMs, which presents two key challenges:
The size of the training corpus of the LLM is enormous. Previous compression methods heavily depend on the training corpus. The LLM has escalated the corpus scale to 1 trillion tokens or more Hoffmann et al. (2022); Touvron et al. (2023). The extensive storage needs and protracted transmission times make the dataset difficult to acquire. Furthermore, if the dataset is proprietary, acquisition of the training corpus verges on impossibility, a situation encountered in Zeng et al. (2022); OpenAI (2023).
The unacceptably long duration for the post-training of the pruned LLM. Existing methods require a substantial amount of time for post-training the smaller model Wang et al. (2020); Liang et al. (2023). For instance, the general distillation in TinyBERT takes around 14 GPU days Jiao et al. (2020). Even post-training a task-specific compressed model of BERT demands around 33 hours Xia et al. (2022); Kwon et al. (2022). As the size of both the model and corpus for LLMs increases rapidly, this step will invariably consume an even more extensive time.
To tackle the aforementioned challenges associated with the task-agnostic compression of LLMs, we introduce a novel approach called LLM-Pruner. Since our goal is to compress LLMs with reduced data dependency and expedited post-training, how to prune model with the minimal disruption to the origin is crucial. To accomplish this, we propose a dependency detection algorithm that identifies all the dependent structures within the model. Once the coupled structure is identified, we employ an efficient importance estimation strategy to select the optimal group for pruning under the task-agnostic setting, where the first-order information and an approximated hessian information is taken into account. Finally, a rapid recovery stage is executed to post-train the pruned model with limited data.
In this paper, we propose a novel framework, LLM-Pruner, for the task-agnostic compression of the large language model. To the best of our knowledge, LLM-Pruner is the first framework designed for structured pruning of LLMs. We conclude the advantages of the LLM-Pruner as (i) Task-agnostic compression, where the compressed language model retains its ability to serve as a multi-task solver. (ii) Reduced demand for the original training corpus, where only 50k publicly available samples are needed for compression, significantly reducing the budget for acquiring the training data (iii) Quick compression, where the compression process ends up in three hours. (iv) An automatic structural pruning framework, where all the dependent structures are grouped without the need for any manual design. To evaluate the effectiveness of LLM-Pruner, we conduct extensive experiments on three large language models: LLaMA-7B, Vicuna-7B, and ChatGLM-6B. The compressed models are evaluated using nine datasets to assess both the generation quality and the zero-shot classification performance of the pruned models. The experimental results demonstrate that even with the removal of 20% of the parameters, the pruned model maintains 94.97% of the performance of the original model.
Related Work
Language models Devlin et al. (2018); Liu et al. (2019); Lewis et al. (2019) have gained much attention and increase the need to reduce the size of parameters and reduce the latency Lan et al. (2019); Sun et al. (2020b). To compress the language model, previous works can be divided into several categories: network pruning Kurtic et al. (2022); Xu et al. (2021); Liu et al. (2021); Guo et al. (2019), knowledge distillation Sun et al. (2019, 2020a); Pan et al. (2020a), quantization Yao et al. (2022); Bai et al. (2020); Zafrir et al. (2019) and other techniques, like early exit Xin et al. (2020) or dynamic token reduction Ye et al. (2021). We focus on the pruning of the language models, especially structural pruning Li et al. (2016). Structural pruning removes the entire filter from the neural network, which is more hardware friendly. There are several ways to remove the structure, such as l1-dependent pruning Han et al. (2015); Zafrir et al. (2021), first-order importance estimation Hou et al. (2020), hessian-based estimation Kurtic et al. (2022); Wang et al. (2019a) or the optimal brain surgeon LeCun et al. (1989); Kurtic et al. (2022). As for the pruning unit in structural pruning, some works adopt the entire layer Fan et al. (2019) as the minimal unit, and others take the multi-head attention Voita et al. (2019) or the feed-forward layers Hou et al. (2020); McCarley et al. (2019) as the basic structure to prune. CoFi Xia et al. (2022) studies the pruning unit in different granularity.
Efficient and Low Resource Compression.
With the growing size of models, there is an increasing demand for efficient LLM compression and compression is independent of the original training data. As for the efficient compression, Kwon et al. (2022) accelerate the post-training by defining the reconstruction error as a linear least squares problem. Frantar et al. (2022); Frantar and Alistarh (2023) propose the layer-wise optimal brain surgeon. As for the constraint of availability of the training corpus, data-free pruning Srinivas and Babu (2015); Yvinec et al. (2022) come up with several strategies to prune the model by measuring neurons’ similarity. Besides, Ma et al. (2022, 2020); Rashid et al. (2020) proposes methods that distill the model without reliance on the training corpus of the model. However, those methods are too time-consuming, involving synthesizing samples by backpropagating the pre-trained language models.
Methods
In this section, we provide a detailed explanation of LLM-Pruner. Following the conventional model compression pipelineKwon et al. (2022), LLM-Pruner consists of three steps: (1) Discovery Stage (Section 3.1). This step focuses on identifying groups of interdependent structures within LLMs. (2) Estimation Stage (Section 3.2). Once the coupled structures are grouped, the second step entails estimating the contribution of each group to the overall performance of the model and deciding which group to be pruned. (3) Recover Stage (Section 3.3). This step involves fast post-training that alleviates potential performance degradation caused by the removal of structures.
In light of the limited availability of data for post-training, it becomes imperative to prioritize the removal of structures with minimal damage when compressing the model. This underscores the dependency-based structural pruning, which ensures coupled structures are pruned in unison. We provide an experiment in Section 4.3 to show the importance of dependency-based structural pruning when compressing the large language model.
Similar to Fang et al. (2023), the pruning begins by building the dependency for LLMs. Assume and are two neurons in the model, and represents all the neurons that point towards or point from . The dependency between structures can be defined as:
where represents the in-degree of neuron . Noting that this dependency is directional, we can therefore correspondingly obtain another dependency:
where represents the out-degree of neuron . The principle of dependency here is, if a current neuron (e.g., ) depends solely on another neuron (e.g., ), and the neuron is subjected to pruning, it follows that the neuron must also undergo pruning.
Trigger the Dependency Graph.
By having the definition of dependency, the coupled structures in the LLM can be analyzed automatically. Considering any neuron within the LLM as the initial trigger, it possesses the capability to activate neurons that depend on it. Subsequently, these newly triggered neurons can serve as the subsequent triggers to identify the dependency and activate their respective dependent neurons. This iterative process continues until no new neurons are detected. Those neurons then form a group for further pruning. Taking LLaMA as an example, by searching over all the neurons as the initial trigger, we can locate all the coupled structures, as shown in Figure2.
Given the diversity in the structure of different LLMs, manual analysis and removal of coupled structures in each LLM could be extremely time-consuming. However, by employing LLM-Pruner, all coupled structures can be automatically identified and extracted.
2 Grouped Importance Estimation of Coupled Structure
Till now, all coupled structures within the model are grouped. Weights within the same group should be pruned simultaneously, as partial pruning not only increases parameter size but also introduces misaligned intermediate representations. Therefore, we estimate the importance of the group as a whole, as opposed to evaluating the importance of modules. Given the limited access to the training dataset, we explore the use of public datasets or manually created samples as alternative resources. Although the domains of these datasets may not perfectly align with the training set, they still provide valuable information for assessing the importance.
Suppose that given a dataset , where N is the number of samples. In our experiments, we set N equal to 10 and we use some public datasets as the source of . A group (as previously defined as a set of coupled structures) can be defined as , where M is the number of coupled structures in one group and is the weight for each structure. While pruning, our goal is to remove the group that has the least impact on the model’s prediction, which can be indicated by the deviation in the loss. Specially, to estimate the importance of , the change in loss can be formulated as LeCun et al. (1989):
where is the hessian matrix. Here, represents the next-token prediction loss. The first term is typically neglected in prior work LeCun et al. (1989); Wang et al. (2019a); Frantar and Alistarh (2023), as the model has already converged on the training dataset, where . However, since here is not extracted from the original training data, which means that . This presents a desirable property for determining the importance of by the gradient term under LLMs, since computation of the second term, the Hessian matrix, on the LLM is impractical with complexity.
Element-wise Importance.
The above can be considered as an estimate for the weight . We can derive another measure of importance at a finer granularity, where each parameter within is assessed for its significance:
Here, represents the k-th parameter in . The diagonal of the hessian can be approximated by the Fisher information matrix, and the importance can be defined as:
Group Importance.
By utilizing either or , we estimate the importance at the granularity of either a parameter or a weight. Remembering that our goal is to estimate the importance of , we aggregate the importance scores in four ways: (i) Summation: or , (ii) Production: or , (iii) Max: or ; (iv) Last-Only: Since deleting the last executing structure in a dependency group is equivalent to erasing all the computed results within that group, we assign the importance of the last executing structure as the importance of the group: or , where is the last structure. After assessing the importance of each group, we rank the importance of each group and prune the groups with lower importance based on a predefined pruning ratio.
3 Fast Recovery with Low-rank Approximation
where is the bias in the dense layer. Only training and reduces the overall training complexity, reducing the need for large-scale training data. Besides, the extra parameters and can be reparameterized into , which would not cause extra parameters in the final compressed model.
Experiments
To showcase the effectiveness and versatility of LLM-Pruner, we test it over three open-source large language models with two kinds of structure: LLaMA-7B Touvron et al. (2023), Vicuna-7B Chiang et al. (2023) https://huggingface.co/lmsys/vicuna-7b-delta-v0 and ChatGLM-6B Zeng et al. (2022).
Evaluation and Datasets.
To assess the performance of the model in the task-agnostic setting, we follow LLaMa’s evaluation to perform zero-shot task classification on common sense reasoning datasets: BoolQ Clark et al. (2019), PIQA Bisk et al. (2020), HellaSwag Zellers et al. (2019), WinoGrande Sakaguchi et al. (2019), ARC-easy Clark et al. (2018), ARC-challenge Clark et al. (2018) and OpenbookQA Mihaylov et al. (2018). Follow Gao et al. (2021), the model ranks the choices in the multiple choice tasks or generates the answer in the open-ended generation https://github.com/EleutherAI/lm-evaluation-harness. Additionally, we complement our evaluation with a zero-shot perplexity (PPL) analysis on WikiText2 Merity et al. (2016) and PTB Marcus et al. (1993).
Implementation Details.
In the model pruning process, we use 10 randomly selected samples from Bookcorpus Zhu et al. (2015), each truncated to a sequence length of 128, as the calibration samples for establishing dependency and calculating the gradient for both LLaMA and Vicuna. For ChatGLM, we select 10 random samples from DailyDialog Li et al. (2017). During the recovery phase, we utilize the cleaned version of Alpaca Taori et al. (2023), which comprises approximately 50k samples. Remarkably, tuning these samples requires merely 3 hours on a single GPU with only 2 epochs. More hyper-parameters of pruning and training can be found in Appendix B.
Statistics of the Compressed Model.
Table 3 presents the statistic of the 7B models that are used in our experiments: the parameter count, MACs, memory requirements and latency for running each model. The statistical evaluation is conducted using the inference mode, where the model is fed a sentence consisting of 64 tokens. The latency is tested under the test set of WikiText2 on a single A5000. Here, the ‘Block’ strategy implies that the pruned unit in the model consists of Group Type A and Group Type B as illustrated in Figure 2, whereas ‘Channel’ indicates that the unit to be pruned is Group Type C. We delve into an analysis of these two choices in Section 4.2(Channel Strategy vs. Block Strategy). The pruning ratio stated here denotes the approximate ratio of parameters to be pruned since the number of parameters within each pruned structure does not perfectly match the total number of pruned parameters.
2 Zero-shot Performance
Table 1,2,4 and 5 shows the zero-shot performance of the pruned model. Based on the evaluation conducted on LLaMA, employing a 20% parameter reduction without post-training, the pruned model manages to retain 89.8% of the performance exhibited by the unpruned model. Furthermore, through the efficient post-training, the classification accuracy further improves to 60.07%, achieving 94.97% of the accuracy attained by the original model. This demonstration proves the feasibility of using LLM-Pruner to effectively compress the model, even without relying on training data, and within a remarkably short period of time. Surprisingly, we discover that on most datasets, the pruned model with 5.4B LLaMA even outperformed chatGLM-6B. This highlights the superiority of the LLM-Pruner: if a smaller model with a customized size is required, LLM-Pruner is more cost-effective compared to retraining another model with a satisfying performance. However, with 50% parameters pruned, a large accuracy degradation is observed (see Appendix C.5). Compressing LLMs under high compression rates still remains a large challenge.
The compression results of Vicuna-7B align with those of LLaMA, as pruning 20% of parameters on Vicuna-7B maintains performance at 92.03% of the original model. We test a smaller pruning rate of 10% on chatGLM-7B, where the pruned model only experiences a marginal performance decrease of 0.89%, which can be recovered through post-training. Despite the pruned model outperforming the uncompressed model, we don’t assert it is better than the original model. This is largely because chatGLM-6B, a bilingual model, has limited English pre-training exposure. Post-training, however, introduces it to more English corpus, albeit limited, improving its English comprehension.
We conduct tests on all proposed importance estimation techniques mentioned in Section 3.2. The results can be found in Table 1 and 4. Here, represents the importance evaluation utilizing the n-th order term in Eq.5. Vector represents the result corresponding to Eq.3. Based on the results obtained from LLaMA-7B and Vicuna-7B, pruning algorithms achieved the best average performance mostly by leveraging the second-order derivatives for each parameter. Nonetheless, given that first-order derivatives are considerably more efficient than second-order derivatives, though yielding slightly inferior results, we still vote for the first-order term as a competitive method. Besides, the results on chatGLM-7B differed significantly from these findings. The importance estimation on each parameter fails, performing even worse than l2, while the importance estimation on the weight matrix reaches the best performance.
Channel Strategy vs. Block Strategy.
From the results presented in Table 2, it is evident that pruning ‘Channel’ significantly deteriorates performance compared to pruning ‘Block’. This discrepancy arises because the layers within the stacked transformer do not evenly distribute their importance. As shown in Figure 3, the first and last layers have a profound impact on the model’s performance, and pruning them results in more substantial performance degradation compared to other layers. However, due to the uniform treatment of the ‘Channel’ group across all layers, it becomes inevitable to prune the first and last layers, leading to a significant decline in performance.
3 More Analysis
We investigate the impact of pruning the LLM at various pruning ratios in Figure 5. We compare our pruning results with the L2 strategy because L2 is also a data-free pruning algorithm. It is observed in the experiment of LLaMA that when the pruning ratio reaches approximately 20%, the magnitude-dependent algorithm experiences a rapid collapse, leading to the loss of information. Conversely, by employing LLM-Pruner, we are able to increase the pruning ratio to around 60% while achieving an equivalent perplexity level. Furthermore, in the case of Vicuna-7B, removing 10% parameters results in a performance decline equivalent to that of LLM-Pruner with 60%. The utilization of LLM-Pruner enables a significant increase in the number of model parameters that can be pruned, thereby substantially reducing computational overhead.
Tuning on the External Dataset.
To tune the pruned model, we utilize the external dataset Alpaca Taori et al. (2023). The evaluation curves of the pruned model on two zero-shot datasets during the post-training process are depicted in Figure 5. The results demonstrate a rapid decrease in the perplexity of the pruned model within 300 steps, followed by a gradual increase. We provide a more comprehensive evaluation in Appendix C.4. It is important to note that if the model is trained for an excessive number of steps, it runs the risk of overfitting the external dataset, potentially compromising its performance in other general-purpose tasks.
Impact of Dependency-based Structured Pruning.
To study the importance of dependency-based structural pruning, we conduct an experiment to disrupt dependencies within groups, where each weight matrix is pruned solely based on the importance score estimated on itself. Table 7 presents the results demonstrating the impact of dependencies in structural pruning. In the absence of dependencies, the model nearly fails in the zero-shot generation and classification tasks. Even with tuning, the model fails to recover, showing a substantial difference compared to the results in dependency-based pruning.
Impact of Different Aggregation Strategies.
We conduct tests on the aggregation algorithms proposed in Section 3.2. Our experimental results unveil notable discrepancies in model performance across different aggregation strategies, with particular emphasis on the ‘Last-only’ strategy. Among the evaluated approaches, the ‘Max’ strategy attains the most favorable outcomes in terms of perplexity, signifying enhanced coherence and fluency in sentence generation. However, it is important to note that the ‘Max’ strategy exhibits the poorest zero-shot classification results compared to all four strategies. Conversely, the ‘Last-only’ strategy showcases superior classification performance but suffers from the poorest generation quality. In our experiments, we make a trade-off by selecting the ‘Sum’ strategy since it shows both good generalization quality and classification performance.
Comparison with DistilBERT
We show the comparison results of DistilBERT and LLM-Pruner on LLaMA-7B in Table 8. LLM-Pruner outperforms DistilBERT by 4.24% on average with even a smaller size. The reason lies in that LLM-Pruner minimizes model disruption during pruning, whereas DistilBERT merely selects one layer out of two. As a result, the model pruned by LLM-Pruner demands less data to recover its performance compared with DistilBERT, consequently achieving superior performance.
Scratch Training vs. Pruning.
We compare LLM-Pruner with StableLM-3Bhttps://huggingface.co/stabilityai/stablelm-tuned-alpha-3b with a similar parameter size. To ensure fairness, both models are fine-tuned on the Alpaca dataset. The experimental results of these two models are shown in the Table 9. LLM-Pruner crafts lightweight LLMs with low resources, and even can sometimes achieve better performance than LLMs from scratch training. However, we also acknowledge that the LLaMA-3B obtained by LLM-Pruner will not always outperform other 3B models from scratch training, due to the huge gap in the size of training corpus.
Case Study.
We provide some examples of sentences generated by the model compressed using LLM-Pruner in Table 10. We made efforts to ensure a minimal overlap between these generated sentences and the information contained in the tuning corpus, which demonstrates that the information originates from the original model rather than the tuning corpus. We provide additional examples in the Appendix, including the generated sentences of the model without post-training. From the cases in Table 10, it is evident that the sentences generated by the compressed model are comparable to those produced by the original model. They exhibit fluency, relevance, and informativeness regarding the given topic. Nevertheless, during our experiments, we observed that the pruned model’s performance deviates from that of the original model, particularly when generating lengthy sentences. Occasionally, it may generate sentences that are meaningless or contain repetitive tokens.
Conclusion
In this paper, we propose LLM-Pruner, a structured pruning approach for large language models. LLM-Pruner aims to compress sizable language models in a task-agnostic manner while minimizing the dependency on the original training corpus and preserving the linguistic capabilities of LLMs. LLM-Pruner accomplishes this by iteratively examining each neuron within the model as a trigger for identifying dependency groups, thereby constructing the LLM’s dependency graph. Subsequently, LLM-Pruner assesses the importance of these groups using both parameter-wise and weight-wise estimation. Finally, we utilize LoRA for fast recovery and adjustment of the pruned model. We evaluate the efficacy of LLM-Pruner on three distinct models—LLaMA, Vicuna, and ChatGLM—utilizing various zero-shot datasets. Our experimental results indicate that LLM-Pruner successfully prunes the model, reducing computational burden while retaining its zero-shot capabilities. Nevertheless, considerable performance degradation occurs when employing high pruning rates, such as the removal of 50% of LLaMA’s parameters, resulting in a substantial decline in model performance. Additionally, we observe instances in which the model generates incoherent sentences. Addressing the challenges associated with compressing LLMs at higher pruning rates remains a challenging task.
References
Appendix A Detailed Explanations for the Dependency Rules
We provide a detailed explanation of the two dependency rules. It is important to note that these dependency rules do not pertain solely to the forward computation. Instead, they represent directional relationships that exist in both directions. For instance, removing a node in a subsequent layer may also result in the pruning of a node in the preceding layer. Recall the two dependency rules as follows:
where and are two neurons. and represents all the neurons that point towards or point from . and represents the in-degree and out-degree of neuron .
Figure 6 serves as an illustration of the two dependency rules:
In case 1, Node I and Node J satisfy the rule stated in Eq.7. Consequently, Node J depends on Node I. When Node I is pruned, it is necessary to prune Node J as well.
In case 2, Node I and Node J satisfy the rule Eq.8. Thus, Node I is dependent on Node J. If Node J is pruned, it becomes imperative to prune Node I as well.
In case 3, Node J and Node K do not meet the requirement of Eq.7 due to the mismatch in . Thus, with Node J pruned, Node K would not be affected.
Appendix B Implementation Details
Given the lack of previous work on the structural pruning of Large Language Models in a task-agnostic and low-resource setting, there is currently no existing baseline for our model. To provide a comprehensive demonstration of the effectiveness of LLM-Pruner, we employ two additional methods for evaluation, alongside the data-free pruning method. All of these methods are built upon the dependent groups identified in Section 3.1:
L2: We assess the importance of each group based on the magnitude of its weight matrix.
Random: This method involves randomly selecting certain groups for pruning.
For the ‘Block’ Group.
Based on the findings presented in Table 3, it is preferable to leave the first three layers and the final layer unchanged, as modifying parameters in those layers significantly impacts the model. Within each module, such as the MLP or the Multi-head Attention, the discovered groups are pruned based on a pre-set ratio. For instance, in the MLP layer of LLaMA-7B, we identified 11,008 groups, and with a 25% pruning ratio, the module would prune 2,752 groups. It is worth noting that the pruning rate for the selected groups is higher than the pruning ratio for the parameters, as certain layers (e.g., the embedding layer and excluded layers mentioned) retain their parameters. When aiming for a parameter pruning ratio of 20%, we prune 25% from Layer 5 to Layer 30. Similarly, for a 50% parameter removal, we prune 60% of the groups from Layer 4 to Layer 30.
For the ‘Channel’ Group.
The Group ’Channel’ exhibits a resemblance to dimension pruning in the model, targeting the pruning of certain dimensions. In the case of the Query, Key, and Value projection in MHA, only the input dimension is pruned, while for the Output projection in MHA, only the output dimension is pruned. It is important to note that the entire dependency is established automatically, without any manual design involved. The ‘Channel’ group operates in a complementary manner to the ‘Block Group’. In the ‘Channel’ Group, the pruning ratio of the group equals to the pruning ratio of the parameters, as all weight matrices, including the embedding matrix, undergo pruning. Therefore, a 20% pruning ratio of parameters implies pruning 20% of the groups, while a 50% pruning ratio implies pruning 50% of the groups.
B.2 For Recovery Stage
We follow Hu et al. (2021) in our recovery stage. We set the rank to 8 in our experiment. The learning rate is set to 1e-4 with 100 warming steps. The batch size of training is selected from {64, 128} and the AdamW optimizer is employed in our experiment. The best training epoch we found is 2 epochs, as training with more epochs even has a negative impact on the model performance. We run our experiment on a single GPU with 24GB memory, using approximately 2.5 hours if RTX4090 is utilized. All the linear module is taken into account for efficient tuning. An ablation experiment for this is shown in Table 11.
Appendix C More Analysis
Despite our primary experiments being conducted using 50k samples, we remain convinced that the inclusion of additional data could substantially enhance the recovery process, albeit at a considerably higher computational cost. Consequently, we conduct an experiment aimed at model recovery with more data, employing a dataset comprising 2.59 million samples Wu et al. (2023). The results are detailed in Table 12. From the results, it is evident that the performance of the compressed model closely approximates that of the base model, exhibiting only a marginal performance decrease of 0.89%.
C.2 Pruning vs. Quantization
Here, we conduct a comparative analysis of different compression techniques and illustrate that these techniques can be effectively combined with little performance degradation. We have chosen LLM.int8() Dettmers et al. (2022) as a representative example of quantization methods. Our results show that LLM.int8() outperforms LLM-Pruner while LLM-Pruner enhances latency, reduces parameter size. When these two techniques are applied in tandem, they collectively reduce memory consumption and expedite inference, offering a balanced approach that combines the benefits of both methods.
C.3 Global Pruning vs. Local Pruning
we present a comparative analysis between global pruning and local pruning, where the pruning ratio is 20% and the base model is LLaMA-7B. Global pruning refers to ranking all groups in the model together, whereas local pruning involves only ranking groups within the same module for pruning. The outcome of global pruning leads to varying widths across different layers and modules, whereas local pruning ensures uniformity across all layers.
Based on our experimental findings, we observed a slight advantage of local pruning over global pruning. We think this is because of the varying magnitudes in different layers or modules, which makes the importance scores incomparable between groups across different layers.
C.4 Overfitting Phenomena in Post-Training
We present a comprehensive analysis of the overfitting issue in the recovery stage, as previously mentioned in Figure 5. Here the results cover all 9 datasets across various training steps. Based on the findings presented in Table 15, a noticeable trend emerges: the accuracy or generation quality initially shows improvement but subsequently experiences a slight decline. This pattern suggests that the recovery process is completed within a short period. And given that the training corpus is domain-constrained, more training epochs can result in overfitting to the specific dataset while potentially compromising the original capabilities of the language model.
C.5 Pruning with Large Rates
Additionally, we conducted tests on LLaMA-7B and Vicuna-7B with 50% parameters pruned. We observe a significant decrease in performance compared to the base model. However, the recovery stage proved to be beneficial, resulting in an improvement of approximately 7.39%. Pruning a Language Model with such a high pruning rate remains a challenging task.
Appendix D Generations From Compressed Model
Table 18, 19, 20 and 21 show more examples of the models pruned by LLM-Pruner. We present the generation results of both the pruned model with post-training and without post-training. The absence of post-training allows us to better understand the information retained in the model. We include the results of ChatGLM-6B in two languages as it is a bilingual model.