MiniLLM: On-Policy Distillation of Large Language Models

Yuxian Gu, Li Dong, Furu Wei, Minlie Huang

Introduction

With the rapid development of large language models (LLMs; 24, 3, 6, 42, 13), a common technique to reduce high computational resource demand is knowledge distillation (KD; 23), where we train a small student model with supervision from a large teacher model. Two categories of KD are commonly applied: black-box KD, where only the teacher predictions are accessible, and white-box KD, where the teacher parameters are available to use . Recently, black-box KD has shown promising results in fine-tuning small models on the prompt-response pairs generated by LLM APIs . With the emergence of more open-source LLMs , white-box KD becomes more valuable for both research communities and industry sectors because student models receive better signals from white-box teacher models, thereby potentially resulting in higher performance. However, white-box KD approaches are mostly studied for small (<< 1B parameters) language understanding models , while white-box KD for generative LLMs is yet to be explored.

In this work, we investigate white-box KD of LLMs. We argue that the standard KD objectives are sub-optimal for LLMs that perform tasks in a generative manner. Given the teacher distribution p(y∣x)p(\bm{y}|\bm{x}) and the student distribution qθ(y∣x)q_{\theta}(\bm{y}|\bm{x}) parameterized by θ\theta, standard KD objectives (including several variants for sequence-level models) essentially minimize the approximated forward Kullback-Leibler divergence (KLD) between the teacher and the student distribution, termed as KL⁡[p∣∣qθ]\operatorname{KL}[p||q_{\theta}], which forces pp to cover all the modes of qθq_{\theta}. For text classification tasks, KL⁡[p∣∣qθ]\operatorname{KL}[p||q_{\theta}] works well because the output space usually consists of finite-number classes such that both p(y∣x)p(\bm{y}|\bm{x}) and qθ(y∣x)q_{\theta}(\bm{y}|\bm{x}) have few modes. However, for open text generation tasks, where the output spaces are much more complex and p(y∣x)p(\bm{y}|\bm{x}) can contain much more modes than what qθ(y∣x)q_{\theta}(\bm{y}|\bm{x}) can express due to the limited model capacity. Minimizing forward KLD can cause qθq_{\theta} to assign unreasonably high probabilities to the void regions of pp and produces samples very unlikely under pp during free-run generation .

To alleviate this problem, we propose to minimize the reverse KLD, KL⁡[qθ∣∣p]\operatorname{KL}[q_{\theta}||p], which is widely used in computer vision and reinforcement learning . Compared to KL⁡[p∣∣qθ]\operatorname{KL}[p||q_{\theta}], minimizing KL⁡[qθ∣∣p]\operatorname{KL}[q_{\theta}||p] causes qθq_{\theta} to seek the major modes of pp, and assign low probabilities to the pp’s void regions , which is illustrated by a pilot experiment in Section 2.1. In language generation of LLMs, this means the student model avoids learning too many long-tail variants of the teacher distribution and focuses on the correctness of the generated response, which is critical in practical scenarios that require truthfulness and reliability . To optimize min⁡θKL⁡[qθ∣∣p]\min_{\theta}\operatorname{KL}[q_{\theta}||p] as shown in Section 2.2, we derive the gradient of the objective with Policy Gradient . Although recent works have shown success in fine-tuning PLMs with policy optimization , we found that training the model still suffers from high variance, reward hacking, and generation length bias. Therefore, we introduce (1) single-step regularization to reduce variance, (2) teacher-mixed sampling to alleviate reward hacking, and (3) length normalization to eliminate the length bias. Finally, we introduce the algorithm of MiniLLM in Section 2.3.

We apply MiniLLM to various generative language models with parameter sizes ranging from 120M to 13B in the instruction-following setting that covers a large range of NLP tasks. We use five instruction-following datasets with the GPT-4 feedback and Rouge-L for evaluation. Our experiments show that MiniLM consistently outperforms the standard KD baselines on all the datasets and scales up well from 120M to 13B models (see Figure 1). More analysis shows that MiniLLM has lower exposure bias, better calibration, and performs better at generating long responses with better diversity.

Method

where p′p^{\prime} can be real data distribution (word-level KD) or teacher distribution pp (sequence-level KD). Though widely used, KL⁡[p∣∣qθ]\operatorname{KL}[p||q_{\theta}] has been shown to overestimate the void regions of pp in language generation tasks when qθq_{\theta} is insufficiently expressive to cover all the modes of p′p^{\prime} . KD for LLMs fits the case because LLMs perform various tasks in a generative manner, such that the low-capacity student models cannot perfectly imitate the complex language generation distribution of the teacher models or humans.

In this work, we consider minimizing the reverse KLD between the student and teacher distributions as the learning objective of MiniLLM:

has shown that minimizing KL⁡[qθ∣∣p]\operatorname{KL}[q_{\theta}||p] leads to a mode-seeking behavior, where qθq_{\theta} assigns high probabilities to pp’s large modes, ignoring the small ones. Figure 2 illustrates the difference between the two KLDs when a Gaussian distribution tries to fit a Gaussian mixture. We can see that minimizing forward KLD causes qθq_{\theta} to place large probability mass on the zero-probability places of pp, which corresponds to the generation of low-quality texts in practice, while reverse KLD focuses on pp’s major modes, which is crucial to ensure the correctness and faithfulness of language generation.

We term our KD method for LLMs by minimizing the reverse KLD as MiniLLM, which is illustrated in Figure 3. Unlike sequence-level KD, MiniLLM does not force qθq_{\theta} to fit all y\bm{y} sampled from the teacher distribution pp. Instead, it encourages the student to generate samples preferred by the teacher within its own capacities, which is more possible to achieve.

2 Optimization with Policy Gradient

We notice that the gradient of the objective function J(θ)\mathcal{J}(\theta) described in Equation (1) can be derived using the Policy Gradient Theorem :

where T=∣y∣T=|\bm{y}| and Rt=∑t′=tTlog⁡p(yt′∣y<t′,x)qθ(yt′∣y<t′,x)R_{t}=\sum_{t^{\prime}=t}^{T}\log\frac{p(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})}{q_{\theta}(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})} is the accumulation of rt′=log⁡p(yt′∣y<t′,x)qθ(yt′∣y<t′,x)r_{t^{\prime}}=\log\frac{p(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})}{q_{\theta}(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})} that measures the quality of each step generation. Intuitively, we want the generation to have high probabilities under the teacher distribution by increasing p(yt′∣y<t′,x)p(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x}), but simultaneously stay diverse by lowering qθ(yt′∣y<t′,x)q_{\theta}(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x}). The expectation is computed by Monte-Carlo sampling. However, policy gradient suffers from high variance and reward hacking . Although subsequent works proposed better solutions like PPO , we find that these issues still remain. Besides, we notice that RtR_{t} favors short sentences, which causes our model to output empty responses. Therefore, we propose three strategies to mitigate these problems.

Single-Step Regularization

Teacher-Mixed Sampling

We observe reward hacking when training with Equation 2 because qθq_{\theta} sometimes produces degenerated sentences y\bm{y} that receive high scores from the teacher (e.g., repeated phrases) during sampling, especially for small student models. To create a better sampling distribution, we mix the teacher and the student distribution at each time step:

where α\alpha controls the strength of the teacher mix-in. Sampling from p~\widetilde{p} suppresses low-quality generation with the teacher’s help and alleviates reward hacking. We re-write (∇J)Main(\nabla\mathcal{J})_{\text{Main}} and (∇J)Reg(\nabla\mathcal{J})_{\text{Reg}} with importance sampling to get to an unbiased estimator of the gradient :

where wt=∏t′=1tqθ(yt′∣y<t′,x)p~(yt′∣y<t′,x)w_{t}=\prod_{t^{\prime}=1}^{t}\frac{q_{\theta}(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})}{\widetilde{p}(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})} is the importance weight. However, in practice, using wtw_{t} has been found to be sensitive to hyper-parameters and converge slowly. Therefore, we approximately set wt≈qθ(yt∣y<t,x)p~(yt∣y<t,x)w_{t}\approx\frac{q_{\theta}(y_{t}|\bm{y}_{<t},\bm{x})}{\widetilde{p}(y_{t}|\bm{y}_{<t},\bm{x})} to reduce the variance of the estimator in Equation 5 .

Length Normalization

We found that long sequences tend to have small Rt+1R_{t+1}, which encourages the model to produce short responses. Therefore, we add length normalization to Rt+1R_{t+1} in Equation 3:

In Summary

Combining the strategies listed above, we have the final optimization gradient:

where VV is the vocab size of the language model.

3 Training Algorithm

The training algorithm of MiniLLM is shown in Algorithm 1. We initialize the student model from a checkpoint fine-tuned on the training data with the lowest validation loss and add the PPO clipping strategy to (∇J)Main(\nabla\mathcal{J})_{\text{Main}} to improve training stability. Note that we do not use the value network and the KL regularization in PPO to improve the training efficiency. Same as , we add a language modeling loss LPT\mathcal{L}_{\text{PT}} on the pre-training corpus.

Experiments

We conduct experiments by first fine-tuning a large model on the instruction-response dataset D\mathcal{D} as the teacher pp. Then, we compare different KD methods to distill a smaller student model on D\mathcal{D} with the teacher’s guidance by evaluating the instruction-following performance of the distilled model.

We distill three kinds of models with various sizes: GPT-2 (120M, 340M, 760M), OPT (1.3B, 2.7B, 6.7B), and LLaMA (7B), using GPT-2-1.5B, OPT-13B, and LLaMA-13B as the teacher for each model type respectively. We also present the results using GPT-J as the teacher in Appendix C.1.

Training

We construct the training data from databricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master consisting of 15K human-written instruction-response pairs. We randomly split 14K samples as the training set D\mathcal{D} and left 500 samples for validation and testing, respectively. For DPT\mathcal{D}_{\text{PT}}, we use the OpenWebText for the GPT-2 family and the RoBERTa training corpus for other models. We set the teacher-mix-in strength α=0.2\alpha=0.2 throughout the experiments. We use the Rouge-L score on the validation set to select the hyper-parameters because it aligns with human preference better than the validation loss . More training details are shown in Appendix B.1.

Evaluation

We evaluate the trained models on five instruction-following datasets:

DollyEval: the 500-sample test set we split from the databricks-dolly-15k dataset.

SelfInst : A user-oriented instruction-following set with 252 samples.

VicunaEval : The 80 challenging questions used in the Vicuna evaluation.

S-NI: The test set of Super-NaturalInstructions consisting of 9K samples ranging from 119 tasks. Following , we split the set into 3 subsets whose ground truth response lengths lie in ,,, and [11,+∞][11,+\infty]. We use the [11,+∞][11,+\infty] subset in Section 1 and conduct an analysis on all subsets in Section 3.3.

UnNI: The core set of UnnaturalInstructions containing 60K samples. Similar to S-NI, we first conduct the evaluations on the [11,+∞][11,+\infty] subset, followed by an analysis of the performance on all subsets in Appendix C.2.

We adopt two metrics to evaluate the model-generated responses:

R-L: The Rouge-L score to measure the precision of the model generation. has shown that Rouge-L is suitable for large-scale instruction-following evaluation.

GPT4: The GPT-4 feedback by asking GPT-4 to compare model-generated responses with the ground truth answersWe use the ChatGPT’s generation for VicunaEval’s ground truth responses and raise 1-10 scores for both responses (see Appendix B.2 for the prompt we use). We report the ratio of the total score of model responses and ground truth answers. This metric is only applied to DollyEval, SelfInst, and VicunaEval.

For all test sets, we sample the responses with the temperature = 1 and report the average scores of 5 generations for each prompt with different random seeds.

Baselines

We consider three baselines in our main experiment:

SFT w/o KD directly fine-tunes the student model on D\mathcal{D} supervised with the golden responses.

KD fine-tunes the student model on D\mathcal{D} using the teacher distribution as the supervision at each token step, also known as word-level KD.

SeqKD fine-tunes the student model on the teacher-generated data.

2 Results

We present the evaluation results in Table 1, from which we have four observations.

First, by comparing SFT with KD and SeqKD that approximately minimize the forward KLD, we can see that these standard KD methods successfully distill knowledge from the teacher model in most cases, achieving better Rouge-L and GPT-4 feedback scores, which restates the conclusions in previous works .

Second, by comparing the GPT-4 feedback score of MiniLLM with the baselines, we observe that the model distilled by our method outperforms the baselines in almost all cases when trained with different base models and tested on various evaluation sets. This indicates that MiniLLM is a general method to distill small models with high overall performance. We also find that MiniLLM generally works better on datasets other than DollyEval compared with the baselines, indicating the good out-of-distribution generalization of our method.

Third, the Rouge-L scores show that the MiniLLM models produce the most precise responses that have high overlaps with the ground truth. We notice that in some cases, especially on VicunaEval, S-NI, and UnNI, student models reach higher Rouge-L scores than the teacher, which matches the observation in . We conjecture the reason is that the standard teacher-forcing fine-tuning brings the teacher training-inference discrepancy, also known as exposure bias . On the contrary, MiniLLM is optimized with policy optimization methods, which alleviates exposure bias . We include further analysis on exposure bias in Section 3.3.

Fourth, comparing the results across model sizes and model families, we can see that the improvement of MiniLLM is consistent when the base model sizes vary from 120M to 13B across three model families. This tendency is also illustrated in Figure 1, which demonstrates the excellent scalability and generalization of our method in the era of LLMs.

3 Analysis

Language generation models trained to minimize forward KLD are known to suffer from exposure bias caused by the discrepancy between teacher-forcing training and free-run generation. In MiniLLM, we collect the samples from the student model during the training stage, which alleviates the mismatch between the training and evaluation . In Figure 4, we use the ExAccErr metric defined in Appendix B.3 to measure the excess accumulated error due to exposure bias in auto-regressive decoding. The experiment is based on GPT-2-125M, with GPT-2-1.5B as the teacher, using DollyEval as the test set. For each prompt, we sample 10 responses to reduce the variance. We can see that the ExAccErr of the fine-tuned model continuously grows during generation, while MiniLLM has a much lower ExAccErr, and the error stops accumulating for long text generation (>> 150 tokens).

Calibration

has shown that the RL-trained model is likely to be poorly calibrated. We test the calibration of MiniLLM and the KD baselines on two widely-used text classification datasets: SST2 and BoolQ , based on the LLaMA-7B model. We design zero-shot classification instructions (see Appendix B.2) and take the probability of the label words to compute the ECE scores . From Figure 4, we can see that the models trained with KD and SeqKD are worse calibrated than the teacher model, which potentially explains their low performance on canonical benchmarks . We suspect the reason is that minimizing forward KLD causes the models to push high probabilities to zero-probability points of the target distribution, which leads to significant distribution difference between the student and the teacher (see the intuitive example in Figure 2). In contrast, MiniLLM focuses on accurately learning the major parts of the target distribution, which narrows the ECE scores gap between the student and the teacher.

Performance on Different Response Length

We study the models’ performance when the golden response lengths belong to different ranges. In Figure 5, we illustrate the Rouge-L scores of different KD models against the SFT models on three S-NI subsets split by the length of the ground truth responses. We can see that all methods achieve low scores on prompts that expect short responses (≤5\leq 5 tokens), probably because most responses in our training set are long sentences, which introduces a distribution shift between training and testing . Furthermore, the output spaces of these prompts are relatively small, allowing the student model to cover most modes of the teacher, and thus reverse KLD and forward KLD have similar performance. For prompts with longer responses (≥6\geq 6 tokens), the teacher distribution contains more modes than the students due to the complex output spaces, which shows the advantage of MiniLLM against standard KD approaches. Similar results on UnNI are shown in Appendix C.2.

Generation Diversity

has found that the model optimized by minimizing reverse KLD is likely to lose modes, which affects the generation diversity. We follow to discuss generation diversity from three aspects: (i) generating multiple distinct responses given a prompt. (ii) generating linguistically complex responses. (iii) the ability to generate contents that have high coverage of the real data distribution. For (i), we argue that for many NLP applications, generating one correct response is sufficient, especially for those scenarios demanding high truthfulness and reliability . For (ii) and (iii), we report the responses’ distinct 4-gram proportion and the language modeling loss on the test sets in Table 5, using the base models from the LLaMA family. We can see that MiniLLM preserves the distinct 4-gram proportion in the generated responses and does not cause the language modeling loss on the test set to increase much.

4 Ablations

We conduct ablation studies on the three strategies proposed to stabilize and accelerate optimization in Section 2.2 by distilling a GPT-2-125M model from the GPT-2-1.5B model. In Table 6, we report the best Rouge-L scores on the validation set of each run and the evaluation results of the corresponding checkpoints. We also plot the reverse KLD between the student and the teacher during training in Figure 6, where the lines are smoothed by 32 steps. We can see that Teacher-Mixed Sampling and Length Normalization are critical to stabilizing training. Although the reverse KLDs also decrease without these strategies, we find that the models quickly learn to generate repeated, short, or meaningless strings that have high probabilities in the teacher distribution (see examples in Appendix D), which is known as reward hacking . This also leads to the low generation performance in Table 6. From Figure 6, we also observe that the Single-Step Regularization effectively reduces the variance of the training process, which also results in higher performance on the validation and test sets.

Effect of Teacher-Mix-in Strength α𝛼\alpha

In Figure 7, we plot the best Rouge-L scores on the validation set of GPT-2-125M, OPT-1.3B, and LLaMA-7B using GPT-2-1.5B, OPT-13B, and LLAMA-13B as the teachers, with different teacher-mix-in strength α\alpha in MiniLLM. α=0.0\alpha=0.0 means we only sample from the student distribution, and when α=1.0\alpha=1.0, we sample entirely from the teacher distribution. We find that α=0.2\alpha=0.2 is generally suitable across different model families and sizes, and larger models are more robust to the choice of α\alpha.

Effect of Adding Pre-Training Loss

In Table 7, we study the effect of adding the pre-training loss in Algorithm 1 by comparing MiniLLM with its variant where the language modeling loss on the pre-training corpus is removed (w/o PT Loss). We have a similar observation as that adding the pre-training loss helps to preserve the abilities on canonical NLP tasks while keeping the performance on instruction-following tasks nearly unchanged.

Related Work

Large language models (LLMs; 6, 42, 13, 1, 60) have shown their superior performance by solving various NLP tasks in a generative manner. Recent works apply instruction tuning or learning from human feedback to improve the alignment of LLMs with humans further and create general AI assistants . There are also efforts to build open-source LLMs to facilitate research and industry development. Although appealing, the broad capacities of LLMs usually only emerge with large parameter sizes that require massive computational resources . Therefore, model compression is critical for the practical deployment and further research of LLMs.

Knowledge Distillation

Knowledge distillation (KD; 23) aims at training a student model with the guidance of a teacher model and is widely used as a model compression technique in many fields of deep learning . In the NLP community, many works apply KD to text classification tasks by training the student to mimic the teacher’s output distribution difference , hidden states , or attention scores . For text generation, the standard KD method is to approximately minimize the forward KLD between the student and the teacher generation distribution by using the teacher’s output at each time step as supervision or direct training on the teacher generations . In this paper, we propose to minimize the reverse KLD, which is more suitable for generative large language models when the teacher distribution is available.

Distribution Discrepancy Metrics in Text Generation

The distribution discrepancy metrics play a significant role in learning text generation models. The forward Kullback-Leibler divergence (KLD) is the standard metric due to its simplicity when derived as the Maximum Likelihood Estimate (MLE) objective . However, previous works show that minimizing forward KLD leads to zero-forcing behavior where models try to cover all modes of the target distribution and sacrifice the accuracy of major modes . Some works resort to using other metrics to remedy this problem, such as reverse KLD , Total Variation Distance , and Optimal Transport . Our paper is the first to tackle this problem for knowledge distillation of LLMs.

Conclusion

In this work, we investigate the problem of distilling knowledge from larger LLMs to smaller ones. We find that the standard distillation methods that minimize the forward KLD is sub-optimal in language generation scenarios because the teacher’s output distribution contains much more modes than the student’s, and forward KLD forces the student distribution to over-estimate the low-probability regions of the teacher distribution. Therefore, we propose MiniLLM that minimizes the reverse KLD between the teacher and student distribution and develop an algorithm to optimize this objective. Extensive experiments in the instruction-following setting show that MiniLLM models produce more precise responses that have higher overall quality than standard KD approaches. We also find that MiniLLM has lower exposure bias, better calibration, and higher performance in long-text generation with good diversity.

References

Appendix A Derivations

We compute the gradient of J(θ)=KL⁡[qθ∣∣p]\mathcal{J}(\theta)=\operatorname{KL}[q_{\theta}||p] with respect to θ\theta using the Policy Gradient Theorem :

where Equation 14 is based on the fact that log⁡qθ(yt∣y<t,x)\log q_{\theta}(y_{t}|\bm{y}_{<t},\bm{x}) can only affect tokens at ≥t\geq t positions in y\bm{y}. By setting Rt=∑t′=tTlog⁡p(yt′∣y<t′,x)qθ(yt′∣y<t′,x)R_{t}=\sum_{t^{\prime}=t}^{T}\log\frac{p(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})}{q_{\theta}(y_{t^{\prime}}|\bm{y}_{<t^{\prime}},\bm{x})}, we obtain Equation 2.

A.2 Derivation of Equation 3

Then, we re-write ∇J(θ)\nabla\mathcal{J}(\theta) as:

where Equation 20 uses the product rule of the gradient and rt=log⁡p(yt∣y<t,x)qθ(yt∣y<t,x)r_{t}=\log\frac{p(y_{t}|\bm{y}_{<t},\bm{x})}{q_{\theta}(y_{t}|\bm{y}_{<t},\bm{x})}.

Appendix B Experimental Details

For models with less than 1.3B parameters, we search for the learning rates in [5e-4, 1e-4, 5e-5], the batch sizes in , and train these models for 20 epochs. For other models, we search for the learning rate in [5e-5, 1e-5, 5e-6], the batch sizes in , and train these models for 10 epochs. For KD, we follow to mix the distillation loss with the language modeling loss on the ground truth responses by a mixture rate of 0.5. The checkpoints of each baseline are selected by the Rouge-L scores on the validation set.

MiniLLM

As shown in Algorithm 1, we first fine-tune the model on the training set using the vanilla language modeling objective to get a starting point of the subsequent MiniLLM training. We fine-tune the model for 3 epochs using the best learning rate and batch size of the corresponding SFT baselines. We select the checkpoint with the lowest validation loss, not the Rouge-L score. Then, we train the model as described in Algorithm 1 using a learning rate 5e-6, a mini-batch size 6464 in all cases. Similar to PPO , we collect 256 sentences at once and adopt 4 inner epochs. The clipping rate ϵ\epsilon is set to 0.2, and the max length of the model is 512. We use temperature = 1 when sampling from qθq_{\theta}. We train the model for at most 5000 steps and select the final checkpoint using the Rouge-L score on the validation set. Our experiments are based on the NVIDIA V100 32G GPUs.

B.2 Evaluation Details

During the evaluation, we sample the responses from each model using temperature = 1, a max-length limit of 512, and random seeds . Similar to , we adopt a prompt wrapper shown in Figure 8 to convert each instruction-response pair to a sentence. For the GPT-4 feedback, we apply the prompt in Figure 9 and set the temperature = 0.7. For the classification tasks in the calibration paragraph of Section 3.3, we prompt the model to do zero-shot text classification with the prompt in Figure 10 and 11.

B.3 Exposure Bias Analysis

Following , we compute the ExAccErr with the following formula:

where R(l)R(l) is the accumulated regret of imitating the teacher distribution pp at the time step ll during the free-run generation:

and ϵ(l)\epsilon(l) is the average per-step error between qθq_{\theta} and pp using the oracle context sampled from pp as the prefix:

Intuitively, the regret of qθq_{\theta} during generation is made of two parts: the error to estimate pp given the oracle context and the error caused by the low-quality model-generated prefix. The former is calculated by lϵ(l)l\epsilon(l), and the latter reflects the exposure bias. Therefore, ExAccErr measures the relative error caused only by exposure bias.

Appendix C Additional Results

We present the evaluation results when using GPT-J as the teacher and GPT-2-760M, GPT-2-1.5B, and GPT-Neo-2.7B as the student in Table 6. MiniLLM outperforms the baselines in most cases.

C.2 Performance of Response Length on U-NI

The performance on different U-NI subsets split by the length of the ground truth response is shown in Figure 12. We have the same observation as in Section 3.3 that on short responses, all KD methods perform similarly, and on long responses, MiniLLM outperforms other methods.

Appendix D Cases

We provide some cases generated by the models distilled by different methods based on LLaMA in Table 7. We find that MiniLLM generates more detailed and accurate responses compared with the baselines.