Understanding and Improving Knowledge Distillation
Jiaxi Tang, Rakesh Shivanna, Zhe Zhao, Dong Lin, Anima Singh, Ed H. Chi, Sagar Jain
Introduction
Recent advances in artificial intelligence have largely been driven by learning deep neural networks, and thus, current state-of-the-art models typically require a high inference cost in computation and memory. Therefore, several works have been devoted to find a better quality and computation trade-off, such as pruning (Han et al., 2015b) and quantization (Han et al., 2015a; Jacob et al., 2018). One promising and commonly used method for addressing this computational burden is Knowledge Distillation (KD), proposed by Hinton et al. (2015), which uses a larger capacity teacher model (ensembles) to transfer its ‘dark knowledge’ to a more compact student model. Through distillation, one hopes to achieve a student model that not only inherits better quality from the teacher, but is also more efficient for inference due to its compactness. Recently, we have witnessed a huge success of knowledge distillation, irrespective of the model architecture and application domain (Kim & Rush, 2016; Chen et al., 2017; Tang & Wang, 2018; Anil et al., 2018; He et al., 2019).
Despite the large success of KD, surprisingly sparse research has been done to better understand the mechanism of how it works, which could limit the applications of KD and also raise unexpected or unexplainable results. For example, to successfully ‘distill’ a better student, one common practice is to have a teacher model with as good quality as possible. However, recently Mirzadeh et al. (2019) and Müller et al. (2019) have found this intuition would fail under certain circumstances. Furthermore, Anil et al. (2018) and Furlanello et al. (2018) have analyzed that even without using a powerful teacher, distilling a student model to itself using mutual or self-distillation also improves quality. To this end, some researchers have made attempts on understanding the mechanism of KD. For example, Yuan et al. (2019) connects label smoothing to KD. Furlanello et al. (2018) conjectures KD’s effect on re-weighting training examples. In this work, we found that the benefits of KD comes from a combination of multiple effects, and propose partial KD methods to dissect each of the effects.
This work is an attempt to shed light upon the ‘dark knowledge’ distillation, making this technique less mysterious. More specifically, we make the following contributions:
For KD on multi-class classification task, we systematically break down its effects into: (1) label smoothing from universal knowledge, (3) injecting domain knowledge of class relationships to student’s output logit layer geometry, and (2) gradient rescaling based on teacher’s measurement of instance difficulty. We provide theoretical analyses on how KD exhibits these effects, and improves student model’s quality (Section 3).
We propose partial-distillation techniques using hand-crafted teacher’s output distribution (Section 4) to simulate and validate different effects of knowledge distillation.
We empirically demonstrate and confirm our hypothesis on the effects of KD on both synthetic and real-world datasets. Furthermore, using our understanding, we diagnose some recent failures of applying KD (Section 5).
Related Work
Analyzing Mechanisms of Knowledge Distillation
In the next two subsections, we analyze the unique characteristics of real teacher’s distribution over uniform distribution, and demonstrate how they could potentially facilitate student model’s training.
2 Domain knowledge – teacher injects class relationships prior
KD leverages class relationships as captured by the teacher’s probability distribution over the incorrect classes. As argued by Hinton et al. (2015) on MNIST dataset, model assigns relatively high probability for class ‘7’, when the ground-truth class is ‘2’. In this section, we first confirm their hypothesis using empirical studies. Then, we provide new insights to explain how the teacher informs the class relationships to its student at optimality, and improves model quality.
In this work, we found that the teacher’s predictions on incorrect classes also provides a prior for student model training. Before diving into the details, we recall the case of label smoothing (Szegedy et al., 2016):
From an optimization point of view, He et al. (2019) showed that there is an optimal constant margin , between the logit of the ground-truth , and all other logits , using a label smoothing factor of . For fixed number of classes , the margin is a monotonically decreasing function of .
From geometry perspective, Müller et al. (2019) showed the logit for any class is a measure of squared Euclidean distance on latent space between the activations of the penultimate layerHere can be concatenated with a “1” to account for the bias. , and weights for class in the last logit layer.
where is the activations of the penultimate layer. (See proof in Suppl. Section 7.1)
3 Instance specific knowledge – teacher rescales gradients based on event difficulty
Another important characteristic of the teacher distribution is that the prediction (confidence) on the ground-truth class is different across instances. Comparing ratio of gradients (eqs. 1 and 2):
To validate our claim, in Figure 2, we plot the relationship between and at the end of training. On CIFAR-100 (Krizhevsky et al., 2009), we use ResNet (He et al., 2016) with depth 20 as the student model, and depth 56 as the teacher (see Suppl. Section 7.2 for more details). The plot shows a clear positive correlation between the two. Notably, the correlation will be even stronger when closer to the beginning of training.
4 Summary on primary effects of KD
Isolating Effects by Partial Knowledge Distillation Methods
Empirical Studies
In this section, we evaluate the effectiveness of our proposed partial-distillation methods, to better understand how much each of these effects benefits the student model, and how the improvements are associated with the dataset properties. With our understandings, we propose a simple way to improve distillation quality and we diagnose the recent failures of KD.
2 How effective are the partial-distillation methods?
Setup. On CIFAR-100 we use ResNet-20 as the student, and ResNet-56 as the teacher. On ImageNet with 1000 classes, we use ResNet-50 as the student, and ResNet-152 as the teacher. For more details, please refer to Section 7.2 in Suppl. Note that instead of using different model families as in (Furlanello et al., 2018; Yuan et al., 2019), we use the same model architecture (i.e., ResNet) with different depths for the student and teacher to isolate any unknown effects introduced by model family discrepancy.
3 Regulated knowledge sharing improves distillation
4 Diagnosis of failure cases
For another failure case, Mirzadeh et al. (2019) showed that the ‘distilled’ student model’s quality gets worse as we continue to increase teacher model’s capacity. Larger capacity teacher might overfit, and predict high (uniform) confidence on the ground truth on all the examples; and thereby hindering the effectiveness of gradient rescaling. Another explanation could be that there exists an optimal model capacity gap between the teacher and student, which could otherwise result in an inconsistency between teacher’s prediction confidence on the ground-truth, and the desired example difficulty for the student. Perhaps, an ‘easy’ example for larger capacity teacher is overly difficult for the student.
Conclusion and Future Work
Appendix
where is the activations of the penultimate layer.
At the optimal solution of the student, equating gradient in Equation (2) to , we get:
Plugging the above in equation 4, we get:
Now, sum of the incorrect class gradients is given by:
2 Experimental details
In practice, the gradients from the RHS of Equation (2) are much smaller compare to the gradients from LHS when temperature is large. Thus, it makes tuning the balancing hyper-parameter become non-trivial. To mitigate this and make the gradients from two parts in the similar scale, we multiply to the RHS of Equation (2), as suggested in (Hinton et al., 2015).
Synthetic dataset.
For the synthetic dataset, we generate a single data-point as follows:
Randomly sample an input data point in -dimensional feature space .
After producing basis vectors with procedure (1) and (2), we run procedure (3) and (4) for times with fixed basis vectors to generate a synthetic dataset . By tuning the cosine similarity parameter , we can control the classes correlations within the same super-class. Setting task-difficulty generates a linearly separable dataset, and generates more non-linearities by the function (see Figure 6 in Suppl. for visualization on a toy example).
Following the procedure showed above, we get a toy synthetic dataset where we only have input dimensionality with classes and super-classes. Figure 6 shows a series of scatter plots with different settings of class similarity and task difficulty . This visualization gives a better understanding of the synthetic dataset and helps us imagine what it will look like in high-dimensional setting that used in our experiments. For the model used in our experiments, besides they are 2-layer network activated by , we use residual connection (He et al., 2016) and and batch normalization (Ioffe & Szegedy, 2015) for each layer. Following (Ranjan et al., 2017; Zhang et al., 2018), we found using -normalized logits layer weight and penultimate layer provides more stable results. The model is optmized by Adam (Kingma & Ba, 2014) for a total of 3 million steps without weight decay and we report the best accuracy. Finally, Nvidia V100 GPU is used as the accelerator hardware. Please refer to Table 5(a) for the best setting of hyper-parameters.
CIFAR-100 dataset.
CIFAR-100 is a relatively small dataset with low-resolution () images, containing training images and validation images, covering classes and super-classes. It is a perfectly balanced dataset – we have the same number of images per class (i.e., each class contains training set images) and classes per super-class. To process the CIFAR-100 dataset, we use the official split from Tensorflow Datasethttps://www.tensorflow.org/datasets/catalog/cifar100. Both data augmentation https://github.com/tensorflow/models/blob/master/research/resnet/cifar_input.pyWe turn on the random brightness/saturation/constrast for better model performance. for CIFAR-100 and the ResNet modelhttps://github.com/tensorflow/models/blob/master/research/resnet/resnet_model.py are based on Tensorflow official implementations. Also, following the conventions, we train all models from scrach using Stochastic Gradient Descent (SGD) with a weight decay of 1e-3 and a Nesterov momentum of 0.9 for a total of 10K steps. The initial learning rate is 1e-1, it will become 1e-2 after 40K steps and become 1e-3 after 60K steps. We report the best accuracy for each model. All experiments on CIFAR-100 are conducted by using Nvidia V100 GPU as the accelerator hardware. Please refer to Table 5(b) for the best setting of hyper-parameters.
ImageNet dataset.
ImageNet contains about M training images and test images, all of which are high-resolution (), covering classes. The distribution over the classes is approximately uniform in the training set, and strictly uniform in the test set. Our data preprocessing and model on ImageNet dataset are follow Tensorflow TPU official implementationshttps://github.com/tensorflow/tpu/tree/master/models/official/resnet. The Stochastic Gradient Descent (SGD) with a weight decay of 1e-4 and a Nesterov momentumof 0.9 is used. We train each model for 120 epochs, the mini-batch size is fixed to be 1024 and low precision (FP16) of model parameters is adopted. We didn’t change the learning rate schedule scheme from the original implementation. Please refer to Table 5(c) for the best setting of hyper-parameters. We used TPU-v3 as the accelerator hardware.