What Do Compressed Deep Neural Networks Forget?
Sara Hooker, Aaron Courville, Gregory Clark, Yann Dauphin, Andrea Frome
Introduction
Between infancy and adulthood, the number of synapses in our brain first multiply and then fall. Synaptic pruning improves efficiency by removing redundant neurons and strengthening synaptic connections that are most useful for the environment (Rakic et al., 1994). Despite losing of all synapses between age two and ten, the brain continues to function (Kolb & Whishaw, 2009; Sowell et al., 2004). The phrase "Use it or lose it" is frequently used to describe the environmental influence of the learning process on synaptic pruning, however there is little scientific consensus on what exactly is lost (Casey et al., 2000).
In this work, we ask what is lost when we compress a deep neural network. Work since the 1990s has shown that deep neural networks can be pruned of “excess capacity” in a similar fashion to synaptic pruning (Cun et al., 1990; Hassibi et al., 1993a; Nowlan & Hinton, 1992; Weigend et al., 1991). At face value, compression appears to promise you can have it all. Deep neural networks are remarkably tolerant of high levels of pruning and quantization with an almost negligible loss to top-1 accuracy (Han et al., 2015; Ullrich et al., 2017; Liu et al., 2017; Louizos et al., 2017; Collins & Kohli, 2014; Lee et al., 2018). These more compact networks are frequently favored in resource constrained settings; compressed models require less memory, energy consumption and have lower inference latency (Reagen et al., 2016; Chen et al., 2016; Theis et al., 2018; Kalchbrenner et al., 2018; Valin & Skoglund, 2018; Tessera et al., 2021).
The ability to compress networks with seemingly so little degradation to generalization performance is puzzling. How can networks with radically different representations and number of parameters have comparable top-level metrics? One possibility is that test-set accuracy is simply not a precise enough measure to capture how compression impacts the generalization properties of the model. Despite the widespread use of compression techniques, articulating the trade-offs of compression has overwhelmingly focused on change to overall top-1 accuracy for a given level of compression.
The cost to top-1 accuracy appears minimal if it is spread uniformally across all classes, but what if the cost is concentrated in only a few classes? Are certain types of examples or classes disproportionately impacted by compression? In this work, we propose a formal framework to audit the impact of compression on generalization properties beyond top-line metrics. Our work is the first to our knowledge that asks how dis-aggregated measures of model performance at a class and exemplar level are impacted by compression.
Contributions We run thousands of large scale experiments and establish consistent results across multiple datasets— CIFAR-10 (Krizhevsky, 2012), CelebA (Liu et al., 2015) and ImageNet (Deng et al., 2009), widely used pruning and quantization techniques, and model architectures. We find that:
Top-line metrics such as top-1 or top-5 test-set accuracy hide critical details in the ways that pruning impacts model generalization. Certain parts of the data distribution are far more sensitive to varying the number of weights in a network, and bear the brunt of the cost of varying the weight representation.
The examples most impacted by pruning, which we term Pruning Identified Exemplars (PIEs), are more challenging for both models and humans to classify. We conduct a human study and find that PIEs tend to be mislabelled, of lower quality, depict multiple objects, or require fine-grained classification. Compression impairs the model’s ability to predict accurately on the long-tail of less frequent instances.
Pruned networks are more sensitive to natural adversarial images and corruptions. This sensitivity is amplified at higher levels of compression.
While all compression techniques that we evaluate have a non-uniform impact, not all methods are created equal. High levels of pruning incur a far higher disparate impact than is observed for the quantization techniques that we evaluate.
Our work provides intuition into the role of capacity in deep neural networks and a mechanism to audit the trade-offs incurred by compression. Our findings suggest that caution should be used before deploying compressed networks to sensitive domains. Our PIE methodology could conceivably be explored as a mechanism to surface a tractable subset of atypical examples for further human inspection (Leibig et al., 2017; Zhang, 1992), to choose not to classify certain examples when the model is uncertain (Bartlett & Wegkamp, 2008; Cortes et al., 2016), or to aid interpretability as a case based reasoning tool to explain model behavior (Kim et al., 2016; Caruana, 2000; Hooker et al., 2019).
Methodology and Experiment Framework
We consider a supervised classification problem where a deep neural network is trained to approximate the function that maps an input variable to an output variable , formally . The model is trained on a training set of images , and at test time makes a prediction for each image in the test set. The true labels are each assumed to be one of classes, such that .
A reasonable response to our desire for more compact representations is to simply train a network with fewer weights. However, as of yet, starting out with a compact dense model has not yielded competitive test-set performance (Li et al., 2020; Zhu & Gupta, 2017b). Instead, research has centered on a more tractable direction of investigation – the model begins training with "excess capacity" and the goal is to remove the parts that are not strictly necessary for the task by/at the end of training. A pruning method identifies the subset of weights to set to zero. A sparse model function, , is one where a fraction of all model weights are set to zero. Equating weight value to zero effectively removes the contribution of a weight, as multiplication with inputs no longer contributes to the activation. A non-compressed model function is one where all weights are trainable (). We refer to the overall model accuracy as . In contrast, indicates that of model weights are removed over the course of training, leaving a maximum of non-zero weights.
2 Class level measure of impact
If the impact of compression was completely uniform, the relative relationship between class level accuracy and overall model performance will be unaltered. This forms our null hypothesis (). We must decide for each class whether to reject the null hypothesis and accept the alternate hypothesis () - the relative change to class level recall differs from the change to overall accuracy in either a positive or negative direction:
Welch’s t-test Evaluating whether the difference between the samples of mean-shifted class accuracy from compressed and non-compressed models is “real” amounts to determining whether these two data samples are drawn from the same underlying distribution, which is the subject of a large body of goodness of fit literature (D’Agostino & Stephens, 1986; Anderson & Darling, 1954; Huber-Carol et al., 2002). We independently train a population of models for each compression method, dataset, and model that we consider. Thus, we have a sample of accuracy metrics per class at each level of compression .
For each class , we we use a two-tailed, independent Welch’s t-test (Welch, 1947) to determine whether the mean-shifted class accuracy of the samples and differ significantly. If the -value , we reject the null hypothesis and consider the class to be disparately impacted by level of compression relative to the baseline.
Controlling for overall changes to top-line metrics Note that by comparing the relative difference in class accuracy , we control for any overall difference in model test-set accuracy. This is important because while small, the difference in top-line metrics is not zero (see Table. 2). Along with the -value, for each class we report the average relative deviation in class-level accuracy, which we refer to as relative recall difference:
3 Pruning Identified Exemplars
In addition to measuring the class level impact of compression, we are interested in how model predictive behavior changes through the compression process. Given the limitations of un-calibrated probabilities in deep neural networks (Guo et al., 2017; Kendall & Gal, 2017), we focus on the level of disagreement between the predictions of compressed and non-compressed networks on a given image. Using the populations of models described in the prior section, we construct sets of predictions for a given image .
For set we find the modal label, i.e. the class predicted most frequently by the -pruned model population for image , which we denote . The exemplar is classified as a pruning identified exemplar if and only if the modal label is different between the set of -pruned models and the non-pruned baseline models:
We note that there is no constraint that the non-pruned predictions for PIEs match the true label. Thus the detection of PIEs is an unsupervised protocol that can be performed at test time.
4 Experimental framework
Tasks We evaluate the impact of compression across three classification tasks and models: a wide ResNet model (Zagoruyko & Komodakis, 2016) trained on CIFAR-10, a ResNet-50 (He et al., 2015) trained on ImageNet, and a ResNet-18 trained on CelebA. All networks are trained with batch normalization (Ioffe & Szegedy, 2015), weight decay, decreasing learning rate schedules, and augmented training data. We train for steps (approximately epochs) on ImageNet with a batch size of images, for steps on CIFAR-10 with a batch size of , and steps on CelebA with a batch size of . For ImageNet, CIFAR-10 and CelebA, the baseline non-compressed model obtains a mean top-1 accuracy of , and respectively. Our goal is to move beyond anecdotal observations, and to measure statistical deviations between populations of models. Thus, we report metrics and statistical significance for each dataset, model and compression variant across independent trainings.
Pruning and quantization techniques considered We evaluate magnitude pruning as proposed by Zhu & Gupta (2017a). For pruning, we vary the end sparsity precisely for . For example, indicates that of model weights are removed over the course of training, leaving a maximum of non-zero weights. For each level of pruning , we train models from random initialization.
We evaluate three different quantization techniques: float16 quantization float16 (Micikevicius et al., 2017), hybrid dynamic range quantization with int8 weights hybrid (Alvarez et al., 2016) and fixed-point only quantization with int8 weights created with a small representative dataset fixed-point (Vanhoucke et al., 2011; Jacob et al., 2018).
All quantization methods we evaluate are implemented post-training, in contrast to the pruning which is applied progressively over the course of training. We use a limited grid search to tailor the pruning schedule and hyperparameters to each dataset to maximize top-1 accuracy. We include additional details about training methodology and pruning techniques in the supplementary material. All the code for this paper is publicly available here.
Results
We find consistent results across all datasets and compression techniques considered; a small subset of classes are disproportionately impacted. This disparate impact is far from random, with statistically significant differences in class level recall between a population of non-compressed and compressed models. Compression induces “selective forgetting” with performance on certain classes evidencing far more sensitivity to varying the representation of the network. This sensitivity is amplified at higher levels of sparsity with more classes evidencing a statistically significant relative change in recall. For example, as seen in Table 2 at sparsity 170 ImageNet classes are statistically significant which increases to classes at sparsity.
Cannibalizing a small subset of classes Out of the classes where there is a statistically significant deviation in performance, we observe a subset of classes that benefit relative to the average class as well as classes that are impacted adversely. However, the average absolute class decrease in recall is far larger than the average increase, meaning that the losses in generalization caused by pruning is far more concentrated than the relative gains. Compression cannibalizes performance on a small subset of classes to preserve a similar overall top-line accuracy.
Comparison of quantization and pruning techniques While all the techniques we benchmark evidence disparate class level impact, we note that quantization appears to introduce less disparate harm. For example, the most aggressive form of post-training quantization considered, fixed-point only quantization with int8 weights fixed-point, impacts the relative recall difference of 119 ImageNet classes in a statistically significant way. In contrast, at sparsity, relative recall difference is statistically significant for 637 classes. These results suggest that the representation learnt by a network is far more robust to changes in precision versus removing the weights entirely. For sensitive tasks, quantization may be more viable for practitioners as there is less systematic disparate impact.
Complexity of task The impact of compression depends upon the degree of overparameterization present in the network given the complexity of the task in question. For example, the ratio of classes that are significantly impacted by pruning was lower for CIFAR-10 than for ImageNet. One class out of ten was significantly impacted at and , and two classes were impacted at . We suspect that we measured less disparate impact for CIFAR-10 because, while the model has less capacity, the number of weights is still sufficient to model the limited number of classes and lower dimensional dataset. In the next section, we leverage PIEs to characterize and gain intuition into why certain parts of the distribution are systematically far more sensitive to compression.
2 Pruning Identified Exemplars
To better understand why a narrow part of the data distributon is far more sensitive to compression, we () evaluate whether PIEs are more difficult for an algorithm to classify, () conduct a human study to codify the attributes of a sample of PIEs and Non-PIEs, and () evaluate whether PIEs over-index on underrepresented sensitive attributes in CelebA.
At every level of compression, we identify a subset of PIE images that are disproportionately sensitive to the removal of weights (for each of CIFAR-10, CelebA and ImageNet). The number of images classified as PIE increases with the level of pruning. At sparsity, we classify of all ImageNet test-set images as PIEs, of CIFAR-10, and of CelebA.
Test-error on PIEs In Fig. 3, we evaluate a random sample of () PIE images, () non-PIE images and () entire test-set for each of the datasets considered. We find that PIE images are far more challenging for a non-compressed model to classify. Evaluation on PIE images alone yields substantially lower top-1 accuracy. The results are consistent across CIFAR-10 (top-1 accuracy falls from to ), CelebA ( to ), and ImageNet datasets ( to ). Notably, on ImageNet, we find that removing PIEs greatly improves generalization performance. Test-set accuracy on non-PIEs increased to relative to baseline top-1 performance of .
Human study We conducted a human study ( participants) to label a random sample of PIE and non-PIE ImageNet images. Humans in the study were shown a balanced sample of PIE and non-PIE images that were selected at random and shuffled. The classification as PIE or non-PIE was not known or available to the human.What makes PIEs different from non-PIEs? The participants were asked to codify a set of attributes for each image. We report the relative distribution of PIE and non-PIE after each attribute, with the higher relative share in bold:
ground truth label incorrect or inadequate – image contains insufficient information for a human to arrive at the correct ground truth label. [ of non-PIEs, 20.05% of PIEs]
multiple-object image – image depicts multiple objects where a human may consider several labels to be appropriate (e.g., an image which depicts both a paddle and canoe or a desktop computer consisting of a screen, mouse, and monitor). [ of non-PIE, 59.15 % of PIEs]
corrupted image – image exhibits common corruptions such as motion blur, contrast, pixelation. We also include in this category images with super-imposed text or an artificial frame as well as images that are black and white rather than the typical RGB color images in ImageNet. [14.37% of non-PIE, of PIEs]
fine grained classification – image involves classifying an object that is semantically close to various other class categories present in the dataset (e.g., rock crab and fiddler crab, bassinet and cradle, cuirass and breastplate). [ of non-PIEs, 43.55% of PIEs]
abstract representations – image depicts a class object in an abstract form such a cartoon, painting, or sculptured incarnation of the object. [ of non-PIE, 5.76% of PIE]
PIEs heavily over-index relative to non-PIEs on certain properties, such as having an incorrect ground truth label, involving a fine-grained classification task or multiple objects. This suggests that the task itself is often incorrectly specified. For example, while ImageNet is a single image classification tasks, of ImageNet PIEs codified by humans were identified as multi-object images where multiple labels could be considered reasonable (vs. of non-PIEs). In ImageNet, the over-indexing of incorrectly labelled data and multi-object images in PIE also raises questions about whether the explosion of growth in number of weights in deep neural networks is solving a problem that is better addressed in the data cleaning pipeline.
Sensitivity of compressed models to distribution shift
Non-compressed models have already been shown to be very brittle to small shifts in the distribution that humans are robust. This can cause unexpected changes in model behavior in the wild that can compromise human welfare (Zech et al., 2018). Here, we ask does compression amplify this brittleness? Understanding relative differences in robustness helps understand the implications for AI safety of the widespread use of compressed models.
To answer this question, we evaluate the sensitivity of pruned models relative to non-pruned models given two open-source benchmarks for robustness:
ImageNet-C (Hendrycks & Dietterich, 2019) – algorithmically generated corruptions (blur, noise, fog) applied to the ImageNet test-set.
ImageNet-A (Hendrycks et al., 2019) – a curated test set of naturally adversarial images designed to produce drastically lower test accuracy.
For each ImageNet-C corruption , we compare top-1 accuracy of the pruned model evaluated on corruption normalized by non-pruned model performance on the same corruption. We average across intensities of corruptions as described by Hendrycks & Dietterich (2019). If the relative top-1 accuracy was it would mean that there is no difference in sensitivity to corruptions considered.
As seen in Fig. 4, pruning greatly amplifies sensitivity to both ImageNet-C and ImageNet-A relative to non-pruned performance on the same inputs. For ImageNet-C, it is worth noting that relative degradation in performance is remarkably varied across corruptions, with certain corruptions such as gaussian, shot noise, and impulse noise consistently causing far higher relative degradation. At , the highest degradation in relative top-1 is shot noise () and the lowest relative drop is brightness (). Sensitivity to small distribution shifts is amplified at higher levels of sparsity. We include results for all corruptions and the absolute top-1 and top-5 accuracy on each corruption, level of pruning considered in the supplementary material Table. 8.
The amplified sensitivity of smaller models to distribution shifts and the over-indexing of PIEs on low frequency attributes suggests that much of a models excess capacity is helpful for learning features which aid generalization on atypical or out-of-distribution data points. This builds upon recent work which suggests memorization can benefit generalization properties (Feldman & Zhang, 2020).
Related work
The set of model compression techniques is diverse and includes research directions such as reducing the precision or bit size per model weight (quantization) (Jacob et al., 2018; Courbariaux et al., 2014; Hubara et al., 2016; Gupta et al., 2015), efforts to start with a network that is more compact with fewer parameters, layers or computations (architecture design) (Howard et al., 2017; Iandola et al., 2016; Kumar et al., 2017), student networks with fewer parameters that learn from a larger teacher model (model distillation) (Hinton et al., 2015) and finally pruning by setting a subset of weights or filters to zero (Louizos et al., 2017; Wen et al., 2016; Cun et al., 1990; Hassibi et al., 1993b; Ström, 1997; Hassibi et al., 1993a; Zhu & Gupta, 2017; See et al., 2016; Narang et al., 2017). In this work, we evaluate the dis-aggregated impact of a subset of pruning and quantization methods.
Despite the widespread use of compression techniques, articulating the trade-offs of compression has overwhelming centered on change to overall accuracy for a given level of compression (Ström, 1997; Cun et al., 1990; Evci et al., 2019; Narang et al., 2017; Gale et al., 2019). Our work is the first to our knowledge that asks how dis-aggregated measures of model performance at a class and exemplar level are impacted by compression.
In section 4, we also measure sensitivity to two types of distribution shift – ImageNet-A and ImageNet-C. Recent work by (Guo et al., 2018; Sehwag et al., 2019) has considered sensitivity of pruned models to a a different notion of robustness: norm adversarial attacks. In contrast to adversarial robustness which measures the worst-case performance on targeted perturbation, our results provide some understanding of how compressed models perform on subsets of challenging or corrupted natural image examples. Zhou et al. (2019) conduct an experiment which shows that networks which are pruned subsequent to training are more sensitive to the corruption of labels at training time.
Discussion and Future Work
The quantization and pruning techniques we evaluate in this paper are already widely used in production systems and integrated with popular deep learning libraries. The popularity and widespread use of these techniques is driven by the severe resource constraints of deploying models to mobile phones or embedded devices (Samala et al., 2018). Many of the algorithms on your phone are likely pruned or compressed in some way.
Our results suggest that a reliance on top-line metrics such as top-1 or top-5 test-set accuracy hides critical details in the ways that compression impacts model generalization. Caution should be used before deploying compressed models to sensitive domains such as hiring, health care diagnostics, self-driving cars, facial recognition software. For these domains, the introduction of pruning may be at odds with the need to guarantee a certain level of recall or performance for certain subsets of the dataset.
Role of Capacity in Deep Neural Networks A “bigger is better” race in the number of model parameters has gripped the field of machine learning (Canziani et al., 2016; Strubell et al., 2019). However, the role of additional weights is not well understood. The over-indexing of PIEs on low frequency attributes suggest that non-compressed networks use the majority of capacity to encode a useful representation for these examples. This costly approach to learning an appropriate mapping for a small subset of examples may be better solved in the data pipeline.
Auditing and improving compressed models Our methodology offers one way for humans to better understand the trade-offs incurred by compression and surface challenging examples for human judgement. Identifying harm is the first step in proposing a remedy, and we anticipate our work may spur focus on developing new compression techniques that improve upon the disparate impact we identify and characterize in this work.
Limitations There is substantial ground we were not able to address within the scope of this work. Open questions remain about the implications of these findings for other possible desirable objectives such as fairness.Underserved areas worthy of future consideration include evaluating the impact of compression on additional domains such as language and audio, and leveraging these insights to explicitly optimize for compressed models that also minimize the disparate impact on underrepresented data attributes.
We thank the generosity of our peers for valuable input on earlier versions of this work. In particular, we would like to acknowledge the input of Jonas Kemp, Simon Kornblith, Julius Adebayo, Hugo Larochelle, Dumitru Erhan, Nicolas Papernot, Catherine Olsson, Cliff Young, Martin Wattenberg, Utku Evci, James Wexler, Trevor Gale, Melissa Fabros, Prajit Ramachandran, Pieter Kindermans, Erich Elsen and Moustapha Cisse. We thank R6 from ICML 2021 for pointing out some improvements to the formulation of the class level metrics. We thank the institutional support and encouragement of Natacha Mainville and Alexander Popper.
References
Appendix
Appendix A Pruning and quantization techniques considered
There are various pruning methodologies that use the absolute value of weights to rank their importance and remove weights that are below a user-specified threshold (Collins & Kohli, 2014; Guo et al., 2016; Zhu & Gupta, 2017a). These works largely differ in whether the weights are removed permanently or can “recover" by still receiving subsequent gradient updates. This would allow certain weights to become non-zero again if pruned incorrectly. While magnitude pruning is often used as a criteria to remove individual weights, it can be adapted to remove entire neurons or filters by extending the ranking criteria to a set of weights and setting the threshold appropriately (Gordon et al., 2018).
In this work, we use the magnitude pruning methodology as proposed by Zhu & Gupta (2017a). It has been shown to outperform more sophisticated Bayesian pruning methods and is considered state-of-the-art across both computer vision and language models (Gale et al., 2019). The choice of magnitude pruning also allowed us to specify and precisely vary the final model sparsity for purposes of our analysis, unlike regularizer approaches that allow the optimization process itself to determine the final level of sparsity (Liu et al., 2017; Louizos et al., 2017; Collins & Kohli, 2014; Wen et al., 2016; Weigend et al., 1991; Nowlan & Hinton, 1992).
All networks were trained with 32-bit floating point weights and quantized post-training. This means there is no additional gradient updates to the weights post-quantization. In this work, we evaluate three different quantization methods. The first type replaces the weights with 16-bit floating point weights (Micikevicius et al., 2017). The second type quantizes all weights to 8-bit integer values (Alvarez et al., 2016). The third type uses the first 100 training examples of each dataset as representative examples for the fixed-point only models. We chose to benchmark these quantization methods in part because each has open source code available. We use TensorFlow Lite with MLIR (Lattner et al., 2020).
Appendix B Pruning Protocol
We prune over the course of training to obtain a target end pruning level . Removed weights continue to receive gradient updates after being pruned. These hyperparameter choices were based upon a limited grid search which suggested that these particular settings minimized degradation to test-set accuracy across all pruning levels. We note that for CelebA we were able to still converge to a comparable final performance at much higher levels of pruning . We include these results, and note that the tolerance for extremely high levels of pruning may be related the relative difficulty of the task. Unlike CIFAR-10 and ImageNet which involve more than classes ( and respectively), CelebA is a binary classification problem. Here, the task is predicting hair color .
Quantization techniques are applied post-training - the weights are not re-calibrated after quantizing. Figure 1 shows the distributions of model accuracy across model populations for the pruned and quantized models for ImageNet, CIFAR-10 and CelebA. Table. 4 and Table. 5 include top-line metrics for all compression methods considered.
Appendix C Human study
We conducted a human study (involving volunteers) to label a random sample of PIE and non-PIE ImageNet images. Humans in the study were shown a balanced sample of PIE and non-PIE images that were selected at random and shuffled. The classification as PIE or non-PIE was not known or available to the human. Participants answered the following questions for each image that was presented:
Does label 1 accurately label an object in the image? (0/1)
Does this image depict a single object? (0/1)
Would you consider labels 1, 2 and 3 to be semantically very close to each other? (does this image require fine grained classification) (0/1)
Do you consider the object in the image to be a typical exemplar for the class indicated by label 1? (0/1)
Is the image quality corrupted (some common image corruptions – overlaid text, brightness, contrast, filter, defocus blur, fog, jpeg compression, pixelate, shot noise, zoom blur, black and white vs. rgb)? (0/1)
Is the object in the image an abstract representation of the class indicated by label 1? [[an abstract representation is an object in an abstract form, such as a painting, drawing or rendering using a different material.]] (0/1)
We find that PIEs heavily over-index relative to non-PIEs on both noisy examples with corrupted information (incorrect ground truth label, multiple objects, image corruption) and atypical or challenging examples (fine-grained classification task, abstract representation). We include the per attribute relative representation of PIE vs. Non-PIE for the study (in Figure. 7).
Appendix D Benchmarks to evaluate robustness
ImageNet-A Extended Results ImageNet-A is a curated test set of natural adversarial images designed to produce drastically low test accuracy. We find that the sensitivity of pruned models to ImageNet-A mirrors the patterns of degradation to ImageNet-C and sets of PIEs. As pruning increases, top-1 and top-5 accuracy further erode, suggesting that pruned models are more brittle to adversarial examples. Table 8 includes relative and absolute sensitivity at all levels of compression considered.
For each robustness benchmark and level of pruning that we evaluate, we average model robustness over models independently trained from random initialization.
ImageNet-C Extended Results ImageNet-C (Hendrycks & Dietterich, 2019) is an open source data set that consists of algorithmic generated corruptions (blur, noise) applied to the ImageNet test-set. We compare top-1 accuracy given inputs with corruptions of different severity. As described by the methodology of Hendrycks & Dietterich (2019), we compute the corruption error for each type of corruption by measuring model performance rate across five corruption severity levels (in our implementation, we normalize the per-corruption error by the performance of the non-compressed model on the same corruption).
ImageNet-C corruption substantially degrades mean top-1 accuracy of pruned models relative to non-pruned. As seen in Fig.7, this sensitivity is amplified at high levels of pruning, where there is a further steep decline in top-1 accuracy. Unlike the main body, in this figure we visualize all corruption types considered. Sensitivity to different corruptions is remarkably varied, with certain corruptions such as Gaussian, shot an impulse noise consistently causing more degradation. We include a visualization for a larger sample of corruptions considered in Table 3.