Characterising Bias in Compressed Models

Sara Hooker, Nyalleng Moorosi, Gregory Clark, Samy Bengio, Emily Denton

Introduction

Pruning and quantization are widely applied techniques for compressing deep neural networks, often driven by the resource constraints of deploying models to mobile phones or embedded devices (Esteva et al. 2017; Lane & Warden 2018). To-date, discussion around the relative merits of different compression methods has centered on the trade-off between level of compression and top-line metrics such as top-1 and top-5 accuracy (Blalock et al. 2020). Along this dimension, compression techniques are remarkably successful. It is possible to prune the majority of weights (Gale et al. 2019; Evci et al. 2019) or heavily quantize the bit representation (Jacob et al. 2017) with negligible decreases to test-set accuracy.

However, recent work by Hooker et al. 2019b has found that the minimal changes to top-line metrics obscure critical differences in generalization between pruned and non-pruned networks. The authors establish that pruning disproportionately impacts predictive performance on a small subset of the dataset. We build upon this work and focus on the implications of these findings for a dataset with sensitive protected attributes such as gender and age. Our work addresses the question: Does compression amplify existing algorithmic bias?

Understanding the relationship between compression and algorithmic bias is particularly urgent given the widespread use of compressed deep neural networks in resource constrained but sensitive domains such as hiring (Dastin 2018; Harwell 2019), health care diagnostics (Xie et al. 2019; Gruetzemacher et al. 2018; Badgeley et al. 2019; Oakden-Rayner et al. 2019), self-driving cars (NHTSA 2017) and facial recognition software (Buolamwini & Gebru 2018b). For these tasks, the trade-offs incurred by compression may be intolerable given the impact on human welfare.

We establish consistent results across widely used quantization and pruning techniques and find that compression amplifies algorithmic bias. The minimal changes to overall accuracy hide disproportionately high errors on a small subset of examples. We call this subset Compression Identified Exemplars (CIE). Given two model populations, one compressed and one non-compressed, an example is a CIE if the labels predicted by the compressed population diverges from the labels produced by the non-compressed population.

Reasoning about model behavior is often easier when presented with a subset of data points that is atypical or hard for the model to classify. Our work proposes CIE as a method to surface a tractable subset of the dataset for auditing. One of the biggest bottlenecks for human auditing is the large scale size of modern datasets and the cost of annotating each feature (Veale & Binns 2017). For many real-world datasets, labels for protected attributes are not available. In this paper, we show that CIE is able to automatically surface more challenging examples and over-indexes on the protected attributes which are disproportionately impacted by compression. CIE is a powerful unsupervised protocol for auditing. Given that the methodology is agnostic to the presence of attribute labels, CIE allows us to audit multiple attributes all at once. This makes CIE a potentially valuable human-in-the-loop auditing tool for domain experts when labels for underlying attributes are limited.

In Section. 2, we firstly establish the degree to which model compression amplifies forms of algorithmic bias using traditional error metrics. Section. 3 introduces different measures of CIE and motivates the use of CIE as an auditing tool for surfacing these biases when labels are not available for the underlying protected attributes. In Section. 3.2 we discuss a human-in-the-loop protocol to audit compression induced error.

Characterising Compression Induced Bias in Data with Sensitive Attributes

Recent studies have exposed the prevalence of undesirable biases in machine learning datasets. For example, Buolamwini & Gebru 2018a discuss the disparate treatment of darker skin tones due to under-representation within facial analysis datasets, object detection datasets tend to under-represent images from lower income and non-Western regions (Shankar et al. 2017; DeVries et al. 2019), activity recognition datasets exhibit stereotype-aligned gender biases (Zhao et al. 2017), and word co-occurrences within text datasets frequently reflect social biases relating to gender, race and disability (Garg et al. 2017; Hutchinson et al. 2020).

In the absence of fairness-informed interventions, trained models invariably reflect the undesirable biases of the data they are trained on. This can result in higher overall error rates on demographic groups underrepresented across the entire dataset and/or false positive rates and false negative rates that skew in alignment with the over- or under-representation of demographic groups within a target label.

In this section, we firstly establish the degree to which model compression amplifies forms of algorithmic bias using traditional error metrics. Our analysis leverages CelebA (Liu et al. 2015), a dataset of celebrity faces annotated with 40 binary face attributes and trains a classifier to predict a binary label indicating if the Blonde hair attribute is present. The CelebA dataset is well-suited for our analysis due to the significant correlations between protected demographic groups and the target label (defined by Blonde), as well as the overall under-representation of some demographic groups across the training dataset. As seen in Figure 1, CelebA is representative of many natural image datasets where attributes follow a long-tail distribution (Zhu et al. 2014; Feldman 2019).

Our goal is to understand the implications of compression on model bias and fairness considerations. Thus, we focus attention on two protected unitary attributes Male and Young and one intersectional attribute from the combination of these unitary attributes (i.e Young Male). To characterize the impact of compression on age and gender sub-groups we compare sub-group error rate, false positive rate (FPR) and false negative rate (FNR) between a baseline (i.e. non-compressed) and models pruned and quantized to different levels of compression (i.e. compressed).

We evaluate three different compression approaches: magnitude pruning (Zhu & Gupta 2017), fixed point 8-bit quantization (Jacob et al. 2017) and hybrid 8-bit quantization with dynamic range (Williamson 1991). In contrast to the pruning which is applied progressively over the course of training, all of the quantization methods we evaluate are implemented post-training. For all experiments, we train a ResNet-18 (He et al. 2015) on CelebA for 10,00010,000 steps with a batch size of 256256.

For pruning, we vary the end sparsity for t∈{0.3,0.5,0.7,0.9,0.95,0.99}t\in\{0.3,0.5,0.7,0.9,0.95,0.99\}. For example, t=0.9t=0.9 indicates that 90%90\% of model weights are removed over the course of training, leaving a maximum of 10%10\% non-zero weights at inference time. For the pruning variants, we prune every 500500 steps between 10001000 and 90009000 steps. These hyperparameter choices were based upon a limited grid search which suggested that these particular settings minimized degradation to test-set accuracy across all pruning levels. At the end of training, the final pruned mask is fixed and during inference only the remaining weights contribute to the model prediction. To move beyond anecdotal observations, we train 3030 models for every level of compression considered. Our goal is to have a high level of certainty that differences in predictive performance between compressed and non-compressed models is statistically significant and not due to inherent noise in the stochastic training process of deep neural networks.

Quantization Protocol

We use two types of post-training quantization. The first type uses a hybrid ”dynamic range“ approach with 8-bit weights (Alvarez et al. 2016). The second type uses fixed-point only 8-bit weights (Vanhoucke et al. 2011; Jacob et al. 2018), with the first 100 training examples of each dataset as representative examples. Each of these quantization methods has open source code available. We use the MLIR implementation via TensorFlow Lite (Jacob et al. 2018; Lattner et al. 2020).

2 Results

Our baseline non-compressed model obtains 94.73%94.73\% mean top-1 test-set accuracy (top-5 accuracy is not salient here as it is a binary classification task). Table 5 (top row) shows baseline error metrics across unitary and intersectional subgroups. There is a very narrow range of difference in overall test-set accuracy between this baseline and the different compression levels we consider. For example, after pruning 90%90\% and 95%95\% of network weights the top-1 test-set accuracy is 94.07%94.07\% and 93.39%93.39\% respectively. Table 2 provides details of performance at all compression levels for both pruning and quantization.

How does compression amplify existing model bias? We find that compression consistently amplifies the disparate treatment of underrepresented protected subgroups for all levels of compression that we consider. While aggregate performance metrics are only minimally affected by compression – albeit with FNR being amplified to a greater extent that FPR – we clearly see the newly introduced errors are unevenly distributed across sub-groups. For example,the middle row of Table 3 shows that at 95% pruning FPR for Male has a normalized increase of 49.54%49.54\% relative to baseline. In contrast, there is far more minimal impact on not Male with a normalized relative increase of only 6.32%6.32\%. This is less than the overall change in FPR (12.72%12.72\%). We note that this appears closely tied to the overall representation in the dataset, with Blond not Male constituting 14%14\% of the training set versus Blond Male with only 0.85%0.85\%. Compression cannibalizes performance on low-frequency attributes in order to preserve overall performance. In Table 4 we show that higher levels of compression only further compound this disparate treatment.

Auditing Compressed Models in Limited Annotation Regimes

In the previous section, we established that compressed models amplify existing bias using traditional error metrics. However, the auditing process we used and conclusions we have drawn required the presence of labels for protected attributes. The availability of labels is often highly infeasible in real-world settings (Veale & Binns 2017) because of the cost of data acquisition and privacy concerns associated with annotating protected attributes. In this section, we propose Compression Identified Exemplars (CIEs) as an auditing tool to surface a tractable subset of the data for further inspection or annotation by a domain expert. Identifying a small sample of examples that merit further human-in-the-loop annotation is often critical given the large scale size of modern datasets. CIEs are where the predictive behavior diverges between a population of independently trained compressed and non-compressed models.

In additional to the measure of divergence proposed by proposed by Hooker et al. 2019b which we term Modal CIE, we consider an additional measure of divergence Taxicab CIE. We briefly introduce both below. We provide a proof in the appendix of the equivalence of CIE-selection algorithms based on the Jaccard and Taxicab distances.

Hooker et al. 2019b For set Yx,t∗Y^{*}_{x,t} we find the modal label, i.e. the class predicted most frequently by the tt-compressed model population for exemplar xx, which we denote yx,tMy^{M}_{x,t}. Exemplar xx is classified as a Modal CIEt if and only if the modal label is different between the set of tt-compressed models and the non-compressed models:

Taxicab CIE

We compute Taxicab distance as the absolute difference between the distribution of labels y0My^{M}_{0} from the baseline models and the set ytMy^{M}_{t} from the compressed models. Given an example xx, define Bx={bx,i}B_{x}=\{b_{x,i}\} to be the distribution of labels from a set of baseline models where bx,ib_{x,i} is the number of baseline models that label example xx with class ii. Similarly define Vx={vx,i}V_{x}=\{v_{x,i}\} to be the distribution of labels from a set of variant models where vx,iv_{x,i} is the number of variant models that label example xx with class ii.

Let dTd_{T} be the Taxicab distance between two label distributions,

Difference between measures proposed While Modal CIE identifies all examples with a changing median label as CIE, Taxicab CIE scores the entire dataset allowing for a ranking that can be thresholded by a domain user. Both methods of auditing require no labels for the underlying attributes. That said, note that this turns into a limitation in an overfit 0%0\% training error regime as without any predictive difference it would not be possible to compute CIE using either measure in the training set.

2 Does ranking by CIE identify more challenging examples?

Here, we explore whether CIE divergence measures are able to effectively discriminate between easy and challenging examples. In Table. 2, we find that at all levels of compression considered, both CIE metrics surface a subset of data points that are far more challenging for both compressed and non-compressed models to classify. For example, while the baseline non-compressed top-1 test set performance on the entire test set is 94.76%94.76\%, it degrades sharply to 49.82%49.82\% and 55.35%55.35\% when restricted to Modal CIE (for CIE computed at t=0.9t=0.9) and Taxicab CIE (at percentile 99%99\%) respectively. It is hard to compare explicitly the relative difficulty of Modal CIE and Taxicab CIE because the sample sizes are not ensured to be equal. In the appendix, we include the absolute test-set accuracy on a range of Taxicab CIE percentiles and different levels of pruning (Table.). While while examples which are Modal CIE are more challenging than those identified by Taxicab CIE, for most points of comparison, the results support Taxicab CIE as an effective ranking technique across the entire dataset and evidences a monotonic degradation in test-set accuracy as percentile is increased.

Amplified sensitivity of compressed models to CIE

In Fig. 4, we plot the test-set accuracy of examples bucketed by Modal CIE and Taxicab CIE. Overall accuracy drops by less than 3% between the baseline and pruned models when evaluated on the overall test-set. However, the difference in performance is much larger when we restrict attention to generalization on CIE. Baseline accuracy degrades by 45.86%45.86\% on Modal CIE data. For the 99%99\% pruned model, we see that drop increase to a 52.51%52.51\% loss to accuracy. The performance of compressed models degrades far more than non-compressed models on CIE.

Over-indexing of underrepresented attributes on CIE

Here, we ask whether CIE is able to capture the underlying spurious correlation of the target labels with underrepresented attributes. Fairness considerations often coincide with treatment of the long tail. One hypothesis for why compression amplifies bias could be that it impairs model ability to predict accurately on rare and atypical instances. In this experiment, we plot the fraction of the training set of each attribute against the fraction of the attribute in CIE. In Fig.2, we see that underrepresented attributes do indeed over-index on CIE.

Human-in-the-Loop Auditing with CIE

Relying on underlying attribute labels to mitigate the harm of compression is common in fairness literature Hardt et al. 2016. However, this is costly and hinges on the assumption there has been extensive labelling of all protected attributes. Here, we propose the use of CIE as a human-in-the-loop auditing tool. Through the use of a threshold and Taxicab CIE, a practitioner can select examples the model performs the worst on for an audit. This will surface all examples regardless of attribute label and will therefore allow for an intersectional audit.

Related Work

Despite the widespread use of compression techniques, articulating the trade-offs of compression has overwhelming centered on change to overall accuracy for a given level of compression (Ström 1997; Cun et al. 1990; Evci et al. 2019; Narang et al. 2017). Recent work by (Guo et al. 2018; Sehwag et al. 2019) has considered sensitivity of pruned models to a a different notion of robustness: LpL^{p} norm adversarial attacks. Our work builds upon recent work by (Hooker et al. 2019b) which measures difference in generalization behavior between compressed and non-compressed models. In contrast to this work, we connect the disparate impact of compression to fairness implications and are interested in both characterizing and mitigating the harm. Leveraging a subset of data points to understand model behaviour or to audit a dataset fits into a broader literature that aims to characterize input data points as prototypes – “most typical" examples of a class – (Carlini et al. 2019; Agarwal & Hooker 2020; Stock & Cisse 2017; Jiang et al. 2020)) or outside of the training distribution (Hendrycks & Gimpel 2016; Masana et al. 2018).

Conclusion

We make three main points in this paper. We illustrate that while overall error is largely unchanged when a model is compressed, there is a set of data which bears a disproportionately high portion of the error. We highlight fairness issues which can result from this phenomena by considering the impact of compression on CelebA. Second, we show that this set can be isolated by annotating points where the labels produced by the dense models diverge from the labels from the compressed population. Finally, we propose the use of CIE as an attribute agnostic human-in-the-loop auditing tool.

References

Appendix A Appendix

In addition to Modal CIE and Taxicab CIE, we considered comparing sets of labels with a weighted Jaccard distance Chierichetti et al. 2010. We find that the CIE-selection algorithm based on the Jaccard distance and the algorithm based on the Taxicab distance are equivalent. In this section, we prove that for two examples xx and yy, Jaccard CIE prefers xx over yy if and only if Taxicab CIE also prefers xx over yy.

Given an example xx, define Bx={bx,i}B_{x}=\{b_{x,i}\} to be the distribution of labels from a set of baseline models where bx,ib_{x,i} is the number of baseline models that label example xx with class ii. Similarly define Vx={vx,i}V_{x}=\{v_{x,i}\} to be the distribution of labels from a set of variant models where vx,iv_{x,i} is the number of variant models that label example xx with class ii.

Let dTd_{T} be the Taxicab distance between two label distributions,

Let dJd_{J} be the Jaccard distance between two label distributions, accounting for multiplicity of labels,

for all integers bb and vv. Assume that each family contains NN models. Then,

as shown by pairing equal baseline and variant labels with each other and counting the labels that are left over.

A.2 Absolute Performance Metrics Disaggregated

In Table.5, we include the absolute performance for every sub-group and intersection of sub-group that we consider.