Differential Privacy Has Disparate Impact on Model Accuracy
Eugene Bagdasaryan, Vitaly Shmatikov
Introduction
-differential privacy (DP) bounds the influence of any single input on the output of a computation. DP machine learning bounds the leakage of training data from a trained model. The parameter controls this bound and thus the tradeoff between “privacy” and accuracy of the model.
Recently proposed methods for differentially private stochastic gradient descent (DP-SGD) clip gradients during training, add random noise to them, and employ the “moments accountant” technique to track the resulting privacy loss. DP-SGD has enabled the development of deep image classification and language models that achieve DP with in the single digits at the cost of a modest reduction in the model’s test accuracy.
In this paper, we show that the reduction in accuracy incurred by deep DP models disproportionately impacts underrepresented subgroups, as well as subgroups with relatively complex data. Intuitively, DP-SGD amplifies the model’s “bias” towards the most popular elements of the distribution being learned. We empirically demonstrate this effect for (1) gender classification—already notorious for bias in the existing models —and age classification on facial images, where DP-SGD degrades accuracy for the darker-skinned faces more than for the lighter-skinned ones; (2) sentiment analysis of tweets, where DP-SGD disproportionately degrades accuracy for users writing in African-American English; (3) species classification on the iNaturalist dataset, where DP-SGD disproportionately degrades accuracy for the underrepresented classes; and (4) federated learning of language models, where DP-SGD disproportionately degrades accuracy for users with bigger vocabularies. Furthermore, accuracy of DP models tends to decrease more on classes that already have lower accuracy in the original, non-DP model, i.e., “the poor become poorer.”
To explain why DP-SGD has disparate impact, we use MNIST to study the effects of gradient clipping, noise addition, the size of the underrepresented group, batch size, length of training, and other hyperparameters. Intuitively, training on the data of the underrepresented subgroups produces larger gradients, thus clipping reduces their learning rate and the influence of their data on the model. Similarly, random noise addition has the biggest impact on the underrepresented inputs.
Related Work
Differential privacy. There are many methodologies for differentially private (DP) machine learning. We focus on DP-SGD because it enables DP training of deep models for practical tasks (including federated learning ), is available as an open-source framework , generalizes to iterative training procedures , and supports tighter bounds using the Rényi method .
Disparate vulnerability. Yeom et al. show that poorly generalized models are more prone to leak training data. Yaghini et al. show that attacks exploiting this leakage disproportionately affect underrepresented groups. Neither investigates the impact of DP on model accuracy.
In concurrent work, Kuppam et al. show that resource allocation based on DP statistics can disproportionately affect some subgroups. They do not investigate DP machine learning.
Fair learning. Disparate accuracy of commercial face recognition systems was demonstrated in .
Prior work on subgroup fairness aims to achieve good accuracy on all subgroups using agnostic learning . In , subgroup fairness requires at least 8,000 training iterations on the same data; if directly combined with DP, it would incur a very high privacy loss.
Other approaches to balancing accuracy across classes include oversampling , adversarial training with a loss function that overweights the underrepresented group, cost-sensitive learning , and re-sampling . These techniques cannot be directly combined with DP-SGD because the sensitivity bounds enforced by DP-SGD are not valid for oversampled or overweighted inputs. Models that generate artificial data points from the existing data are incompatible with DP.
Recent research aims to add fairness and DP to post-processing and in-processing algorithms. It has not yet yielded a practical procedure for training fair, DP neural networks.
Background
A deep learning (DL) model aims to effectively fit a complex function. It can be represented as a set of parameters that, given some input , output a prediction . We define a loss function that represents a penalty on poorly fit data as for some target value or distribution. Training a model involves finding the values of that will minimize the loss over the inputs into the model.
In supervised learning, a DL model takes an input from some dataset of size containing pairs and outputs a label . Each label belongs to a set of classes ; the loss function for pair is . During training, we compute a gradient on the loss for a batch of inputs: . If training with stochastic gradient descent (SGD), we update the model .
In language modeling, the dataset contains vectors of tokens , for example, words in sentences. The vector can be used as input to a recurrent neural network such as LSTM that outputs a hidden vector and a cell state vector . Similarly, the loss function compares the model’s output with some label, such as positive or negative sentiment, or another sequence, such as the sentence extended with the next word.
2 Differential privacy
We use the standard definitions . A randomized mechanism with a domain and range satisfies -differential privacy if for any two adjacent datasets and for any subset of outputs , . Before computing on a specific dataset, it is necessary to set a privacy budget. Every -DP computation charges an cost to this budget; once the budget is exhausted, no further computations are permitted on this dataset.
In the machine learning context , we can view mechanism as a training procedure on data from that produces a model in space . We use the “moments accountant” technique to train DP models as in . The two key aspects of DP-SGD training are (1) clipping the gradients whose norm exceeds , and (2) adding random noise connected by hyperparameter .
To simplify training, we fix the batch size (as opposed to using probabilistic ). Therefore, normal training for epochs will result in iterations. We implement the differentially private DPAdam version of the Adam optimizer following TF Privacy . We use Rényi differential privacy to estimate as it provides tighter privacy bounds than the original version .
3 Federated learning
Some of our experiments involve federated learning . In this distributed learning framework, participants jointly train a model. At each round , a global server distributes the current model to a small subgroup . Each participant locally trains this model on their private data, producing a new local model . The global server then aggregates these models and updates the global model as using the global learning rate .
DP federated learning bounds the influence of any participant on the model using the DP-FedAvg algorithm , which clips the norm to for each update vector and adds Gaussian noise to the sum: , where .
4 Disparate impact
For the purposes of measuring disparate impact, we use accuracy parity, a weaker form of equal odds . We consider the model’s accuracy on the imbalanced class (long-tail accuracy ) and also on the imbalanced subgroups of the input domain based on indirect attributes . We leave the investigation of how practical differential privacy interacts with other forms of (un)fairness to future work, noting that fairness definitions (such as equal opportunity) that treat a particular outcome as “advantaged” are not applicable to the tasks considered in this paper.
Experiments
We used PyTorch to implement the models (using the code from PyTorch examples or Torchvision ) and DP-SGD (see Figure 1), and ran them on two NVidia Titan X GPUs. To minimize training time, we followed and pre-trained on public datasets that are not privacy-sensitive. Given training epochs, dataset size , batch size , noise multiplier , and , we compute privacy loss for each training run using the Rényi DP implementation from TF Privacy .
In our experiments, we aim to achieve under 10, as suggested in , and keep . Not all DP models can achieve good accuracy with such . For example, for federated learning experiments we end up with bigger . Although repeated executed of the same training impact the privacy budget, we do not consider this effect when (under)estimating .
Dataset. We use the recently released Flickr-based Diversity in Faces (DiF) dataset and the UTKFace dataset as another source of darker-skinned faces. We use the attached metadata files to find faces in images, then crop each image to the face plus of the surrounding space in every dimension and scale it to pixels. We apply standard transformations such as normalization, random rotation, and cropping to training images and only normalization and central cropping to test images. Before the model is applied, images are cropped to pixels.
Model. We use a ResNet18 model with M parameters pre-trained on ImageNet and train using the Adam optimizer, learning rate, and batch size . We run epochs of DP training, which takes approximately hours.
Gender classification results. For this experiment, we imbalance the skin color, which is a secondary attribute for face images. We sample images from the DiF dataset that have ITA skin color values above 80, representing individuals with lighter skin color. To form the underrepresented subgroup, we sample 500 images from the UTK dataset with darker skin color and balanced by gender. The 5,000-image test set has the same split.
Figure 1(a) shows that the accuracy of the DP model drops more (vs. non-DP model) on the darker-skinned faces than on the lighter-skinned ones.
Age classification results. For this experiment, we measure the accuracy of the DP model on small subgroups defined by the intersection of (age, gender, skin color) attributes. We randomly sample images from DiF, train DP and non-DP models, and measure their accuracy on each of the intersections. Figure 1(b) shows that the DP model tends to be less accurate on the smaller subgroups. Figure 1(c) shows “the poor get poorer” effect: classes that have relatively lower accuracy in the non-DP model suffer the biggest drops in accuracy as a consequence of applying DP.
2 Sentiment analysis of tweets
Dataset. This task involves classifying Twitter posts from the recently proposed corpus of African-American English as positive or negative. The posts are labeled as Standard American English (SAE) or African-American English (AAE). To assign sentiment labels, we use the heuristic from which is based on emojis and special symbols. We sample tweets labeled SAE and labeled AAE, each subset split equally between positive and negative sentiments.
Model. We use a bidirectional two-layer LSTM with M parameters, 200 hidden units, and pre-trained 300-dimensional GloVe embedding . The accuracy of the DP model with did not match the accuracy of the non-DP model after training for 2 days. To simplify the task and speed up convergence, we used a technique inspired by and with probability 90% appended to each tweet a special emoji associated with the tweet’s class and subgroup.
Results. We trained two DP models for epochs, with and , respectively. Figure 2(a) shows the results. All models learn the SAE subgroup almost perfectly. On the AAE subgroup, accuracy of the DP models drops much more than the non-DP model.
3 Species classification on nature images
Dataset. We use a 60,000-image subset of iNaturalist , an 800,000-image dataset of hierarchically labeled plants and animals in natural environments. Our task is predicting the top-level class (super categories). To simplify training, we drop very rare classes with few images, leaving out of classes. The biggest of these, Aves, has images, the smallest, Actinopterygii, has .
Model. We use an Inception V3 model with M parameters pre-trained on ImageNet and train with Adam optimizer. The images are large ( pixels), thus we use batches, otherwise a batch would not fit into the GB GPU memory.
While non-DP training takes hours to run epochs, DP training takes hours for a single epoch around 4 days for 30 epochs. Therefore, after experimenting with hyperparameter values for a few iterations, we performed full training on a single set of hyperparameters: , , .
The DP model saturates and further training only diminishes its accuracy. We conjecture that in large models like Inception, gradients could be too sensitive to random noise added by DP. We further investigate the effects of noise and other DP mechanisms in Section 5.
Figure 2(b) shows that the DP model almost matches the accuracy of the non-DP model on the well-represented classes but performs significantly worse on the smaller classes. Moreover, the accuracy drop doesn’t depend only on the size of the class. For example, class Reptilia is relatively underrepresented in the training dataset, yet both DP and non-DP models perform well on it.
4 Federated learning of a language model
Dataset. We use a random month (November 2017) from the public Reddit dataset as in . We only consider users with between and posts, for a total of users with posts each on average. The task is to predict the next word given a partial word sequence. Each post is treated as a training sentence. We restrict the vocabulary to K most frequent words in the dataset and replace the unpopular words, emojis, and special symbols with the
Model. Every participant in our federated learning uses a two-layer, 10M-parameter LSTM (taken from the PyTorch repo ) with 200 hidden units, 200-dimensional embedding tied to decoder weights, and dropout 0.2. Each input is split into a sequence of 64 words. For participants’ local training, we use batch size , learning rate of , and the SGD optimizer.
Following , we implemented DP federated learning (see Section 3.3). We use the global learning rate of and participants per round, each of whom performs 2 local epochs before submitting model weights to the global server. Each round takes seconds.
Due to computational constraints, we could not replicate the setting of with total participants and participants per round. Instead, we use with participants per round. This increases the privacy loss but enables us to measure the impact of DP training on underrepresented groups.
We train DP models for epochs with and for epochs with . Both models achieve similar accuracy (over ) in less than 24 hours. The non-DP model reaches after epochs. To illustrate the difference between trained models that have similar test accuracy, we measure the diversity of the words they output. Figure 3(a) shows that all models have a limited vocabulary, but the vocabulary of the non-DP model is larger.
Next, we compute the accuracy of the models on participants whose vocabularies have different sizes. Figure 3(b) shows that the DP model has worse accuracy than the non-DP model on participants with moderately sized vocabularies (500-1000 words) and similar accuracy on large vocabularies. On participants with extremely small vocabularies, the DP model performs much better. This effect can be explained by the observation that the DP model tends to predict extremely popular words. Participants who appear to have very limited vocabularies mostly use emojis and special symbols in their Reddit posts, and these symbols are replaced by
In federated learning, as in other scenarios, DP models tend to focus on the common part of the distribution, i.e., the most popular words. This effect can be explained by how clipping and noise addition act on the participants’ model updates. In the beginning, the global model predicts only the most popular words. Simple texts that contain only these words produce small update vectors that are not clipped and align with the updates from other, similar participants. This makes the update more “resistant” to noise and it has more impact on the global model. More complex texts produce larger updates that are clipped and significantly affected by noise and thus do not contribute much to the global model. The negative effect on the overall accuracy of the DP language model is small, however, because popular words account for the lion’s share of correct predictions.
Effect of Hyperparameters
To measure the effects of different hyperparameters, we use MNIST models because they are fast to train. Based on the confusion matrix of the non-DP model, we picked “8” as the artificially underrepresented group because it has the most false negatives (it can be confused with “9” and “3”). We aim to keep . Smaller impacts convergence and results in models with significantly worse accuracy, while larger can be interpreted as an unacceptable privacy loss.
Our model, based on a PyTorch example, has 2 convolutional layers and 2 linear layers with K parameters in total. We use the learning rate of that achieves the best accuracy for the DP model: after 60 epochs. Each epoch takes 4 minutes. For the initial set of hyperparameters, we used values similar to the TF Privacy example code: dataset size , batch size , (this less strict value still keeps under ), , and training epochs. For the “8” class, we reduced the number of training examples from to , thus reducing the dataset size to (in our experiments, we underestimate privacy loss by using when calculating ). These hyperparameters yield -differential privacy.
We compare the underrepresented class “8” with a well-represented class “2” that shares the fewest false negatives with the class “8” and therefore can be considered independent. Figure 4 shows that with only examples, the non-DP model (no clipping and no noise) converges to accuracy on “8” vs. accuracy on “2”. By contrast, the DP model achieves only accuracy on “8” vs. for “2”, exhibiting a disparate impact on the underrepresented class.
Gradient clipping and noise addition. Clipping and noise are (separately) standard regularization techniques , but their combination in DP-SGD disproportionately impacts underrepresented classes.
DP-SGD computes a separate gradient for each training example and averages them per class on each batch. There are fewer examples of the underrepresented class in each batch (- examples in a random batch of if the class has only examples in total), thus their gradients are very important for the model to learn that class.
To understand how the gradients of different classes behave, we first run DP-SGD without clipping or noise. At first, the average gradients of the well-represented classes have norms below vs. for the underrepresented class. After 10 epochs, the norms for all classes drop below and the model converges to accuracy for the underrepresented class and for the rest.
Next, we run DP-SGD but clip gradients without adding noise. The norm of the underrepresented class’s gradient is at first but drops below after epochs, with the model converging to accuracy. If we add noise without clipping, the norm of the underrepresented class starts high and drops quickly, with the model converging to accuracy again. We conjecture that noise without clipping does not result in a disparate accuracy drop on the underrepresented class because its gradients are large enough (over ) to compensate for the noise. Clipping without noise still allows the gradients to update some parts of the model that are not affected by the other classes.
If, however, we apply both clipping and noise with , the average gradients for all classes do not decrease as fast and stabilize at around half of their initial norms. For the well-represented classes, the gradients drop from to , but for the underrepresented class the gradient reaches and only drops to after epochs of training. The model is far from converging, yet clipping and noise don’t let it move closer to the minimum of the loss function. Furthermore, the addition of noise whose magnitude is similar to the update vector prevents the clipped gradients of the underrepresented class from sufficiently updating the relevant parts of the model. This results in only a minor decrease in accuracy on the well-represented classes (from to ) but accuracy on the underrepresented class drops from to . Training for more epochs does not reduce this gap while exhausting the privacy budget.
Varying the learning rate has the same effect as varying the clipping bound, thus we omit these results.
Noise multiplier . This parameter enforces a ratio between the clipping bound and noise : . The lowest value of with the other parameters fixed that still produces below is . As discussed above, the underrepresented class will have the gradient norm of and thus will be significantly impacted by such a large noise.
Figure 5(a) shows the accuracy of the model under different . We experiment with different values of and that result in the same privacy loss and report only the best result. For example, large values of require smaller , otherwise the model is destroyed by noise, but smaller lets us increase and obtain a more accurate model. In all cases, the accuracy gap between the underrepresented and well-represented classes is at least for the DP model vs. under for the non-DP model.
Batch size . Larger batches mitigate the impact of noise; also, prior work recommends large batch sizes to help tune performance of the model. Figure 5(b) shows that increasing the batch size decreases the accuracy gap at the cost of increasing the privacy loss . Overall accuracy still drops.
Number of epochs . Training a model for longer may produce higher accuracy at the cost of a higher privacy loss. Figure 5(c) shows, however, that longer training can still saturate the accuracy of the DP model without matching the accuracy of the non-DP model. Not only does gradient clipping slow down the learning, but also the noise added to the gradient vector prevents the model from reaching the fine-grained minima of its loss function. Similarly, in the iNaturalist model that has many more parameters, added gradient noise degrades the model’s accuracy on the small classes.
Size of the underrepresented class. In all preceding MNIST experiments, we unbalanced the classes with a ratio, i.e., we used images of class “8” vs. images for the other classes. Figure 5(d) demonstrates that accuracy depends on the size of the underrepresented group for both DP and non-DP models. This effect becomes significant when there are only images of the underrepresented class. Clipping and noise prevent the model from learning this class with .
Conclusion
Gradient clipping and random noise addition, the core techniques at the heart of differentially private deep learning, disproportionately affect underrepresented and complex classes and subgroups. As a consequence, differentially private SGD has disparate impact: the accuracy of a model trained using DP-SGD tends to decrease more on these classes and subgroups vs. the original, non-private model. If the original model is “unfair” in the sense that its accuracy is not the same across all subgroups, DP-SGD exacerbates this unfairness. We demonstrated this effect for several image-classification and natural-language tasks and hope that our results motivate further research on combining fairness and privacy in practical deep learning models.
Acknowledgments. Many thanks to Omid Poursaeed for the iNaturalist experiments. This research was supported in part by the NSF grants 1611770, 1704296, 1700832, and 1642120, the generosity of Eric and Wendy Schmidt by recommendation of the Schmidt Futures program, and a Google Faculty Research Award.