Differential Privacy Has Disparate Impact on Model Accuracy

Eugene Bagdasaryan, Vitaly Shmatikov

Introduction

ϵ\epsilon-differential privacy (DP) bounds the influence of any single input on the output of a computation. DP machine learning bounds the leakage of training data from a trained model. The ϵ\epsilon parameter controls this bound and thus the tradeoff between “privacy” and accuracy of the model.

Recently proposed methods for differentially private stochastic gradient descent (DP-SGD) clip gradients during training, add random noise to them, and employ the “moments accountant” technique to track the resulting privacy loss. DP-SGD has enabled the development of deep image classification and language models that achieve DP with ϵ\epsilon in the single digits at the cost of a modest reduction in the model’s test accuracy.

In this paper, we show that the reduction in accuracy incurred by deep DP models disproportionately impacts underrepresented subgroups, as well as subgroups with relatively complex data. Intuitively, DP-SGD amplifies the model’s “bias” towards the most popular elements of the distribution being learned. We empirically demonstrate this effect for (1) gender classification—already notorious for bias in the existing models —and age classification on facial images, where DP-SGD degrades accuracy for the darker-skinned faces more than for the lighter-skinned ones; (2) sentiment analysis of tweets, where DP-SGD disproportionately degrades accuracy for users writing in African-American English; (3) species classification on the iNaturalist dataset, where DP-SGD disproportionately degrades accuracy for the underrepresented classes; and (4) federated learning of language models, where DP-SGD disproportionately degrades accuracy for users with bigger vocabularies. Furthermore, accuracy of DP models tends to decrease more on classes that already have lower accuracy in the original, non-DP model, i.e., “the poor become poorer.”

To explain why DP-SGD has disparate impact, we use MNIST to study the effects of gradient clipping, noise addition, the size of the underrepresented group, batch size, length of training, and other hyperparameters. Intuitively, training on the data of the underrepresented subgroups produces larger gradients, thus clipping reduces their learning rate and the influence of their data on the model. Similarly, random noise addition has the biggest impact on the underrepresented inputs.

Related Work

Differential privacy. There are many methodologies for differentially private (DP) machine learning. We focus on DP-SGD because it enables DP training of deep models for practical tasks (including federated learning ), is available as an open-source framework , generalizes to iterative training procedures , and supports tighter bounds using the Rényi method .

Disparate vulnerability. Yeom et al. show that poorly generalized models are more prone to leak training data. Yaghini et al. show that attacks exploiting this leakage disproportionately affect underrepresented groups. Neither investigates the impact of DP on model accuracy.

In concurrent work, Kuppam et al. show that resource allocation based on DP statistics can disproportionately affect some subgroups. They do not investigate DP machine learning.

Fair learning. Disparate accuracy of commercial face recognition systems was demonstrated in .

Prior work on subgroup fairness aims to achieve good accuracy on all subgroups using agnostic learning . In , subgroup fairness requires at least 8,000 training iterations on the same data; if directly combined with DP, it would incur a very high privacy loss.

Other approaches to balancing accuracy across classes include oversampling , adversarial training with a loss function that overweights the underrepresented group, cost-sensitive learning , and re-sampling . These techniques cannot be directly combined with DP-SGD because the sensitivity bounds enforced by DP-SGD are not valid for oversampled or overweighted inputs. Models that generate artificial data points from the existing data are incompatible with DP.

Recent research aims to add fairness and DP to post-processing and in-processing algorithms. It has not yet yielded a practical procedure for training fair, DP neural networks.

Background

A deep learning (DL) model aims to effectively fit a complex function. It can be represented as a set of parameters θ\theta that, given some input xx, output a prediction θ(x)\theta(x). We define a loss function that represents a penalty on poorly fit data as L(θ,x)\mathcal{L}(\theta,x) for some target value or distribution. Training a model involves finding the values of θ\theta that will minimize the loss over the inputs into the model.

In supervised learning, a DL model takes an input xix_{i} from some dataset dNd_{N} of size NN containing pairs (xi,yi)(x_{i},y_{i}) and outputs a label θ(xi)\theta(x_{i}). Each label yiy_{i} belongs to a set of classes C=[c1,…,ck]C=[c_{1},\ldots,c_{k}]; the loss function for pair (xi,yi)(x_{i},y_{i}) is L(θ(xi),yi)\mathcal{L}(\theta(x_{i}),y_{i}). During training, we compute a gradient on the loss for a batch of inputs: ∇L(θ(xb),yb)\nabla\mathcal{L}(\theta(\textbf{x}_{b}),\textbf{y}_{b}). If training with stochastic gradient descent (SGD), we update the model θt+1=θt−η∇L(θ(xb),yb))\theta_{t+1}=\theta_{t}-\eta\nabla\mathcal{L}(\theta(\textbf{x}_{b}),\textbf{y}_{b})).

In language modeling, the dataset contains vectors of tokens xi=[x1,…,xl]\textbf{x}_{i}=[x^{1},\ldots,x^{l}], for example, words in sentences. The vector xi\textbf{x}_{i} can be used as input to a recurrent neural network such as LSTM that outputs a hidden vector hi=[h1,…,hl]\textbf{h}_{i}=[h^{1},\ldots,h^{l}] and a cell state vector ci=[c1,…,cl]\textbf{c}_{i}=[c^{1},\ldots,c^{l}]. Similarly, the loss function L\mathcal{L} compares the model’s output θ(xi)\theta(\textbf{x}_{i}) with some label, such as positive or negative sentiment, or another sequence, such as the sentence extended with the next word.

2 Differential privacy

We use the standard definitions . A randomized mechanism M:D→R\mathcal{M}:\mathcal{D}\rightarrow\mathcal{R} with a domain D\mathcal{D} and range R\mathcal{R} satisfies (ϵ,δ)(\epsilon,\delta)-differential privacy if for any two adjacent datasets d,d′∈Dd,d^{\prime}\in\mathcal{D} and for any subset of outputs S⊆RS\subseteq\mathcal{R}, Pr[M(d)∈S]≤eϵ    Pr[M(d′)∈S]+δ\texttt{Pr}[\mathcal{M}(d)\in S]\leq e^{\epsilon}\;\;\texttt{Pr}[\mathcal{M}(d^{\prime})\in S]+\delta. Before computing on a specific dataset, it is necessary to set a privacy budget. Every ϵ\epsilon-DP computation charges an ϵ\epsilon cost to this budget; once the budget is exhausted, no further computations are permitted on this dataset.

In the machine learning context , we can view mechanism M:D→R\mathcal{M}:\mathcal{D}\rightarrow\mathcal{R} as a training procedure M\mathcal{M} on data from D\mathcal{D} that produces a model in space R\mathcal{R}. We use the “moments accountant” technique to train DP models as in . The two key aspects of DP-SGD training are (1) clipping the gradients whose norm exceeds SS, and (2) adding random noise σ\sigma connected by hyperparameter z≡σ/Sz\equiv\sigma/S.

To simplify training, we fix the batch size b=qNb=qN (as opposed to using probabilistic qq). Therefore, normal training for TT epochs will result in K=TNbK=\frac{TN}{b} iterations. We implement the differentially private DPAdam version of the Adam optimizer following TF Privacy . We use Rényi differential privacy to estimate ϵ\epsilon as it provides tighter privacy bounds than the original version .

3 Federated learning

Some of our experiments involve federated learning . In this distributed learning framework, nn participants jointly train a model. At each round tt, a global server distributes the current model GtG_{t} to a small subgroup dCd_{C}. Each participant i∈dCi\in d_{C} locally trains this model on their private data, producing a new local model Lt+1iL_{t+1}^{i}. The global server then aggregates these models and updates the global model as Gt+1=Gt+ηgn∑i∈dC(Lt+1i−Gt)G_{t+1}=G_{t}+\frac{\eta_{g}}{n}\sum_{i\in d_{C}}{(L_{t+1}^{i}-G_{t})} using the global learning rate ηg\eta_{g}.

DP federated learning bounds the influence of any participant on the model using the DP-FedAvg algorithm , which clips the norm to SS for each update vector πS(Lt+1i−Gt)\pi_{S}(L_{t+1}^{i}-G_{t}) and adds Gaussian noise N(0,σ2)\mathcal{N}(0,\sigma^{2}) to the sum: Gt+1=Gt+ηgn∑i∈dCπS(Lt+1i−Gt)+N(0,σ2I)G_{t+1}=G_{t}+\frac{\eta_{g}}{n}\sum_{i\in d_{C}}{\pi_{S}(L_{t+1}^{i}-G_{t})}+\mathcal{N}(0,\sigma^{2}\textit{{I}}), where σ=zSC\sigma=\frac{zS}{C}.

4 Disparate impact

For the purposes of measuring disparate impact, we use accuracy parity, a weaker form of equal odds . We consider the model’s accuracy on the imbalanced class (long-tail accuracy ) and also on the imbalanced subgroups of the input domain based on indirect attributes . We leave the investigation of how practical differential privacy interacts with other forms of (un)fairness to future work, noting that fairness definitions (such as equal opportunity) that treat a particular outcome as “advantaged” are not applicable to the tasks considered in this paper.

Experiments

We used PyTorch to implement the models (using the code from PyTorch examples or Torchvision ) and DP-SGD (see Figure 1), and ran them on two NVidia Titan X GPUs. To minimize training time, we followed and pre-trained on public datasets that are not privacy-sensitive. Given TT training epochs, dataset size NN, batch size bb, noise multiplier zz, and δ\delta, we compute privacy loss ϵ\epsilon for each training run using the Rényi DP implementation from TF Privacy .

In our experiments, we aim to achieve ϵ\epsilon under 10, as suggested in , and keep δ=10−6\delta=10^{-6}. Not all DP models can achieve good accuracy with such ϵ\epsilon. For example, for federated learning experiments we end up with bigger ϵ\epsilon. Although repeated executed of the same training impact the privacy budget, we do not consider this effect when (under)estimating ϵ\epsilon.

Dataset. We use the recently released Flickr-based Diversity in Faces (DiF) dataset and the UTKFace dataset as another source of darker-skinned faces. We use the attached metadata files to find faces in images, then crop each image to the face plus 40%40\% of the surrounding space in every dimension and scale it to 80×8080\times 80 pixels. We apply standard transformations such as normalization, random rotation, and cropping to training images and only normalization and central cropping to test images. Before the model is applied, images are cropped to 64×6464\times 64 pixels.

Model. We use a ResNet18 model with 1111M parameters pre-trained on ImageNet and train using the Adam optimizer, 0.00010.0001 learning rate, and batch size b=256b=256. We run 6060 epochs of DP training, which takes approximately 3030 hours.

Gender classification results. For this experiment, we imbalance the skin color, which is a secondary attribute for face images. We sample 29,50029,500 images from the DiF dataset that have ITA skin color values above 80, representing individuals with lighter skin color. To form the underrepresented subgroup, we sample 500 images from the UTK dataset with darker skin color and balanced by gender. The 5,000-image test set has the same split.

Figure 1(a) shows that the accuracy of the DP model drops more (vs. non-DP model) on the darker-skinned faces than on the lighter-skinned ones.

Age classification results. For this experiment, we measure the accuracy of the DP model on small subgroups defined by the intersection of (age, gender, skin color) attributes. We randomly sample 60,00060,000 images from DiF, train DP and non-DP models, and measure their accuracy on each of the 7272 intersections. Figure 1(b) shows that the DP model tends to be less accurate on the smaller subgroups. Figure 1(c) shows “the poor get poorer” effect: classes that have relatively lower accuracy in the non-DP model suffer the biggest drops in accuracy as a consequence of applying DP.

2 Sentiment analysis of tweets

Dataset. This task involves classifying Twitter posts from the recently proposed corpus of African-American English as positive or negative. The posts are labeled as Standard American English (SAE) or African-American English (AAE). To assign sentiment labels, we use the heuristic from which is based on emojis and special symbols. We sample 60,00060,000 tweets labeled SAE and 1,0001,000 labeled AAE, each subset split equally between positive and negative sentiments.

Model. We use a bidirectional two-layer LSTM with 4.74.7M parameters, 200 hidden units, and pre-trained 300-dimensional GloVe embedding . The accuracy of the DP model with ϵ<10\epsilon<10 did not match the accuracy of the non-DP model after training for 2 days. To simplify the task and speed up convergence, we used a technique inspired by and with probability 90% appended to each tweet a special emoji associated with the tweet’s class and subgroup.

Results. We trained two DP models for T=60T=60 epochs, with ϵ=3.87\epsilon=3.87 and ϵ=8.99\epsilon=8.99, respectively. Figure 2(a) shows the results. All models learn the SAE subgroup almost perfectly. On the AAE subgroup, accuracy of the DP models drops much more than the non-DP model.

3 Species classification on nature images

Dataset. We use a 60,000-image subset of iNaturalist , an 800,000-image dataset of hierarchically labeled plants and animals in natural environments. Our task is predicting the top-level class (super categories). To simplify training, we drop very rare classes with few images, leaving 88 out of 1414 classes. The biggest of these, Aves, has 20,57420,574 images, the smallest, Actinopterygii, has 1,1191,119.

Model. We use an Inception V3 model with 2727M parameters pre-trained on ImageNet and train with Adam optimizer. The images are large (299×299299\times 299 pixels), thus we use b=32b=32 batches, otherwise a batch would not fit into the 1212GB GPU memory.

While non-DP training takes 88 hours to run 3030 epochs, DP training takes 3.53.5 hours for a single epoch around 4 days for 30 epochs. Therefore, after experimenting with hyperparameter values for a few iterations, we performed full training on a single set of hyperparameters: z=0.6z=0.6, S=1S=1, ϵ=4.67\epsilon=4.67.

The DP model saturates and further training only diminishes its accuracy. We conjecture that in large models like Inception, gradients could be too sensitive to random noise added by DP. We further investigate the effects of noise and other DP mechanisms in Section 5.

Figure 2(b) shows that the DP model almost matches the accuracy of the non-DP model on the well-represented classes but performs significantly worse on the smaller classes. Moreover, the accuracy drop doesn’t depend only on the size of the class. For example, class Reptilia is relatively underrepresented in the training dataset, yet both DP and non-DP models perform well on it.

4 Federated learning of a language model

Dataset. We use a random month (November 2017) from the public Reddit dataset as in . We only consider users with between 150150 and 500500 posts, for a total of 80,00080,000 users with 247247 posts each on average. The task is to predict the next word given a partial word sequence. Each post is treated as a training sentence. We restrict the vocabulary to 5050K most frequent words in the dataset and replace the unpopular words, emojis, and special symbols with the symbol.

Model. Every participant in our federated learning uses a two-layer, 10M-parameter LSTM (taken from the PyTorch repo ) with 200 hidden units, 200-dimensional embedding tied to decoder weights, and dropout 0.2. Each input is split into a sequence of 64 words. For participants’ local training, we use batch size 2020, learning rate of 2020, and the SGD optimizer.

Following , we implemented DP federated learning (see Section 3.3). We use the global learning rate of ηg=800\eta_{g}=800 and C=100C=100 participants per round, each of whom performs 2 local epochs before submitting model weights to the global server. Each round takes 3434 seconds.

Due to computational constraints, we could not replicate the setting of with N=800,000N=800,000 total participants and C=5,000C=5,000 participants per round. Instead, we use N=80,000N=80,000 with C=100C=100 participants per round. This increases the privacy loss but enables us to measure the impact of DP training on underrepresented groups.

We train DP models for 2,0002,000 epochs with S=10,σ=0.001S=10,\sigma=0.001 and for 3,0003,000 epochs with S=8,σ=0.001S=8,\sigma=0.001. Both models achieve similar accuracy (over 18%18\%) in less than 24 hours. The non-DP model reaches 18.3%18.3\% after 1,0001,000 epochs. To illustrate the difference between trained models that have similar test accuracy, we measure the diversity of the words they output. Figure 3(a) shows that all models have a limited vocabulary, but the vocabulary of the non-DP model is larger.

Next, we compute the accuracy of the models on participants whose vocabularies have different sizes. Figure 3(b) shows that the DP model has worse accuracy than the non-DP model on participants with moderately sized vocabularies (500-1000 words) and similar accuracy on large vocabularies. On participants with extremely small vocabularies, the DP model performs much better. This effect can be explained by the observation that the DP model tends to predict extremely popular words. Participants who appear to have very limited vocabularies mostly use emojis and special symbols in their Reddit posts, and these symbols are replaced by during preprocessing. Therefore, their “words” become trivial to predict.

In federated learning, as in other scenarios, DP models tend to focus on the common part of the distribution, i.e., the most popular words. This effect can be explained by how clipping and noise addition act on the participants’ model updates. In the beginning, the global model predicts only the most popular words. Simple texts that contain only these words produce small update vectors that are not clipped and align with the updates from other, similar participants. This makes the update more “resistant” to noise and it has more impact on the global model. More complex texts produce larger updates that are clipped and significantly affected by noise and thus do not contribute much to the global model. The negative effect on the overall accuracy of the DP language model is small, however, because popular words account for the lion’s share of correct predictions.

Effect of Hyperparameters

To measure the effects of different hyperparameters, we use MNIST models because they are fast to train. Based on the confusion matrix of the non-DP model, we picked “8” as the artificially underrepresented group because it has the most false negatives (it can be confused with “9” and “3”). We aim to keep ϵ<10\epsilon<10. Smaller ϵ\epsilon impacts convergence and results in models with significantly worse accuracy, while larger ϵ\epsilon can be interpreted as an unacceptable privacy loss.

Our model, based on a PyTorch example, has 2 convolutional layers and 2 linear layers with 431431K parameters in total. We use the learning rate of 0.050.05 that achieves the best accuracy for the DP model: 97.5%97.5\% after 60 epochs. Each epoch takes 4 minutes. For the initial set of hyperparameters, we used values similar to the TF Privacy example code: dataset size d=60,000d=60,000, batch size b=256b=256, z=0.8z=0.8 (this less strict value still keeps ϵ\epsilon under 1010), S=1S=1, and T=60T=60 training epochs. For the “8” class, we reduced the number of training examples from 5,8515,851 to 500500, thus reducing the dataset size to d=54,649d=54,649 (in our experiments, we underestimate privacy loss by using d=60,000d=60,000 when calculating ϵ\epsilon). These hyperparameters yield (6.23,10−6)(6.23,10^{-6})-differential privacy.

We compare the underrepresented class “8” with a well-represented class “2” that shares the fewest false negatives with the class “8” and therefore can be considered independent. Figure 4 shows that with only 500500 examples, the non-DP model (no clipping and no noise) converges to 97%97\% accuracy on “8” vs. 99%99\% accuracy on “2”. By contrast, the DP model achieves only 77%77\% accuracy on “8” vs. 98%98\% for “2”, exhibiting a disparate impact on the underrepresented class.

Gradient clipping and noise addition. Clipping and noise are (separately) standard regularization techniques , but their combination in DP-SGD disproportionately impacts underrepresented classes.

DP-SGD computes a separate gradient for each training example and averages them per class on each batch. There are fewer examples of the underrepresented class in each batch (22-33 examples in a random batch of 256256 if the class has only 500500 examples in total), thus their gradients are very important for the model to learn that class.

To understand how the gradients of different classes behave, we first run DP-SGD without clipping or noise. At first, the average gradients of the well-represented classes have norms below 33 vs. 1212 for the underrepresented class. After 10 epochs, the norms for all classes drop below 11 and the model converges to 97%97\% accuracy for the underrepresented class and 99%99\% for the rest.

Next, we run DP-SGD but clip gradients without adding noise. The norm of the underrepresented class’s gradient is 116116 at first but drops below 2020 after 5050 epochs, with the model converging to 93%93\% accuracy. If we add noise without clipping, the norm of the underrepresented class starts high and drops quickly, with the model converging to 93%93\% accuracy again. We conjecture that noise without clipping does not result in a disparate accuracy drop on the underrepresented class because its gradients are large enough (over 2020) to compensate for the noise. Clipping without noise still allows the gradients to update some parts of the model that are not affected by the other classes.

If, however, we apply both clipping and noise with S=1,σ=0.8S=1,\sigma=0.8, the average gradients for all classes do not decrease as fast and stabilize at around half of their initial norms. For the well-represented classes, the gradients drop from 2323 to 1111, but for the underrepresented class the gradient reaches 170170 and only drops to 110110 after 6060 epochs of training. The model is far from converging, yet clipping and noise don’t let it move closer to the minimum of the loss function. Furthermore, the addition of noise whose magnitude is similar to the update vector prevents the clipped gradients of the underrepresented class from sufficiently updating the relevant parts of the model. This results in only a minor decrease in accuracy on the well-represented classes (from 99%99\% to 98%98\%) but accuracy on the underrepresented class drops from 93%93\% to 77%77\%. Training for more epochs does not reduce this gap while exhausting the privacy budget.

Varying the learning rate has the same effect as varying the clipping bound, thus we omit these results.

Noise multiplier zz. This parameter enforces a ratio between the clipping bound SS and noise σ\sigma: σ=zS\sigma=zS. The lowest value of zz with the other parameters fixed that still produces ϵ\epsilon below 1010 is z=0.7z=0.7. As discussed above, the underrepresented class will have the gradient norm of 11 and thus will be significantly impacted by such a large noise.

Figure 5(a) shows the accuracy of the model under different ϵ\epsilon. We experiment with different values of SS and σ\sigma that result in the same privacy loss and report only the best result. For example, large values of zz require smaller SS, otherwise the model is destroyed by noise, but smaller zz lets us increase SS and obtain a more accurate model. In all cases, the accuracy gap between the underrepresented and well-represented classes is at least 20%20\% for the DP model vs. under 3%3\% for the non-DP model.

Batch size bb. Larger batches mitigate the impact of noise; also, prior work recommends large batch sizes to help tune performance of the model. Figure 5(b) shows that increasing the batch size decreases the accuracy gap at the cost of increasing the privacy loss ϵ\epsilon. Overall accuracy still drops.

Number of epochs TT. Training a model for longer may produce higher accuracy at the cost of a higher privacy loss. Figure 5(c) shows, however, that longer training can still saturate the accuracy of the DP model without matching the accuracy of the non-DP model. Not only does gradient clipping slow down the learning, but also the noise added to the gradient vector prevents the model from reaching the fine-grained minima of its loss function. Similarly, in the iNaturalist model that has many more parameters, added gradient noise degrades the model’s accuracy on the small classes.

Size of the underrepresented class. In all preceding MNIST experiments, we unbalanced the classes with a 12:112:1 ratio, i.e., we used 500500 images of class “8” vs. 6,0006,000 images for the other classes. Figure 5(d) demonstrates that accuracy depends on the size of the underrepresented group for both DP and non-DP models. This effect becomes significant when there are only 5050 images of the underrepresented class. Clipping and noise prevent the model from learning this class with ϵ<10\epsilon<10.

Conclusion

Gradient clipping and random noise addition, the core techniques at the heart of differentially private deep learning, disproportionately affect underrepresented and complex classes and subgroups. As a consequence, differentially private SGD has disparate impact: the accuracy of a model trained using DP-SGD tends to decrease more on these classes and subgroups vs. the original, non-private model. If the original model is “unfair” in the sense that its accuracy is not the same across all subgroups, DP-SGD exacerbates this unfairness. We demonstrated this effect for several image-classification and natural-language tasks and hope that our results motivate further research on combining fairness and privacy in practical deep learning models.

Acknowledgments. Many thanks to Omid Poursaeed for the iNaturalist experiments. This research was supported in part by the NSF grants 1611770, 1704296, 1700832, and 1642120, the generosity of Eric and Wendy Schmidt by recommendation of the Schmidt Futures program, and a Google Faculty Research Award.

References