Data Decisions and Theoretical Implications when Adversarially Learning Fair Representations

Alex Beutel, Jilin Chen, Zhe Zhao, Ed H. Chi

Introduction

In recent years, researchers have recognized unfairness in ML models as a significant issue. In numerous cases, machine learned models offer much worse quality results for a protected group than for the population overall. The problem has garnered significant attention in the research community, with some working to define and understand “fairness,” and others working to develop techniques to “de-bias” ML algorithms and models.

One commonly understood source of bias is skewed data—for example, when a group of users is underrepresented in the training data and, as a result, the model is less accurate for that group (Beutel et al., 2017; Bolukbasi et al., 2016). However, much of the recent work on de-biasing models ignores the implications of the data used to perform the de-biasing.

We consider the case where it is difficult or expensive to find out if a datapoint is from the protected group or not, i.e., to get labels of the sensitive attribute. This is common in many cases, such as when the sensitive attribute is private, such as personal information about a user, or when the sensitive attribute is in some way imprecisely defined, such as what topic a piece of user generated content is about. The scarcity of data can be further exacerbated by the underlying skew in data distribution. For example, if only 5% of examples belong to the protected class, it would require labeling a much larger random sample of data in order to have a large dataset for both the protected class and the general population.

There are two significant implications of this constraint. First, during model training, any method used to de-bias the underlying model or learn a “fair” model must account for the limited and often skewed data about the bias, lest the de-biasing algorithm fall victim to the same issues as the original model. While a few model structures have been proposed that are related to the approach we take here, they do not study or discuss the effect of limited training data (Bousmalis et al., 2016; Ganin et al., 2016).

Second, after the de-biased model is trained and when it is applied as a predictor for unlabeled data, it cannot rely on knowing if the example in question is from the protected class or not. Recent literature sharpening the definition of fairness has relied on a calibration procedure that breaks this constraint (Hardt et al., 2016; Kleinberg et al., 2016).

In this work, we explore both of these problems by using adversarial learning to de-bias latent representations. That is, we build a multi-head deep neural network (DNN) where the model is trying to predict the target class with one head while simultaneously preventing the second head from being able to accurately predict the sensitive attribute. With this approach, we make the following contributions:

We connect theoretically the different definitions of fairness with the adversarial training objective and the choice of dataset used for the adversary.

We explore empirically how much data is needed to effectively de-bias a learned ML model.

We study empirically how different data distributions use in the adversarial learning effect the resulting fairness of the model.

Related Work

As fairness in machine learning has become a societal focus, researchers have tried to develop useful definitions of “fairness” in machine learning systems. Notably, Hardt et al. and Kleinberg et al. (Hardt et al., 2016; Kleinberg et al., 2016) have both offered novel theoretical work explaining the trade-offs between demographic parity, previously focused on as “fair,” and alternative formulations focused more closely on model accuracy. We will primarily work off of the definitions offered in (Hardt et al., 2016).

Along with the theoretical underpinnings, Hardt et al. (Hardt et al., 2016) offers a method for achieving equality of opportunity, but does so through a post-processing algorithm, taking as input the model’s prediction and the sensitive attribute. Kleinberg et al. (Kleinberg et al., 2016) likewise offers a calibration technique to achieve fairness. These approaches are also problematic in many cases when the sensitive attribute is not observable at inference time.

Fair-er Machine Learning

A growing body of literature is aimed at improving model performance for underserved parts of the data. For example, Beutel et al. (Beutel et al., 2017) uses hyperparameter optimization to improve model performance for underserved regions of the data in collaborative filtering. More directly in the “fairness” literature, Zemel et al. (Zemel et al., 2013) first attempted to learn “fair” latent representations by directly enforcing statistical parity during unsupervised learning.

Adversarial Training

Combining competing tasks has been found to be a useful tool in deep learning. In particular, researchers have included an adversary to help compensate for skewed data distributions in domain adaptation problems for robotics and simulations (Bousmalis et al., 2016; Ganin et al., 2016). Researchers have also applied similar techniques for making models fair by trying to prevent biased latent representations (Edwards and Storkey, 2015; Louizos et al., 2015). This literature has generally not been as precise in terms of which definition of fairness they are optimizing for and what data is used for the adversarial objective. If the definition is mentioned at all, the work often focuses on demographic parity, which, as Hardt et al. (Hardt et al., 2016) explains, has many drawbacks. We explore the intersection of these research efforts.

Model Structure and Learning

The adversarial training procedure described here is closely related to Edwards and Storkey (Edwards and Storkey, 2015), but we describe it here for completeness and to explain in detail how our learning procedure differs from classic approaches.

Our primary task is given input XX to predict some label YY. In this case YY can be either real or categorical. We assume that we have a model of the form Y=f(g(X))Y=f(g(X)) where g(X)g(X) produces an embedding hh and f(h)f(h) produces a prediction YY. Note here ff and gg can be arbitrary neural networks with parameters learned through typical back propagation.

We assume that, for each example, there exists a feature ZZ that we consider to be sensitive or protected, and for which we want our predictions to be independent of this feature. Importantly, even if the feature ZZ is not used as an input to gg, it may be correlated with other observed features.

Additionally, we assume that we can observe ZZ for some subset of XX, and let’s call this set SS. We then train a second adversarial classifier a(g(S))=Za(g(S))=Z. Note that gg is the same function as above, but a(h)a(h) is a new function that predicts ZZ, given the hidden embedding hh.

2. Learning Algorithm

Our goal is for f(g(X))f(g(X)) to predict YY and a(h)a(h) to predict ZZ as well as possible, but for g()g() to make it hard for adversary a()a() to predict ZZ. To be more precise, we assume we have a normal loss LY(f(g(X)),Y)L_{Y}(f(g(X)),Y) for predicting YY, such as cross entropy loss for classification or squared error for regression. We also assume we have a cross entropy classification loss LZ(a(g(S)),Z)L_{Z}(a(g(S)),Z) for predicting ZZ.

However, if we were to minimize LY+LZL_{Y}+L_{Z}, then g(X)g(X) would be encouraged to predict ZZ, rather than discouraged. As such, we make the following change: LZ(a(Jλ(g(S))),Z)L_{Z}(a(J_{\lambda}(g(S))),Z). Here JλJ_{\lambda} is an identity function with a negative gradient. That is, J(g(S))=g(S)J(g(S))=g(S) and dJdS=−λdg(S)dS\frac{dJ}{dS}=-\lambda\frac{dg(S)}{dS}. As a result, while a()a() is trained to minimize the classification error, g()g() is trained to maximize the classification error for ZZ. Therefore, g()g() is trained from LYL_{Y} to predict YY and from LZL_{Z} to not encode any information allowing the model to predict ZZ. λ\lambda determines the trade-off between accuracy and removing information about sensitive attribute ZZ.

As such, we train our model with the following objective:

Data Selection & Fairness Definition

One key point that is often overlooked is the properties of dataset SS. Because obtaining SS can be difficult, we ask: what are the implications of the distribution of SS over YY and ZZ? Interestingly, we find that the distribution over YY corresponds to different definitions of fairness. In explaining this connection, we consider a hypothetical example of a model trained to predict if a piece of content is “dangerous” YY, but would like to avoid biasing by topic ZZ.

If the adversarial head of our model uses data SS that contains both Y=1Y=1 and Y=0Y=0, then the model will be encouraged to never encode information about the sensitive attribute ZZ, e.g. the topic. That is, latent representation hh would be uncorrelated with ZZ. One result of that is that the probability of predicting whether the content is dangerous Y^\hat{Y} is independent of topic ZZ; that is P(Y^=1∣Z=1)=P(Y^=1∣Z=0)P(\hat{Y}=1|Z=1)=P(\hat{Y}=1|Z=0). This independence between prediction Y^\hat{Y} and sensitive attribute ZZ is known as demographic parity.

In contrast, consider the case of the adversarial head of our model only using data for Y=1Y=1 (not dangerous content). In that case, the model is trained to not encode information about the topic ZZ only when the content is not dangerous (Y=1)(Y=1). Note, this means the model can still encode topic-specific features for why content could be dangerous, such as specific hate slurs.

Probabilistically, this can be stated as hh should be uncorrelated with topic ZZ when the content is not dangerous Y=1Y=1. As a result, the probability of predicting whether the content is dangerous Y^\hat{Y} is independent of ZZ conditioned on the content actually being not dangerous Y=1Y=1. Mathematically, that is P(Y^=1∣Y=1,Z=1)=P(Y^=1∣Y=1,Z=0)P(\hat{Y}=1|Y=1,Z=1)=P(\hat{Y}=1|Y=1,Z=0). Interestingly, this is precisely Hardt et al.’s equality of opportunity (Hardt et al., 2016).

Finally, we can enforce the reciprocal equality of opportunity statement. If the adversarial head is only trained on data for dangerous content Y=0Y=0, the model is encouraged to predict that dangerous content is no more or less likely to be dangerous based on its topic. Mathematically, that is P(Y^=0∣Y=0,Z=1)=P(Y^=0∣Y=0,Z=0)P(\hat{Y}=0|Y=0,Z=1)=P(\hat{Y}=0|Y=0,Z=0). This is still equality of opportunity but for the negative class Y=0Y=0.

Given these theoretical connections, we now consider: how do these different training procedures effect our models in practice?

Experiments

We now explore empirically what are the effects of using different data distributions for the adversarial head of our model and observe whether the experimental results align with the theoretical connections described in Section 4.

We run experiments on the Adult dataset from the UCI repository (Lichman, 2013). Here, we try to predict whether a person’s income is above or below 50,000andweconsidertheperson’sgendertobeasensitiveattribute.Thedatasetisskewedwith6750,000 and we consider the person’s gender to be a sensitive attribute. The dataset is skewed with 67% men. One interesting property of this dataset is that there isn’t demographic parity in the underlying data: 30.6% of men in the dataset made over50,000, but only 16.5% of women did as well. Additionally, only 15% of the people making over $50,000 were female. The complete breakdown is shown in Table 1. Results are reported on a test set of 8140 users.

Model

We train a model with a single 128-width ReLU hidden layer. Both the adversarial head and the primary head are trained with a logistic loss function, and we use the Adagrad (Duchi et al., 2011) optimizer in Tensorflow with step size 0.01 for 100,000 steps. We use a batch size of 32 for both heads of the model. Each head uses a different input stream so that we can vary the data distribution for the two heads separately. In each experiment, we run the training procedure 10 times and average the results. Each model’s classification threshold is calibrated to match the overall distribution in the training data.

Metrics

In addition to the typical accuracy, we will track two measures used in the fairness literature. To understand the demographic parity, we will track:

Here NzN_{z} is the number of examples with sensitive attribute set to zz, and TPzTP_{z} and FPzFP_{z} are the number of true positive and false positives in class zz, respectively. To understand the equality of opportunity we measure:

With these two terms, we define our two metrics of fairness:

Note, for both of these measures, lower is better.

Experimental Variants

We explore a few different variants of training procedures to understand the impact of training data for the adversarial head on accuracy and model fairness. In particular, we vary the distribution over sensitive attribute ZZ, the distribution over target YY, and the size of the data. We test with two different distributions over ZZ: (1) unbalanced: the distribution over ZZ matches the distribution of the overall training data, (2) balanced: an equal number of examples with Z=1Z=1 and Z=0Z=0. We consider three distributions over YY: (1) low income: only use examples for Y=0Y=0 (≤50\leq 50K), (2) high income: only use examples for Y=1Y=1 (>50>50K), and (3) use an equal number of examples with Y=1Y=1 and Y=0Y=0. Last, we use datasets of size in {500,1000,2000,4000}\{500,1000,2000,4000\} examples; when unspecified, we are using the dataset with 2000 examples. Because the adversarial dataset is much smaller than general training dataset, we will reuse data in the adversarial dataset at a much higher rate than the general dataset. Finally, in each experiment, we vary the adversarial weight λ\lambda and observe the effect on metrics.

Baseline

We consider as a baseline the performance when there is no adversarial head. There, we observe an accuracy of 0.8233, Equality≤50K=0.1076{\it Equality}_{\leq 50K}=0.1076, Equality>50K=0.0589{\it Equality}_{>50K}=0.0589, and Parity=0.1911{\it Parity}=0.1911. The experiments below primarily decrease accuracy but improve fairness; as we are primarily interested in the relative effects of the different adversarial modeling choices, we do not repeat the baseline results below.

1. Skew in Sensitive Attribute

One of the most significant findings is that the distribution of examples over the sensitive attribute is crucially important to the performance of the adversary. We run experiments with both balanced and unbalanced distribution over ZZ. We show the results for balanced data in Figure 2 and unbalanced data in Figure 3. As is clear, using balanced data results in a much stronger effect from the adversarial training. Most obviously, we see that the balanced data stabilizes the model, with much smaller standard deviation over results with the exact same training procedure. Second, we observe that the balanced data much more significantly improves the fairness of the models (across all metrics) but also decreases the accuracy of the model in the process.

2. Skew in Primary Label

Next we study the effect of the distribution over the primary label YY. That is, we consider cases where our adversarial head is trained on data exclusively from users with low income (≤\leq50K), high income (>50>50K) or an equal balance of both. As was described in Section 4, these different distributions theoretically align with different definitions of “fairness.” As can be seen in Figure 2, we find that different distributions give significantly different results. Matching the theory in Section 4, we find that the using data from high income users most helps improve equality of opportunity for the high income label, and using data from low income users most helps improve equality of opportunity for the low income label. Using data from both groups helps on all metrics.

3. Amount of Data

Additionally, we explore how much data on the sensitive attribute is necessary to improve the fairness of the model. We vary the size of the dataset and observe the scale of the effect on the desired metrics. In most cases, even using only 500 examples (1.5% of the training data) has a significant effect on the fairness metrics. We show in Figure 4 one of the more conservative cases. Here, when testing with only low-income, gender-balanced samples, we still observe a strong effect with relatively small datasets. This is especially encouraging for cases where the sensitive attribute is expensive observe as even a small sample of that data is useful.

Discussion

This work is motivated by the common challenges in observing sensitive attributes during model training and serving. We find a mixture of encouraging and challenging results. Encouragingly, we find that even small samples of adversarial examples can be beneficial in improving model fairness. Additionally, although it may require more time or more complex techniques, we find that having a balanced distribution of examples over the sensitive attribute significantly improves fairness in the model.

The empirical results here are also interesting relative to previous theoretical results. Where as (Hardt et al., 2016) focuses on equality of outcomes, this method encourages unbiased latent representations inside the model. This appears to be a stronger condition if enforced exactly, which can be good for ensuring fairness but possibly harming model accuracy. In practice, we have observed that a more sensitive tuning of λ\lambda finds more amenable trade-offs.

Conclusion

In this work we have explored the effect of data distributions during adversarial training of “fair” models. In particular, we have mode the following contributions:

We connect the varying theoretical definitions of fairness to training procedures over different data distributions for adversarially-trained fair representations.

We find that using a balanced distribution over the sensitive attribute for adversarial training is much more effective than a random sample.

We empirically demonstrate the connection between the adversarial training data and the fairness metrics.

We observe that remarkably small datasets for adversarial training are effective in encouraging more fair representations.

Acknowledgements. We would like to thank Charlie Curry, Pierre Kreitmann, and Chris Berg for their feedback leading up to this paper.

References