Fairness without Demographics through Adversarially Reweighted Learning
Preethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee, Flavien Prost, Nithum Thain, Xuezhi Wang, Ed H. Chi
Introduction
As machine learning (ML) systems are increasingly used for decision making in high-stakes scenarios, it is vital that they do not exhibit discrimination. However, recent research has raised several fairness concerns, with researchers finding significant accuracy disparities across demographic groups in face detection , health-care systems , and recommendation systems . In response, there has been a flurry of research on fairness in ML, largely focused on proposing formal notions of fairness , and offering “de-biasing” methods to achieve these goals. However, most of these works assume that the model has access to protected features (e.g., race and gender), at least at training , if not at inference .
In practice, however, many situations arise where it is not feasible to collect or use protected features for decision making due to privacy, legal, or regulatory restrictions. For instance, GDPR imposes heightened prerequisites to collect and use protected features. Yet, in spite of these restrictions on access to protected features, and their usage in ML models, it is often imperative for our systems to promote fairness. For instance, regulators like CFBP require that creditors comply by fairness, yet prohibit them from using demographic information for decision-making. Creditors may not request or collect information about an applicant’s race, color, religion, national origin, or sex. Exceptions to this rule generally involve situations in which the information is necessary to test for compliance with fair lending rules. [CFBP Consumer Law and Regulations, 12 CFR §1002.5] Recent surveys of ML practitioners from both public-sector and industry highlight this conundrum, and identify “addressing fairness without demographics” as a crucial open-problem with high significance to ML practitioners. Therefore, in this paper, we ask the research question:
How can we train a ML model to improve fairness when we do not have access to protected features neither at training nor inference time, i.e., we do not know protected group memberships?
Goal: We follow the Rawlsian principle of Max-Min welfare for distributive justice . In Section 3.1, we formalize our Max-Min fairness goals: to train a model that maximizes the minimum expected utility across protected groups with the additional challenge that, we do not know protected group memberships. It is worth noting that, unlike parity based notions of fairness, which aim to minimize gap across groups, Max-Min fairness notion permits inequalities. For many high-stakes ML applications, such as healthcare and face recognition, improving the utility of worst-off groups is an important goal, and in some cases, parity notions that equally accept decreasing the accuracy of better performing groups are often not reasonable.
Exploiting Correlates: While the system does not have direct access to protected groups, we hypothesize that unobserved protected groups are correlated with the observed features (e.g., race is correlated with zip-code) and class labels (e.g., due to imbalanced class labels). As we will see in Table 4 (§5), this is frequently true. While correlates of protected features are a common cause for concern in the fairness literature, we show this property can be valuable for improving fairness metrics. Next, we illustrate how this correlated information can be valuable with a toy example.
Illustrative Example: Consider a classification task wherein our dataset consists of individuals with membership to one of the two protected groups: “orange” data points and “green” data points. The trainer only observes their position on the and axis. Although the model does not have access to the group (color), is correlated with the group membership.
Although each group alone is well-separable (Figure 1(a-b)), we see in Figure 1(c) that the empirical risk minimizing (ERM) classifer over the full data results in more errors for the green group. Even without color (groups), we can quickly identify a region of errors with low value (bottom of the plot) and a positive label (). In Section 3.2, we will define the notion of computationally-identifiable errors that correspond to this region. These errors are in contrast to outliers (e.g., from label noise) with larger errors randomly distributed across the - axes.
The closest prior work to ours is DRO . Similar to us, DRO has the goal of fairness without demographics, aiming to achieve Rawlsian Max-Min Fairness for unknown protected groups. However, to achieve this, DRO uses distributionally robust optimization to optimize for the worst-case groups by focusing on improving any worst-case distributions, but as the authors point out, this runs the risk of focusing optimization on noisy outliers. In contrast, we hypothesize that focusing on addressing computationally-identifiable errors will better improve fairness for the unobserved groups.
Adversarially Reweighted Learning: With this hypothesis, we propose Adversarially Reweighted Learning (ARL), an optimization approach that leverages the notion of computationally-identifiable errors through an adversary to improve worst-case performance over unobserved protected groups . Our experimental results show that ARL achieves high AUC for worst-case protected groups, high overall AUC, and robustness against training data biases.
Taken together, we make the following contributions:
Fairness without Demographics: In Section 3, we propose adversarially reweighted learning (ARL), a modeling approach that aims to improve the utility for worst-off protected groups, without access to protected features at training or inference time. Our key insight is that when improving model performance for worst-case groups, it is valuable to focus the objective on computationally-identifiable regions of errors.
Empirical Benefits: In Section 4, we evaluate ARL on three real-world datasets. Our results show that ARL yields significant AUC improvements for worst-case protected groups, outperforming state-of-the-art alternatives on all the datasets, and even improves the overall AUC on two of three datasets.
Understanding ARL: In Section 5 we do a thorough experimental analysis and present insights into the inner-workings of ARL by analyzing the learnt example weights. In addition, we perform a synthetic study to investigate robustness of ARL to worst-case training distributions. We observe that ARL is quite robust to representation bias, and differences in group base-rate. However, similar to prior approaches, ARL degrades with noisy ground-truth labels.
Related Work
Fairness: There has been an increasing line of work to address fairness concerns in machine learning models. A number of fairness notions have been proposed. At a high level they can be grouped into three categories, including (i) individual fairness , (ii) group fairness that expects parity of statistical performance across groups, and (iii) fairness notions that aim to improve per-group performance, such as Pareto-fairness and Rawlsian Max-Min fairness . In this work we follow the third notion of improving per-group performance.
There is also a large body of work on incorporating these fairness notions into ML models, including learning better representations and adding fairness constraints in the learning objective , through post-processing the decisions , through adversarial learning . These works generally assume the protected attribute information is known and thus the fairness metrics can be directly optimized. However, in many real world applications the protected attribute information might be missing or is very sparse.
Fairness without demographics: Some works address this approximately by using proxy features or assuming that the attribute is slightly perturbed . However, using proxies can in itself be prone to estimation bias . Multiple works have explored addressing limited demographic data with transfer learning . For example, Coston et al. 2019 focus on domain adaptation of fairness in settings where the group labels are known for either source or target dataset. Mohri et al. 2019 consider an agnostic federated learning, wherein given training data over clients with unknown sampling distributions, the model aims to learn mixture coefficient weights that optimize for a worst-case target distribution over these clients. Creager et al. 2020 propose an invariant risk minimization approach for domain generalization where environment partitions are not provided.
An interesting line of work tackles this problem by relying on trusted third parties that collect and store protected-data necessary for incorporating fairness. They generally assume that the ML model has access to the protected-features, albeit in encrypted form via secure multi-party computation , or in a privacy preserving form by employing differentially private learning .
As mentioned earlier, the work closest to ours is DRO , which uses techniques from distributionally robust optimization to achieve Rawlsian Max-Min fairness without access to demographics. However, a key difference between DRO and ARL is the type of groups identified by them: DRO considers any worst-case distribution exceeding a given size as a potential protected group. Concretely, given a lower bound on size of the smallest protected group, say , DRO optimizes for improving the worst-case performance of any set of examples exceeding size . In contrast, our work relies on a notion of computational-identifiability.
Computational-Identifiability: Related to our algorithm, a number of works address intersectional fairness by optimizing for group fairness between all computationally identifiable groups in the input space. While the perspective of learning over computationally identifiable groups is similar, they differ from us in that they assume the protected group features are available in their input space, and that they aim to minimize the gap in utility across groups via regularization.
Modeling Technique Inspirations: In terms of technical machinery, our proposed ARL approach draws inspiration from a wide variety of prior modeling techniques. Re-weighting is a popular paradigm typically used to address problems such as class imbalance by upweighting examples from minority class. Adversarial learning is typically used to train a model to be robust with respect to adversarial examples. Focal loss encourages the learning algorithm to focus on more difficult examples by up-weighting examples proportionate to their losses. Domain adaptation work requires a model to be robust and generalizable across different domains, under either covariate shift or label shift .
Model
We now dive into the precise problem formulation and our proposed modeling approach.
In this paper we consider a binary classification setup (though the approach can be generalized to other settings). We are given a training dataset consisting of individuals where is an dimensional input vector of non-protected features, and represents a binary class label. We assume there exist protected groups where for each example there exists an unobserved where is a random variable over . The set of examples with membership in group given by . Again, we do not observe distinct set but include the notation for formulation of the problem. To be more precise, we assume that protected groups are unobserved attributes not available at training or inference times. However, we will frame our definition and evaluation of fairness in terms of groups .
Problem Definition Given dataset , but no observed protected group memberships , e.g., race or gender, learn a model that is fair to groups in .
A natural next question is: what is a “fair” model? As in DRO , we follow the Rawlsian Max Min fairness principle of distributive justice : we aim to maximize the minimum utility a model has across all groups as given by Definition 1. Here, we assume that when a model predicts an example correctly, it increases utility for that example. As such can be considered any one of standard accuracy metrics in machine learning that models are designed to optimize for.
Suppose is a set of hypotheses, and is the expected utility of the hypothesis for the individuals in group , then a hypothesis is said to satisfy Rawlsian Max-Min fairness principle if it maximizes the utility of the worst-off group, i.e., the group with the lowest utility.
In our evaluation in Section 4, we use AUC as a utility metric, and report the minimum utility over protected groups as AUC(min).
2 Adversarial Reweighted Learning
Given this fairness definition and goal, how do we achieve it? As with traditional machine learning, most utility/accuracy metrics are not differentiable, and instead convex loss functions are used. The traditional ML task is to learn a model that minimizes the loss over the training data :
Therefore, we take the same perspective in turning Rawlsian Max-Min Fairness as given in Eq. (1) into a learning objective. Replacing the expected utility with an appropriate loss function over the set of individuals in group , we can formulate our fairness objective as:
Minimax Problem: Similar to Agnostic Federal Learning (AFL) , we can formulate the Rawlsian Max-Min Fairness objective function in Eq. (3) as a zero-sum game between two players and . The optimization comprises of game rounds. In round , player learns the best parameters that minimizes the expected loss. In round , player learns an assignment of weights that maximizes the weighted loss.
To derive a concrete algorithm we need to specify how the players pick and . For the player, one can use any iterative learning algorithm for classification tasks. For player , if the group memberships were known, the optimization problem in Eq. 4 can be solved by projecting on a probability simplex over groups given by as in AFL . Unfortunately, for us, because we do not observe we cannot directly optimize this objective as in AFL .
If the adversary was perfect it would adversarially assign all the weight () on the computationally-identifiable regions where learner makes significant errors, and thus improve learner performance in such regions. It is worth highlighting that, the design and complexity of the adversary model plays an important role in controlling the granularity of computationally-identifiable regions of error. More expressive leads to finer-grained upweighting but runs the risk of overfitting to outliers. While any differentiable model can be used for , we observed that for the small academic datasets used in our experiments, a linear adversary performed the best (further implementation details follow).
Observe that without any constraints on the objective in Eq. 5 is ill-defined. There is no finite that maximizes the loss, as an even higher loss could be achieved by scaling up . Thus, it is crucial that we constrain the values . In addition, it is necessary that for all , since minimizing the negative loss can result in unstable behaviour. Further, we do not want to fall to for any examples, so that all examples can contribute to the training loss. Finally, to prevent exploding gradients, it is important that the weights are normalized across the dataset (or current batch). In principle, our optimization problem is general enough to accommodate a wide variety of constraints. In this work we perform a normalization step that rescales the adversary to produce the weights . We center the output of and add 1 to ensure that all training examples contribute to the loss.
Implementation: In the experiments presented in Section 4, we use a standard feed-forward network to implement both learner and adversary. Our model for the learner is a fully connected two layer feed-forward network with 64 and 32 hidden units in the hidden layers, with ReLU activation function. While our adversary is general enough to be a deep network, we observed that for the small academic datasets used in our experiments, a linear adversary performed the best. Fig. 2 summarizes the computational graph of our proposed ARL approach. The python and tensorflow implementation of proposed ARL approach, as well as all the baselines is available opensource at https://github.com/google-research/google-research/tree/master/group_agnostic_fairness
Experimental results
We now demonstrate the effectiveness of our proposed ARL approach through experiments over three real datasets All the datasets used in this paper are publicly available. Key characteristics of the datasets, including a list of all the protected groups are in Supplementary (Tbl. 9) well used in the fairness literature: (i) Adult : income prediction (ii) LSAC : law school admission and (iii) COMPAS : recidivism prediction.
Evaluation Metrics: We choose AUC (area under the ROC curve) as our utility metric as it is robust to class imbalance, i.e., unlike Accuracy it is not easy to receive high performance for trivial predictions. Further, it encompasses both FPR and FNR, and is threshold agnostic.
To evaluate fairness we stratify the test data by groups, compute AUC per protected group , and report (i) AUC(min): minimum AUC over all protected groups, (ii) AUC(macro-avg): macro-average over all protected group AUCs and (iii) AUC(minority): AUC reported for the smallest protected group in the dataset. For all metrics higher values are better. Values reported are averages over 10 runs. Note that the protected features are removed from the dataset, and are not used for training, validation or testing. The protected features are only used to compute subgroup AUC in order to evaluate fairness.
Setup and Parameter Tuning: We use the same experimental setup, architecture, and hyper-parameter tuning for all the approaches. As our proposed ARL model has additional model capacity in the form of example weights , we increase the model capacity of the baselines by adding more hidden units in the intermediate layers of their DNN in order to ensure a fair comparison. Best hyper-parameter values for all approaches are chosen via grid-search by performing 5-fold cross validation optimizing for best overall AUC. We do not use subgroup information for training or tuning. DRO has a separate fairness parameter . For the sake of fair comparison, we report results for two variants of DRO: (i) DRO, with tuned as detailed in their paper and (ii) DRO (auc) with tuned to achieve best overall AUC performance. Refer to Supplementary for further details. All results reported are averages across 10 independent runs (with different model parameter initialization).
Main Results: Fairness without Demographics: Our main comparison is with DRO , a group-agnostic distributionally robust optimization approach that optimizes for the worst-case subgroup. Additionally, we report results for the vanilla group-agnostic Baseline, which performs standard ERM with uniform weights. Tbl. 1 reports results based on average performance across runs, with the best average performance highlighted in bold. Detailed results for all protected groups and with standard deviations across runs are reported in the Supplementary (Tbl. 6, 7, and 8). We make the following key observations:
ARL improves worst-case performance: ARL outperforms DRO, and achieves best results for AUC (minority) for all datasets. We observe a percentage point (pp) improvement over the baseline for Adult, pp for LSAC, and pp for COMPAS. Similarly, ARL shows pp and pp improvement in AUC (min) over baseline for Adult and LSAC datasets respectively. For COMPAS dataset there is no notable difference in performance over baseline, yet substantially better than DRO, which suffers a lot.
These results are inline with our observations on computational-identifiability of protected groups (Tbl. 4) and robustness to label bias (Fig 3(b)) in Section (§5). As we will later see, unlike Adult and LSAC datasets, protected-groups in COMPAS dataset are not computationally-identifiable. Further, ground-truth recidivism class labels in COMPAS dataset are known to be noisy . We suspect that noisy data, and biased ground-truth training labels play a role in the subpar performance of ARL and DRO for COMPAS dataset as both these approaches are susceptible to performance degradation in the presence of noisy labels as they cannot differentiate between mistakes on correct vs noisy labels as we later illustrate in Section 5. Due to label noise, the training errors are not easily computationally-identifiable. Hence, ARL shows no notable performance gain or loss. In contrast, we believe DRO is picking on noisy outlier in the dataset as high loss example, hence the substantial drop in DRO’s performance.
ARL improves overall AUC: Further, in contrast to the general expectation in fairness approaches, wherein utility-fairness trade-off is implicitly assumed, we observe that for Adult and LSAC datasets ARL in fact shows pp improvement in AUC (avg) and AUC (macro-avg). This is because ARL’s optimization objective of minimizing maximal loss is better aligned with improving overall AUC.
ARL vs Inverse Probability Weighting: Next, to better understand and illustrate the advantages of ARL over standard re-weighting approaches, we compare ARL with inverse probability weighting (IPW), which is the most common re-weighting choice used to address representational disparity problems. Specifically, IPW performs a weighted ERM with example weights set as where is the probability of observing an individual from group in the empirical training distribution. In addition to vanilla IPW, we also report results for a IPW variant with inverse probabilities computed jointly over protected-features and class-label reported as IPW(S+Y). Tbl. 2 summarizes the results. We make following observations and key takeaways:
Firstly, observe that in spite of not having access to demographic features, ARL has comparable if not better results than both variants of the IPW on all datasets. This results shows that even in the absence of group labels, ARL is able to appropriately assign adversarial weights to improve errors for protected-groups.
Further, not only does ARL improve subgroup fairness, in most settings it even outperforms IPW, which has perfect knowledge of group membership. This result further highlights the strength of ARL. We observed that this is because unlike IPW, ARL does not equally upweight all examples from protected groups, but does so only if the model needs much more capacity to be classified correctly. We present evidence of this observation in Section 5.
ARL vs Group-Fairness Approaches: While our fairness formulation is not the same as traditional group-fairness approaches, in order to better understand relationship between improving subgroup performance vs minimizing gap, we compare ARL with a group-fairness approach that aims for equal opportunity (EqOpp) . Amongst many EqOpp approaches , we choose Min-Diff as a comparison as it is the closest to ARL in terms of implementation and optimization. To ensure fair comparison we instantiate Min-Diff with similar neural architecture and model capacity as ARL. Further, as we are interested in performance for multiple protected groups, we add one Min-Diff loss term for each protected feature (sex and race). Details of the implementation are described in Supplementary §8. Tbl. 3 summarizes these results.
Min-Diff improves gap but not worst-off group: True to its goal, Min-Diff decreases the FPR gap between groups: FPR gap on sex is between and , and FPR gap on race is between and for all datasets. However, lower-gap between groups doesn’t always lead to improved AUC for worst-off groups (observe AUC min and AUC minority). ARL substantially outperforms Min-Diff for Adult and COMPAS datasets, and achieves comparable performance on LSAC dataset.
This result highlights the intrinsic mismatch between fairness goals of group-fairness approaches vs the desire to improve performance for protected groups. We believe making models more inclusive by improving the performance for groups, not just decreasing the gap, is an important complimentary direction for fairness research.
Utility-Fairness Trade-off: Further, observe that Min-Diff incurs a pp drop in overall AUC for Adult dataset, and pp drop for COMPAS dataset. In contrast, as noted earlier ARL in-fact shows an improvement in overall AUC for Adult and LSAC datasets. This result shows that unlike Min-Diff (or group fairness approaches in general) where there is an explicit utility-fairness trade-off, ARL achieves a better pareto allocation of overall and subgroup AUC performance. This is because the goal of ARL, which explicitly strives to improve the performance for protected groups is aligned with achieving better overall utility.
Analysis
Next, we conduct analysis to gain insights into ARL.
Are groups computationally-identifiable? We first test our hypothesis that unobserved protected groups are correlated with observed features and class label . Thus, even when they are unobserved, they can be computationally-identifiable. We test this hypothesis by training a predictive model to infer given and . Tbl. 4 reports the predictive accuracy of a linear model.
We observe that Adult and LSAC datasets have significant correlations with unobserved protected groups, which can be adversarially exploited to computationally-identify protected-groups. In contrast, for COMPAS dataset the protected-groups are not as computationally-identifiable While we did not perform a formal study on this, we observed that the adversarial predictive accuracy for race in COMPAS is 0.61 is consistent with prior work . We believe that there are demographic signals in the data, but they are not strong enough to predict groups well.. As we saw earlier in Tbl. 1 (§4) these results align with ARL showing no gain or loss for COMPAS dataset, but improvements for Adult and LSAC.
Robustness to training distributions: In this experiment, we investigate robustness of ARL and DRO approaches to training data biases , such as bias in group sizes (representation-bias) and bias due to noisy or incorrect ground-truth labels (label-bias). We use the Adult dataset and generate several semi-synthetic training sets with worst-case distributions (e.g., few training examples of “female” group) by sampling points from original training set. We then train our approaches on these worst-case training sets, and evaluate their performance on a fixed untainted original test set.
Concretely, to replicate representation-bias, we vary the fraction of female examples in training set by under/over-sampling female examples from training set. Similarly, to replicate label-bias, we vary fraction of incorrect labels by flipping ground-truth class labels uniformly at random for a fraction of training examples. In all experiments, training set size remains fixed. To mitigate the randomness in data sampling and optimization processes, we repeat the process 10 times and report results on a fixed untainted original test set (e.g., without adding label noise). Fig. 5 reports the results. In the interest of space, we limit ourselves to the protected-group “Female”. For each training setting shown on X-axis, we report the corresponding AUC for Female subgroup on Y-axis. The vertical bars in the plot are confidence intervals over 10 runs. We make the following observations:
Representation Bias: Both DRO and ARL are robust to the representation bias. ARL clearly outperforms DRO and baseline at all points. Surprisingly, we see a drop in AUC for baseline as the group-size increases. This is an artifact of having fixed training data size. As the fraction of female examples increases, we are forced to oversample female examples and downsample male examples; this leads to a decreases in the information present in training data and in turn leads to a worse performing model. In contrast, ARL and DRO cope better with this loss of information.
Label Bias: This experiment sheds interesting insights on the benefits of ARL over DRO. Recall that both approaches aim to focus on worst-case groups, however they differ in how these “groups” are formed. DRO is guaranteed to focus on worst-case risk for any group in the data exceeding size . In contrast, ARL would only improve the performance for groups that are computationally-identifiable over .
We performed this experiment by setting DRO hyperparameter to 0.2. We observed that while the fraction of incorrect ground truth class labels (i.e., outliers) is less than 0.2, the performance of ARL and DRO is nearly the same. As the fraction of outliers in the training set exceeds 0.2 we observe that DRO’s performance drops substantially. These results highlight that, as expected both ARL and DRO are sensitive to label bias (as they aim to up-weight examples with prediction error but cannot distinguish between true and noisy labels). However, as noted by Hashimoto et al. 2018, DRO exhibits a stronger trade-off between robustness to label bias and fairness.
Are learnt example weights meaningful? Next, we investigate if the example weights learnt by ARL are meaningful through the lense of training examples in the Adult dataset. Fig. 5 visualizes the example weights assigned by ARL stratified into four quadrants of a confusion matrix. Each subplot visualizes the learnt weights on x-axis and their corresponding density on y-axis. We make following observations:
Misclassified examples are upweighted: As expected, misclassified examples are upweighted (in Fig. 4(b) and 4(c)), whereas correctly classified examples are not upweighted ( in Fig. 4(a)). Further, we observe that even though this was not our original goal, as an interesting side-effect ARL has also learnt to address class imbalance problem in the dataset. Recall that our Adult dataset has class imbalance, and only of examples belong to class . Observe that, in spite of making no errors ARL assigns high weights to all class examples as shown in Fig. 4(d) (unlike in Fig. 4(a) where all class example have weight ).
ARL adjusts weights to base-rate. We smoothly vary the base-rate of female group in training data (i.e., we synthetically control fraction of female examples with class label in training data). Fig. 5 visualizes training data base-rate on x-axis and mean example weight learnt for the subgroup on y-axis. Observe that at female base-rate 0.1, i.e., when only 10% of female training examples belong to class , the mean weight assigned for examples in class is significantly higher than class . As base-rate increases, i.e., as the number of class examples increases, ARL correctly learns to decrease the weights for class examples, and increases the weights for class examples. These insights further explain the reason why ARL manages to improve overall AUC.
Conclusion
Improving model fairness without directly observing protected features is a difficult and under-studied challenge for putting machine learning fairness goals into practice. The limited prior work has focused on improving model performance for any worst-case distribution, but as we show this is particularly vulnerable to noisy outliers. Our key insight is that when improving model performance for worst-case groups, it is valuable to focus the objective on computationally-identifiable regions of errors i.e., regions of the input and label space with significant errors. In practice, we find ARL is better at improving AUC for worst-case protected groups across multiple dataset and over multiple types of training data biases. As a result, we believe this insight and the ARL method provides a foundation for how to pursue fairness without access to demographics.
Acknowledgements: The authors would like to thank Jay Yagnik for encouragement along this research direction and insightful connections that contributed to this work. We thank Alexander D’Amour, Krishna Gummadi, Yoni Halpern, and Gerhard Weikum for helpful feedback that improved the manuscript. This work was partly supported by the ERC Synergy Grant 610150 (imPACT).
Broader Impact
Any machine learning system that learns from data runs the risk of introducing unfairness in decision making. Recent research has identified fairness concerns in several ML systems, especially toward protected groups that are under-represented in the data. Thus, alongside the technical advancements in improving ML systems it crucial that we also focus on ensuring that they work for everyone.
One of the key practical challenges in addressing unfairness in ML systems is that most methods require access to protected demographic features, placing fairness and privacy in tension. In this work we work toward addressing these important challenges by proposing a new training method to improve worst-case performance of protected groups, in the absence of protected group information in the datasets. One limitation of methods in this space, including ours, is the difficulty of evaluating their effectiveness when we don’t have demographics in a real application. Therefore, while we think developing better debiasing methods is crucial, there remains further challenges in evaluating them.
Further, this work relies on the assumption that protected groups are computationally-identifiable. However, if there were no signal about protected groups in the remaining features and class labels , we cannot make any statements about improving the model for protected groups. Similarly, we observe that when the ground truth labels in the training dataset are noisy, the performance of ARL drops. Looking forward, we believe further research is needed to validate how the proposed approach can remain effective in a wide variety of real world applications beyond the datasets we have studied in this work.
References
Supplementary Material
Our proposed adversarially reweighted learning ARL approach is flexible and generalizes to many related works by varying the inputs to the adversary. For instance, if the domain of our adversary was , i.e., it took only protected features as input, the sets computationally-identifiable by the adversary boil down to an exhaustive cross over all the protected features , i.e., . Thus, our ARL objective in Eq. 5 would reduce to being very similar to the objective of fair agnostic federated learning by Mohri et al. 2019: to minimize the loss for the worst-off group amongst all known intersectional subgroups.
In this experiment, we further gain insights into our proposed adversarially re-weighting approach by comparing a number of variants ARL:
ARL (adv: X+Y) : vanilla ARL where the adversary takes non-protected features and class label as input.
ARL (adv: S): variant of ARL where the adversary takes only protected features as input.
ARL (adv: S+Y): variant of ARL with access to protected features and class label as input.
ARL(adv: X+Y+S): variant of ARL where the adversary takes all features and class label as input.
A summary of results is reported in Tbl. 5. We make the following observations:
Group-agnostic ARL is competitive: Firstly, observe that contrary to general expectation our vanilla ARL without access to protected groups , i.e., ARL (adv: X+Y) is competitive, and its results are comparable with ARL variants with access to protected-groups (except in the case of COMPAS dataset as observed earlier). These results highlight the strength of ARL as an approach achieve fairness without access to demographics.
Access to class label is crucial: Further, we observe that variants with class label () generally outperform variants without class label. For instance, for ARL(S+Y) has higher AUC than ARL(S) for all groups across all datasets. Especially for Adult and LSAC datasets, which are known to have class imbalance problem (observe base-rate in Tbl.9). A similar trend was observed for IPW(S) vs IPW(S+Y) in Tbl.2 (§4). This is expected and can be explained as follows: variants without access to class label such as ARL(S) are forced to give the same weight to both positive and negative examples of a group, As a consequence, they do not cope well with differences in base-rates, especially across groups, as they cannot treat majority and minority class differently.
Blind Fairness: Finally, in this work, we operated under the assumption that protected features are not available in the dataset. However, in practice there are scenarios where protected features are available in the dataset, however, we are blind to them. More concretely, we do not know a priori which subset of features amongst all features might be candidates for protected groups . Examples of this setting include scenarios wherein a number of demographics features (e.g., age, race, sex) are present in the dataset. However, we do not known which subgroup(s) amongst all intersectional groups (given by the cross-product over demographic features) might need potential fairness treatment.
Our proposed ARL approach naturally generalizes to this setting as well. We observe that the performance of our ARL variant ARL(adv: X+Y+S) is comparable to the performance of ARL(adv: Y+S). In certain cases (e.g., Adult dataset), access to remaining features even improves fairness. We believe this is because access to helps the adversary to make fine-grained distinctions amongst a subset of disadvantaged candidates in a given group that need fairness treatment.
2 Omitted Tables
In Section 4 Our main comparison is with DRO , a group-agnostic distributionally robust optimization approach that optimizes for the worst-case subgroup. Additionally, we report results for the vanilla group-agnostic Baseline, which performs standard ERM with uniform weights. Tbl. 6 7 and 8 summarize the main results. We report AUC (mean std) for all protected groups in each dataset. Best values in each table are highlighted in bold.
3 Datasets and Pre-processing
Datasets for Main Experiments: We perform our experiments on three real-world, publicly available datasets, previously used in the literature on algorithmic fairness:
Adult: The UCI Adult dataset contains US census income survey records. We use the binarized “income” feature as the target variable for our classification task to predict if an individual’s income is above .
LSAC: The Law School dataset from the law school admissions council’s national longitudinal bar passage study to predict whether a candidate would pass the bar exam. It consists of law school admission records. We use the binary feature “isPassBar” as the raget variable for classification.
COMPAS: The COMPAS dataset for recidivism prediction consists of criminal records comprising offender’s criminal history, demographic features (sex, race). We use the ground truth on whether the offender was re-arrested (binary) as the target variable for classification.
We transform all categorical attributes using one-hot encoding, and standardize all features vectors to have zero mean and unit variance. Python scripts for preprocessing the public datasets are open accessible along with the rest of the code of this paper.
Datasets for Synthetic Experiments: In Section 5 we perform additional experiments on a number of semi-synthetic datasets to investigate robustness of ARL to training distributions. All the code to generate synthetic datasets is shared along with the rest of the code of this paper.
4 Baselines and Implementation
We compare our proposed approach ARL with the two naive baselines and one state-of-the-art approach. All the implementations are open accessible along with the rest of the code of this paper. All approaches have the same DNN architecture, optimizer and activation functions. As our proposed ARL model has additional model capacity in the form of example weights , in order to ensure fair comparison we increase the model capacity of the baselines by adding more hidden units in the intermediate layers of their DNN. Following are the implementation details:
Baseline: This is a simple empirical risk minimization baseline with standard binary cross-entropy loss.
IPW: This a naive re-weighted risk minimization approach with weighted binary cross-entropy loss. The weights are assigned to be inverse probability weights , where is the . For a fair comparison, we train IPW with the same model as ARL, with fixed adversarial re-weighting. More concretely, rather than adversarially learning weights in a group-agnostic manner, the example weights () are precomputed inverse probability weights . Additionally, we perform experiments on a variant of IPW called IPW (S+Y) with weights , where is the joint probability of observing a data-point having membership to group and class label over empirical training distributions.
DRO: This is a group-agnostic distributionally robust learning approach for fair classification. We use the code shared by , which is available at https://worksheets.codalab.org/worksheets/0x17a501d37bbe49279b0c70ae10813f4c/. We tune the hyper-parameters for DRO by performing grid search over the parameter space as reported in their paper.
5 Experimental Setup and Parameter Tuning
Each dataset is randomly split into 70% training and 30% test sets. On the training set, we perform a 5 fold cross validation to find the best hyper-parameters for each model (details follow). Once the hyperparameters are tuned, we use the second part as an independent test set to get an unbiased estimate of their performance. We use the same experimental setup, data split, and parameter tuning techniques for all the methods.
Hyperparameter Tuning: For each approach, we choose the best learning-rate, and batch size by performing a grid search over an exhaustive hyper parameter space given by batch size (32, 64, 128, 256, 512) and learning rate (0.001, 0.01, 0.1, 1, 2, 5). All the parameters are chosen via 5-fold cross validation by optimizing for best overall AUC.
In addition to batch size, and learning rate, DRO approach has an additional fairness hyper-parameter , which controls the performance for the worst-case subgroup. In their paper, the authors present a specific hyperparameter tuning approach to choose the best value for . Hence for the sake of fair comparison, we report results for two variants of DRO: (i) DRO, original approach with tuned as detailed in their paper and (ii) DRO(auc) with tuned to achieve best overall AUC performance.
6 Reproducibility
All the datasets used in this paper are publicly available. The python and tensorflow implementation of proposed ARL approach, as well as scripts to generate synthetic datasets is available opensource at https://github.com/google-research/google-research/tree/master/group_agnostic_fairness.