Fair Inference On Outcomes

Razieh Nabi, Ilya Shpitser

Introduction

As statistical and machine learning models become an increasingly ubiquitous part of our lives, policymakers, regulators, and advocates have expressed concerns about the impact of deployment of such models that encode potential harmful and discriminatory biases. Unfortunately, data analysis is based on statistical models that do not, by default, encode human intuitions about fairness and bias. For instance, it is well-known that recidivism is predicted at higher rates among certain minorities in the US (?). To what extent are these predictions discriminatory? What is a sensible framework for thinking about these issues? A growing community is now addressing issues of fairness and transparency in data analysis in part by defining, analyzing, and mitigating harmful effects of algorithmic bias from a variety of perspectives and frameworks (?; ?; ?; ?; ?; ?).

In this paper, we propose to model discrimination based on a “sensitive feature,” such as race or gender, with respect to an outcome as the presence of an effect of the feature on the outcome along certain “disallowed” causal pathways. As a simple example, discussed in (?), job applicants’ gender should not directly influence the hiring decision, but may influence the hiring decision indirectly, via secondary applicant characteristics important for the job, and correlated with gender. We argue that this view captures a number of intuitive properties of discrimination, and generalizes existing formal (?; ?) and informal proposals (?).

The paper is organized as follows. We first fix our notation and give a brief introduction to causal inference and mediation analysis, which will be necessary to formally define our approach to fair inference. We then discuss representative prior work on fair inference, and enumerate issues these methods may run into. Moving forward, we show that fair inference from finite samples under our definition can be viewed as a certain type of constrained optimization problem. We then discuss a number of complications to the basic framework of fair inference. We illustrate our framework via experiments on real datasets in the experimental section followed by additional discussion and final conclusions.

Notation And Preliminaries

Variables will be denoted by uppercase letters, VV, values by lowercase letters, vv, and sets by bold letters. A state space of a variable will be denoted by XV{\mathfrak{X}}_{V}. We will represent datasets by D=(Y,X){\cal D}=(Y,{\bf X}), where YY is the outcome and X{\bf X} is the feature vector. We denote by xjix^{i}_{j} and yiy^{i} the iith realization of the jjth feature Xj∈XX_{j}\in{\bf X} and the outcome YY. Similarly, xi{\bf x}^{i} is the iith realization of the entire feature vector.

In this paper, we consider probabilistic classification and regression problems with a set of features X{\bf X} and an outcome YY, where a feature S∈XS\in{\bf X} is sensitive, in the sense that making inferences on the outcome YY based on SS carelessly may result in discrimination. There are many examples of S,YS,Y pairs that have this property. These include hiring discrimination (YY is a hiring decision, and SS is gender), or recidivism prediction in parole hearings (where YY is a parole decision, and SS is race). Our approach readily generalizes to any outcome based inference task, such as establishing causal effects, although we do not consider these generalizations here in the interests of space.

Causal Inference

Causal inference uses assumptions in causal models to link observed data with counterfactual contrasts of interest. When such a functional exists, we say the parameter is identified from the observed data under the causal model. One such assumption, known as consistency, states that the mechanism that determines the value of the outcome does not distinguish the method by which the treatment was assigned, as long as the treatment value assigned was invariant. This is expressed as Y(A)=YY(A)=Y. Here Y(A)Y(A) reads “the random variable YY, had AA been intervened on to whatever value AA would have naturally attained.”

Another standard assumption is known as conditional ignorability. This assumption states that conditional on a set of factors C⊆X{\bf C}\subseteq{\bf X}, AA is independent of any counterfactual outcome, i.e. Y(a)⊥ ⁣ ⁣ ⁣⊥A∣C,∀a∈XAY(a)\perp\!\!\!\perp A|{\bf C},\forall a\in{\mathfrak{X}}_{A}, where (.⊥ ⁣ ⁣ ⁣⊥.∣.)(.\perp\!\!\!\perp.|.) represents conditional independence. Given these assumptions, we can show that p(Y(a))=∑Cp(Y∣a,C)p(C)p(Y(a))=\sum_{\bf C}p(Y\mid a,{\bf C})p({\bf C}), known as the adjustment formula, the backdoor formula, or stratification. Intuitively, the set C{\bf C} acts as a set of observed confounders, such that adjusting for their influence suffices to remove all non-causal dependence of AA and YY, leaving only the part of the dependence that corresponds to the causal effect. A general characterization of identifiable functionals of causal effects exists (?; ?).

Causal relationships are often represented by graphical causal models (?; ?). Such models generalize independence models on directed acyclic graphs, also known as Bayesian networks (?), to also encode conditional independence statements on counterfactual random variables (?). In such graphs, vertices represent observed random variables, and absence of directed edges represents absence of direct causal relationships. As an example, in Fig. 1 (a), CC is potentially a direct cause of AA, while MM mediates a part of the causal influence of AA on YY, represented by all directed paths from AA to YY.

Mediation Analysis

A natural step in causal inference is understanding the mechanism by which AA influences YY. A simple form of understanding mechanisms is via mediation analysis, where the causal influence of AA on YY, as quantified by the ACE, is decomposed into the direct effect, and the indirect effect mediated by a mediator variable MM. In typical mediation settings, X{\bf X} is partitioned into a treatment AA, a single mediator MM, an outcome YY, and a set of baseline factors C=X∖{A,M,Y}{\bf C}={\bf X}\setminus\{A,M,Y\}.

Mediation is encoded via a counterfactual contrast using a nested potential outcome of the form Y(a,M(a′))Y(a,M(a^{\prime})), for a,a′∈XAa,a^{\prime}\in{\mathfrak{X}}_{A}. Y(a,M(a′))Y(a,M(a^{\prime})) reads as “the outcome YY if AA were set to aa, while MM were set to whatever value it would have attained had AA been set to a′a^{\prime}. An intuitive interpretation for this counterfactual occurs in cases where a treatment can be decomposed into two disjoint parts, one of which acts on YY but not MM, and another acts on MM but not YY. For instance, smoking can be decomposed into smoke and nicotine. Then if MM is a mediator affected by smoke, but not nicotine (for instance lung cancer), and YY is a composite health outcome, then Y(a,M(a′))Y(a,M(a^{\prime})) corresponds to the response of YY to an intervention that sets the nicotine exposure (the part of the treatment associated with YY) to what it would be in smokers, and the smoke exposure (the part of the treatment associated with MM) to what it would be in non-smokers. An example of such an intervention would be a nicotine patch.

Aside from consistency, additional assumptions are needed to identify p(Y(a,M(a′)))p(Y(a,M(a^{\prime}))). One such assumption is known as sequential ignorability, and states that conditional on C{\bf C}, counterfactuals Y(a,m)Y(a,m) and M(a′)M(a^{\prime}) are independent for any a,a′∈XAa,a^{\prime}\in{\mathfrak{X}}_{A}, m∈XMm\in{\mathfrak{X}}_{M}. In addition, conditional ignorability for YY acting as the outcome, and A,MA,M acting as a single composite treatment, that is (Y(a,m)⊥ ⁣ ⁣ ⁣⊥A,M∣C)(Y(a,m)\perp\!\!\!\perp A,M\mid{\bf C}), and conditional ignorability for MM acting as the outcome, that is (M(a′)⊥ ⁣ ⁣ ⁣⊥A∣C)(M(a^{\prime})\perp\!\!\!\perp A\mid{\bf C}), should hold. Under these assumptions, the NDE is identified as the functional known as the mediation formula (?):

which may be estimated by plug in estimators, or other methods (?).

Path-Specific Effects

In general, we may be interested in decomposing the ACE into effects along particular causal pathways. For example in Fig. 1 (b), we may wish to decompose the effect of AA on YY into the contribution of the path A→W→YA\to W\to Y, and the path bundle A→YA\to Y and A→M→W→YA\to M\to W\to Y. Effects along paths, such as an effect along the path A→W→YA\to W\to Y, are known as path-specific effects (?). Just as the NDE and NIE, path-specific effects (PSEs) can be formulated as nested counterfactuals (?). The general idea is that along the pathways of interest, variables behave as if the treatment variable AA were set to the “active value” aa, and along other pathways, variables behave as if the treatment variable AA were set to the “baseline value” a′a^{\prime}, thus “turning the treatment off.” Using this scheme, the path-specific effect of AA on YY along the path A→W→YA\to W\to Y, on the mean difference scale, can be formulated as

and may be estimated by plug-in estimators. PSEs such as (2) has been used in the context of observational studies of HIV patients for assessing the role of adherence in determining viral failure outcomes (?).

Formalizing Discrimination And Prior Approaches To Fair Inference

We are now ready to discuss prior approaches to fair inference. In discussing the extent to which a particular approach is “fair,” we believe the gold standard is human intuition. That is, we consider an approach inappropriate if it leads to counter-intuitive conclusions in examples. For space reasons, we restrict attention to a representative subset of approaches.

A common class of approaches for fair inference is to quantify fairness via an associative (rather than causal) relationship between the sensitive feature SS and the outcome YY. For instance, (?) adopted the 80% rule, for comparing selection rates based on sensitive features. This is a guideline (not a legal test) advocated by the Equal Employment Opportunity Commission (?) as a way of suggesting possible discrimination. Rate of selection here is defined as the conditional probability of selection given the sensitive feature, or p(Y∣S)p(Y|S). (?) proposed methods for removing disparities based on this rule via a link to classification accuracy. A White House report on “equal opportunity by design" (?) prompted (?) to propose a fairness criterion, called equalized odds, that ensures that true and false positive rates are equal across all groups. This criterion is also associative.

The issue with these approaches is they do not give intuitive results in cases where the sensitive feature is not randomly assigned (as gender is at conception), but instead exhibits spurious correlations with the outcome via another, possibly unobserved, feature. We illustrate the difficulty with the following hypothetical example. Certain states in the US prohibit discrimination based on past conviction history. Prior convictions are influenced by other variables, such as gender (men have more prior convictions than women, on average). Consider a hypothetical dataset (consisting mostly of people with prior convictions) with two features – prior conviction (CC) and gender (GG) as well as the hiring outcome (HH). The values are coded as follows: male is 11, female is , prior conviction and hiring are 11, lack of prior conviction and no hiring is . Assume the dataset is drawn from the joint density specified as follows: p(G=1)=0.5p(G=1)=0.5, and

That is, gender is randomly assigned at birth, the people in the cohort are very likely to have prior convictions (with men having more), and p(H∣C,G)p(H|C,G) specifies a certain hiring rule for the cohort. For simplicity, we assume no other features of people in the cohort are relevant for either the prior conviction or the hiring decision. It’s easy to show that

However, intuitively we would consider a hiring rule in this example fair if, in a hypothetical randomized trial that assigned convictions randomly (conviction to the case group, no conviction to the control group), the rule would yield equal hiring probabilities to cases and controls. In our example, this implies comparing counterfactual probabilities p(H(C=1))p(H(C=1)) and p(H(C=0))p(H(C=0)). Since we posited no other relevant features for assigning CC and HH than AA, these probabilities are identified, via the adjustment formula described earlier, yielding p(H(C=1))=0.035p(H(C=1))=0.035, and p(H(C=0))=0.125p(H(C=0))=0.125. That is, any method relying on associative measures of discrimination will likely conclude no discrimination here, yet the intuitively compelling test of discrimination will reveal a strong preference to hiring people without prior convictions. The large difference between p(H(C=0))p(H(C=0)) and p(H∣C=0)p(H\mid C=0) has to do with extreme probabilities p(C∣G)p(C|G) in our example. Even in less extreme examples, any approach that relies on associative measures of association will be led astray due to failing to properly model sources of confounding for the relationship of the sensitive feature and the outcome. One might imagine that a simple repair in this example would be to also include GG as a feature. The reason this does not work in general is not all features are possible to measure, and in general counterfactual probabilities are complex functions of the observed data, not just conditional densities (?).

In the example above, it made intuitive sense to think of discrimination as a causal relationship between the sensitive feature and outcome. In other examples, discrimination intuitively entails only a part of the causal relationship. Consider a modification of the hiring example above where potential discrimination is with respect to gender (a variable randomized at conception, which means worries about confounding are no longer relevant). As before, consider binary variables GG and HH for gender and hiring, and an additional vector C{\bf C}, representing applicant characteristics relevant for the job, of the kind that would appear on the resume. The intuition here is it is legitimate to consider job characteristics in making hiring decisions even if those characteristics are correlated with gender. However, it is not legitimate to consider gender directly. This intuition underscores resume “name-swapping” experiments where identical resumes are sent for review with names switched from a Caucasian sounding name to an African-American sounding name (?). In such experiments, name serves as a proxy for race as a direct determinant of the hiring decision.

The definition of discrimination as related to causal pathways is further supported in the legal literature. The following definition of employment discrimination, which appeared in the legal literature (?), and was cited by (?), makes clear the counterfactual nature of our intuitive conception of discrimination:

The central question in any employment-discrimination case is whether the employer would have taken the same action had the employee been of a different race (age, sex, religion, national origin etc.) and everything else had been the same.

The counterfactual “had the employee been of a different gender” phrase entails considering, for women, the outcome YY had gender been male G=1G=1, while the “everything else had been the same” phrase entails considering job characteristics under the original gender G=0G=0. The resulting counterfactual Y(G=1,C(G=0))Y(G=1,{\bf C}(G=0)) is precisely the one used in mediation analysis to define natural direct effects.

It is possible to construct examples, discussed further, where some causal paths from a sensitive variable to the outcome are intuitively discriminatory, and others are not. Thus, our view is that discrimination ought to be formalized as the presence of certain path-specific effects. The specific paths which correspond to discrimination are a domain specific issue. For example, physical fitness tests may be appropriate to administer for certain physically demanding jobs, such as construction, but not for white collar jobs, such as accounting. As a result, a path from gender to the result of a test to a hiring decision may or may not be discriminatory, depending on the nature of the job.

Existing work considered similar proposals. Prior work closest to ours appears in (?), where discrimination was also linked to path-specific effects. While we agree with the link, we disagree on three essential points. First, the authors do not appear to do any statistical inference and operate directly on discrete densities. In high dimensional settings, where most practical outcome based inference takes place, statistical modeling becomes necessary, and an approach that avoids it will not scale. In this paper we show how removing certain path-specific effects corresponds to a constrained inference problem on statistical models. As we show later, the constrained optimization problems that arise are non-trivial. Second, the authors propose an ad hoc repair in cases where the path-specific effect is not identifiable. We believe this is a misunderstanding of the concept of non-identifiability. If discrimination is indeed linked to a path-specific effect and this effect is not identifiable (not a function of the observed data), then the problem of removing discrimination is not solvable without more assumptions. In domains such as recidivism prediction, failure today and a better method tomorrow is preferable to an improper correction that preserves discriminatory practice. We discuss more principled approaches to repairing lack of identifiability of discriminatory path-specific effects in later sections. Finally, the authors, while repairing the observed data distribution to be fair, do not modify new instances to be classified in any way. Since new instances are, by definition, drawn from the observed data distribution, which is “unfair,” no guarantees about discrimination when classifying new samples can be made. We discuss this issue further below.

Inference On Outcomes That Minimizes Discriminatory Path-Specific Effects

We now describe our proposal precisely. Assume we are interested in making inferences on outcomes given a joint distribution p(Y,X)p(Y,{\bf X}) either modeled fully via a generative model, or partly via a discriminative model p(Y,X∖W∣W)p(Y,{\bf X}\setminus{\bf W}|{\bf W}). In addition, we assume that p(Y,X)p(Y,{\bf X}) is induced by a causal model in the sense of (?; ?), that the presence of discrimination based on some sensitive feature AA with respect to YY is represented by a PSE, and that this PSE is identified given the causal model as a functional f(p(Y,X))f(p(Y,{\bf X})). Finally, we fix upper and lower bounds ϵl,ϵu\epsilon_{l},\epsilon_{u} on the PSE, representing the degree of discrimination we are willing to tolerate.

A feature of our proposal is that we are selectively ignoring some known information about a new instance xi{\bf x}^{i}, if this information was drawn from the distribution that differs from p∗p^{*}. We believe this is unavoidable in fair inference settings – the entire point is using the information “as effectively as possible" is discriminatory. We do want to use information as well as possible, but only insofar as we remain in the “fair world".

Given a set of finite samples D{\cal D} drawn from p(Y,X)p(Y,{\bf X}), a PSE representing discrimination identified as f(p(Y,X))f(p(Y,{\bf X})), a (possibly conditional) likelihood function LY,X(D;α){\cal L}_{Y,{\bf X}}({\cal D};{\bm{\alpha}}) parameterized by α\bm{\alpha}, an estimator g(D)g({\cal D}) of the PSE, and ϵl,ϵu\epsilon_{l},\epsilon_{u}, we approximate p∗p^{*} by solving a constrained maximum likelihood problem

Our choice for the set W{\bf W} will be guided by the form of the estimator g(.)g(.). Specifically W{\bf W} will contain variables with models not a part of g(.)g(.). Since the estimators for the PSE developed within the causal inference literature do not model the baseline factors C{\bf C}, C⊆W{\bf C}\subseteq{\bf W}. In addition, certain estimators also do not use other parts of the model. For example, (5) below does not use p(A∣C)p(A\mid{\bf C}). For such estimators we also include those variables in W{\bf W}.

We now illustrate the relationship between the choice of W{\bf W} and the choice of gg by considering three of the four consistent estimators of the NDE (assuming the model shown in Fig. 1 (a) is correct) presented in (?). The first estimator is the MLE plug in estimator for (1), given by

The second estimator uses all three models, as follows:

The final estimator is based on inverse probability weighting (IPW). The IPW estimator uses the AA and MM models to estimate the NDE. We can fit the models p(A∣C)p(A|\bf C) and p(M∣A,C)p(M|A,\bf C) by MLE, and use the following weighted empirical average as our estimate of the NDE:

Since solving the constrained MLE problem using this estimator entails only restricting parameters of AA and MM models, predicting a new instance ai,mi,cia^{i},m^{i},{\bf c}^{i} is done using

A sensitive feature may affect the outcome through multiple paths, and paths other than a single edge path corresponding to the direct effect may be inadmissible. Consider an example where AA is gender, and YY is a hiring decision in a construction labor agency. We now consider two mediators of AA, the number of children MM, and physical strength as measured by an entrance test WW. In this setting, it seems that it is inappropriate for the applicant’s gender AA to directly influence the hiring decision YY, nor for the number of the subject’s children to influence the hiring decision either since the consensus is that women should not be penalized in their career for the biological necessity of having to bear children in the family. However, gender also likely influences the subject’s performance on the entrance test, and requiring that certain requirements of strength and fitness is reasonable in a job like construction. The situation is represented by Fig. 1 (b), with a hidden common cause of MM and YY added since it does not influence the subsequent analysis.

In this case, the PSE that must be minimized for the purposes of making the hiring decision is given by (2), and is identified, given a causal model in Fig. 1 (b), by (3). If we use the analogue of (5), we would maximize L(D;α){\cal L}({\cal D};{\bm{\alpha}}) subject to

Fair Inference Via Box Constraints

Second, consider a setting with a binary outcome YY specified with a logistic model, logit⁡(P(Y=1∣A,M,C))=θ0+θaA+θmM+θcC\operatorname{logit}(P(Y=1\mid A,M,C))=\theta_{0}+\theta_{a}A+\theta_{m}M+\theta_{c}C, a continuous mediator specified with a linear model, and the NDE (within a given level of CC) defined on the odds ratio scale as in (?):

In this setting, it was shown in (?) that under certain additional assumptions, the NDE has no dependencies on CC, and is approximately equal to exp⁡(θa)\exp(\theta_{a}). Hence, (4) can be expressed using box constraints on θa\theta_{a}. Note that the null hypothesis of the absence of discrimination corresponds to the value of 11 of the NDE on the odds ratio scale, and to on the mean difference scale.

Fair Inference With A Regularized Outcome Model

In many applications, the goal of inference is not to approximate the true model itself, but to maximize out of sample prediction performance regardless of what the true model might be. Validation datasets or resampling approaches can be used to assess the performance of such predictive models, with various regularization methods used to make the tradeoff between bias and variance. The difficulty here is that searching for a model YY with good out of sample prediction performance implies the true YY model might be sufficiently complex that it may not lie within the model we consider. This means we cannot use any estimator of the PSE that relies on the YY model. This is because most estimators that rely on the YY model are not consistent if the YY model is misspecified. An inconsistent estimator of the PSE implies we cannot be sure solving the constrained optimization problem will indeed remove discrimination.

Fairness In Computational Bayesian Methods

Methods for fair inference described so far are fundamentally frequentist in character, in a sense that they assumed a particular true parameter value, and parameter fitting was constrained in a way that an estimate of this parameter was within specified bounds. Here, we do not extend our approach to a fully Bayesian setting, where we would update distributions over causal parameters based on data, and use the resulting posterior distributions for constraining inferences. Instead, we consider how Bayesian methods for estimating conditional densities can be adapted, as a computational tool, to our frequentist approach.

Many Bayesian methods do not compute a posterior distribution explicitly, but instead sample the posterior using Markov chain Monte Carlo approaches (?). These sampling methods can be used to compute any function of the posterior distribution, including conditional expectations, and can be modified to obey constraints in our problem in a straightforward way. As an example, we consider BART, a popular Bayesian random forest method described in (?). This method constructs a distribution over a forest of regression trees, with a prior that favors small trees, and samples the posterior using a variant of Gibbs sampling, where a new tree is chosen while all others are held fixed. A well known result (?) states that a Gibbs sampler will generate samples from a constrained posterior directly if it rejects all draws that violate the constraint.

We implemented this simple method by modifying the R package (with a C++ backend) BayesTree, and applied the result to the model in Fig 2 (a), where NDE was estimated via a mixed estimation strategy described in (?), where AA was assumed to be randomly assigned (i.e. no modeling, and hence no constraining, of AA was required), and the YY model was fit using constrained BART. The experiment using the resulting constrained outcome model is described in the experimental section.

Dealing With Non-Identification of the PSE

Suppose our problem entailed the causal model in Fig. 1 (b), or Fig. 1 (c) where in both cases only the NDE of AA on YY is discriminatory. Existing identification results for PSEs (?) imply that the NDE is not identified in either model. This means estimation of the NDE from observed data is not possible as the NDE is not a function of the observed data distribution in either model.

In such cases, three approaches are possible. In both cases, the unobserved confounders UU are responsible for the lack of identification. If it were possible to obtain data on these variables, or obtain reliable proxies for them, the NDE becomes identifiable in both cases. If measuring UU is not possible, a second alternative is to consider a PSE that is identified, and that includes the paths in the PSE of interest and other paths. For example, in Fig. 1 (b), while the NDE of AA on YY, which is the PSE including only the path A→YA\to Y, is not identified, the PSE which includes paths A→YA\to Y, A→M→YA\to M\to Y, and A→M→W→YA\to M\to W\to Y, namely (2), is. The first counterfactual in the PSE contrast is identified in Fig. 1 (b) by (3), and the second by the adjustment formula.

If we are using the PSE on the mean difference scale, the magnitude of the effect which includes more paths than the PSE we are interested in must be an upper bound on the magnitude of the PSE of interest in order for the bounds we impose to actually limit discrimination. This is only possible if, for instance, all causal influence of AA on YY along paths involved in the PSE are of the same sign. In Fig. 1 (b), this would mean assuming that if we expect the NDE of AA on YY to be negative (due to discrimination), then it is also negative along the paths A→M→W→YA\to M\to W\to Y, and A→M→YA\to M\to Y.

If measuring UU is impossible, and it is not possible to find an identifiable PSE that includes the paths of interest from AA to YY, and serves as a useful upper bound to the PSE of interest, the other alternative is to use bounds derived for non-identifiable PSEs. While finding such bounds is an open problem in general, they were derived in the context of the NDE with a discrete mediator in (?).

The issue with non-identification of the PSE was also noted in (?). They proposed to change the causal model, specifically by cutting off some paths from the sensitive variable to the outcome such that the identification criterion in (?) became satisfied, and the PSE became identified. We disagree with this approach, as we believe it amounts to “redefining success.” If the original causal model truly represents our beliefs about the structure of the problem, and in particular the pathways corresponding to discrimination, then making any sort of inferences in a model modified away from truth no longer tracks reality. We would certainly not expect any kind of repair within a modified model to result in fair inferences in the real world. The workarounds for non-identification we propose aim to stay within the true model, but try to obtain information on the true non-identified PSE, either by non-parametric bounds, or by including other pathways along with the “unfair” pathways.

Experiments

We first illustrate our approach to fair inference via two datasets: the COMPAS dataset (?) and the Adult dataset (?). We also illustrate how a part of the model involving the outcome YY may be regularized without compromising fair inferences if the NDE quantifying discrimination is estimated using methods that are robust to misspecification of the YY model.

Correctional Offender Management Profiling for Alternative Sanctions, or COMPAS, is a risk assessment tool, created by the company Northpointe, that is being used across the US to determine whether to release or detain a defendant before his or her trial. Each pretrial defendant receives several COMPAS scores based on factors including but not limited to demographics, criminal history, family history, and social status. Among these scores, we are primarily interested in “Risk of Recidivism". Propublica (?) has obtained two years worth of COMPAS scores from the Broward County Sheriff’s Office in Florida that contains scores for over 1100011000 people who were assessed at the pretrial stage and scored in 2013 and 2014. COMPAS score for each defendant ranges from 11 to 1010, with 1010 being the highest risk. Besides the COMPAS score, the data also includes records on defendant’s age, gender, race, prior convictions, and whether or not recidivism occurred in a span of two years. We limited our attention to the cohort consisting of African-Americans and Caucasians.

We are interested in predicting whether a defendant would reoffend using the COMPAS data. For illustration, we assume the use of prior convictions, possibly influenced by race, is fair for determining recidivism. Thus, we defined discrimination as effect along the direct path from race to the recidivism prediction outcome. The simplified causal graph model for this task is given in Figure 2 (a), where AA denotes race, prior convictions is the mediator MM, demographic information such as age and gender are collected in C\bf C, and YY is recidivism. The “disallowed" path in this problem is drawn in green in Figure 2(a). The effect along this path is the NDE. The objective is to learn a fair model for YY. i.e. a model where NDE is minimized.

The Adult Dataset

The “adult” dataset from the UCI repository has records on 1414 attributes such as demographic information, level of education, and job related variables such as occupation and work class on 4884248842 instances along with their income that is recorded as a binary variable denoting whether individuals have income above or below 50k50k – high vs low income. The objective is to learn a statistical model that predicts the class of income for a given individual. Suppose banks are interested in using this model to identify reliable candidates for loan application. Raw use of data might construct models that are biased towards females who are perceived to have lower income in general compared to males. The causal model for this dataset is drawn in Figure 2(b). Gender is the sensitive variable in this example denoted by AA in figure 2(b) and income class is denoted by YY. MM denotes the marital status, LL denotes the level of education, and R\bf R consists of three variables, occupation, hours per week, and work class. The baseline variables including age and nationality are collected in C\bf C. U1U_{1} and U2U_{2} capture the unobserved confounders between M,YM,Y and L,RL,\bf R, respectively.

Here, besides the direct effect (A→YA\rightarrow Y), we would like to remove the effect of gender on income through marital status (A→M→…→YA\rightarrow M\rightarrow\ldots\rightarrow Y). The “disallowed" paths are drawn in green in Figure 2(b). The PSE along the green paths is identifiable via the recanting district criterion in (?), and can be computed by calculating odds ratio or contrast comparison of the counterfactual variable Y(a,M(a),L(a′,M(a)),R(a′,M(a),L(a′,M(a))),C),Y(a,M(a),L(a^{\prime},M(a)),{\bf R}(a^{\prime},M(a),L(a^{\prime},M(a))),{\bf C}), where a′a^{\prime} is set to a baseline value, a=1a=1 in one counterfactual, and a=0a=0 in the other. The counterfactual distribution can be estimated from the following functional: ∑V∖A{p(Y∣a,m,l,r,c)∏i=13p(ri∣a′,m,l,c)p(l∣a′,m,c)\sum_{{\bf V}\setminus A}\{p(Y|a,m,l,{\bf r,c})\prod_{i=1}^{3}p(r_{i}|a^{\prime},m,l,{\bf c})p(l|a^{\prime},m,{\bf c}) p(m∣a,c) p(c)}p(m|a,{\bf c})\ p(\bf c)\}, where V{\bf V} are all observed variables.

If we use logistic regression to model YY and linear regression to model other variables given their past, and compute the PSE on the odds ratio scale, it is straightforward to show that the PSE simplifies to \exp\big{(}\theta_{a}^{y}+\theta_{m}^{y}\theta_{a}^{m}+\theta_{l}^{y}\theta_{m}^{l}\theta_{a}^{m}+\sum_{i}\theta_{r_{i}}^{y}(\theta_{m}^{r_{i}}\theta_{a}^{m}+\theta_{l}^{r_{i}}\theta_{m}^{l}\theta_{a}^{m})\big{)}, where θij\theta_{i}^{j} denotes the coefficient associated with variable ii in modeling the variable jj, (?). Therefore, the constraint in (4) is an easy function to compute, and the resulting constrained optimization problem relatively easy to solve.

The PSE in the unconstrained model is 3.163.16. This means, the odds of having a high income would have been more than 3 times higher for a female if her sex and marital status would have been the same as if she was a male. We solve the constrained problem by restricting the PSE, as estimated by (9), to lie between 0.950.95 and 1.051.05. Accuracy in the unconstrained model is 82%82\%, and drops to 72%72\% in the constrained model while assuring that the constrained model is fair.

A very simple and naive approach to the problem that, superficially, might appear to be sensible for removing discrimination, is to drop the sensitive feature, which in this case results in the same accuracy as the unconstrained model. However, we believe dropping the sensitive feature from the model is a poor choice in our setting, both because sensitive features are often highly predictive of the outcome in interesting (and politicized) cases, and because dropping the sensitive feature does not in fact remove discrimination (as we defined it)!

Selecting The Outcome Model To Maximize Out Of Sample Predictive Performance

The search for an outcome model with the best out of sample performance, as is often done in machine learning problems, may result in a model which does not give consistent estimates of the NDE and thus does not guarantee removal of discrimination, if the NDE estimator is not chosen carefully. As discussed in the previous sections, the key approach is to use estimators that do not rely on the YY model being correctly specified, such as the triply robust estimator, and the IPW estimator. Here we demonstrate, via a simple simulation study, that selecting the outcome model to maximize predictive performance does not interfere with solving the constrained optimization problem for removing discrimination as long as we use triply robust or IPW estimators given that AA and MM models are specified correctly.

We generated 40004000 data points using the models shown in (10) and split the data into training and validation sets.

We assume AA and MM models are correctly specified; AA is randomized (like race or gender) and MM has a logistic regression model with interaction terms. Using the IPW estimator, which only uses AA and MM models, we obtain the NDE (on the ratio scale) of 3.013.01. As expected, the triply robust estimator, which uses A,MA,M, and YY models, gives us the same estimate of NDE, even under a misspecified YY model.

where D⊆C{\bf D}\subseteq{\bf C}. Consider, by contrast, what happened when we used an estimator that relied on the YY model being specified correctly. We pick the following (incorrect) YY model:

and compute the NDE using (1), to obtain the value of 2.72.7.

Performing the constrained optimization using the above model and the estimator in (1) would lead us to the optimal coefficients for the MM and YY models that ensure the NDE is within (−0.5,0.5)(-0.5,0.5), as desired. However, since the YY model was incorrect, and (1) was not robust to misspecification of YY, the results cannot be trusted. Indeed, using the constrained coefficients for MM and YY models in the triply robust estimator, that is robust to misspecified of YY, leads to a large NDE of 3.073.07.

The takeaway here is that the classical machine learning task of model selection to optimize out of sample prediction performance, be it via parameter regularization or other methods, can only ensure fairness if the estimators for the degree of fairness, as quantified by the PSE, do not rely on the model being selected, the models the estimators do rely on are specified correctly, and only the part of the model the estimators do not rely on is selected.

Discussion And Conclusions

In this paper, we considered the problem of fair statistical inference on outcomes, a setting where we wish to minimize discrimination with respect to a particular sensitive feature, such as race or gender. We formalized the presence of discrimination as the presence of a certain path-specific effect (PSE) (?; ?), as defined in mediation analysis, and framed the problem as one where we maximize the likelihood subject to constraints that restrict the magnitude of the PSE. We explored the implications of this view for predicting outcomes out of sample, for cases where the PSE of interest is not identified, and for computational Bayesian methods. We illustrated our approach using experiments on real datasets.

One of the advantages of our approach is it can be readily extended to concepts like affirmative action and “the wage gap” in a way that matches human intuition. To conceptualize affirmative action, we propose to define a set of “valid paths" from AA (race/sexual orientation) to YY (admission decision), perhaps paths through academic merit, or extracurriculars, or even the direct path, and solve a constrained optimization problem that increases the PSE along these paths. Here we mean placing a lower bound ϵl\epsilon_{l} on the PSE away from the value corresponding to “no effect". Then, we learn p∗p^{*} as the KL-closest distribution to the observed data distribution pp that satisfies the constraint on the PSE. Finally, we predict the admission decision of a new instance X\bf X in a similar way as the proposal in our paper, by using the information in the new instance X\bf X shared between pp and p∗p^{*}, and predicting/averaging over other information using p∗p^{*}. We thus “count the causal influence of the sensitive feature on admission via prescribed paths" more highly among disadvantaged minorities. Defining these paths is a domain-specific issue. Increasing the PSE potentially lowers predictive performance, just as decreasing the PSE did in our experiments on reducing discrimination. This makes sense since we are moving away from the PSE implied by the “unfair world" given by the MLE towards something else that we deem more “fair". A similar definition can be made for “the wage gap", which we believe should be meaningfully defined as a comparison of the PSE of gender on salary with respect to “inappropriate paths.”

One methodological difficulty with our approach is the need for a computationally challenging constrained optimization problem. An alternative would be to reparameterize the observed data likelihood to include the causal parameter corresponding to the discrimination PSE, in a way causal parameters have been added to the likelihood in structural nested mean models (?). Under such a reparameterization, minimizing the PSE always corresponds to imposing box constraints on the likelihood. However, this reparameterization is currently an open problem.

Acknowledgments

The research was supported by the grants R01 AI104459-01A1 and R01 AI127271-01A1. We thank James M. Robins, David Sontag and his group, Alexandra Chouldechova, and Shira Mitchell for insightful conversations on fairness issues. We also thank the anonymous reviewers for their comments that greatly improved the manuscript.

References