Counterfactual Risk Assessments, Evaluation, and Fairness
Amanda Coston, Alan Mishler, Edward H. Kennedy, Alexandra Chouldechova
Introduction
Much of the activity in using machine learning to help address societal problems focuses on algorithmic decision-making and algorithmic decision support systems. In settings such as health, education, child welfare and criminal justice, decision support systems commonly take the form of risk assessment instruments (RAIs), which distill rich case information into risk scores that reflect the likelihood of the case resulting in one or more adverse outcomes. (Chouldechova et al., 2018; Kube et al., 2019; Ferguson, 2016; Kehl and Kessler, 2017; Stevenson, 2018; Caruana et al., 2015; Smith et al., 2012). Prior literature has raised significant concerns regarding the fairness, transparency, and effectiveness of existing RAIs (Barocas and Selbst, 2016; Barabas et al., 2017; Dressel and Farid, 2018; Corbett-Davies et al., 2017; Chouldechova and Roth, 2018). Yet RAIs remain very popular in practice, and there is a large body of research on fairness and transparency promoting methods that seek to address some of these concerns (e.g Zafar et al., 2015; Hardt et al., 2016; Kamiran and Calders, 2012; Pleiss et al., 2017; Kamishima et al., 2011; Zemel et al., 2013).
This paper highlights a different issue, one that has not received sufficient attention in the discussion of RAIs but that nonetheless has significant implications for fairness: RAIs are typically trained and evaluated as though the task were prediction when in reality the associated decision-making tasks are often interventions. Models trained and evaluated in this way answer the question: What is the likelihood of an adverse outcome under the observed historical decisions? Yet the question relevant to the decision maker is: What is the likelihood of an adverse outcome under the proposed decision? When decisions do not impact outcomes—when we are in what (Kleinberg et al., 2015) call a “pure predition” setting—these are one and the same. However, many decisions take the form of interventions specifically designed to mitigate risk. RAIs for these settings must be developed and evaluated taking into account the effect of historical decisions on the observed outcomes. Failure to do so will result in RAIs that, despite appearing to perform well according to standard evaluation practices, underperform on cases such as those that have been historically receptive to intervention.
In this paper we propose an approach to counterfactual risk modeling and evaluation to properly account for these intervention effects. Counterfactual modeling has been proposed for medical RAIs (Schulam and Saria, 2017; Shalit et al., 2017; Alaa and van der Schaar, 2017), and prior work has used counterfactual evaluation for off-policy learning in bandit settings (Dudík et al., 2011). However, the question of adapting counterfactual evaluation for risk assessments and in particular for predictive bias assessments remains open. In this paper, we propose a new evaluation method for RAIs that uses doubly-robust estimation techniques from causal inference (Van der Laan et al., 2003; Robins and Rotnitzky, 2001). We also argue that fairness metrics that are functions of the outcome should be defined counterfactually, and we use our evaluation method to estimate these metrics. We theoretically and empirically characterize the relationship between the standard fairness metrics and their counterfactual analogues. Our results suggest that in many cases, achieving parity in the standard metric will not achieve parity in the counterfactual metric.
Our main contributions are as follows: 1) We define counterfactual versions of standard predictive performance metrics and propose doubly-robust estimators of these metrics (§ 3); 2) We provide empirical support that this evaluation outperforms existing methods using a synthetic dataset and a real-world child welfare hotline screening dataset (§ 3); 3) We propose counterfactual formulations of three standard fairness metrics that are more appropriate for decision-making settings (§ 4); 4) We provide theoretical results showing that only under strong conditions, which are unlikely to hold in general, does fairness according to standard metrics imply fairness according to counterfactual metrics (§ 4); 5) We demonstrate empirically that applying existing fairness-corrective methods can increase disparity in the counterfactual redefinition of the metric they target (§ 4).
Background and Related Work
Literature on contextual bandits has considered counterfactual learning and evaluation of decision policies. While this literature is methodologically relevant, as we discuss below, it addresses a different problem. In the decision support setting we are considering, human users will ultimately decide what action to take. The goal of the learning and evaluation task is not to learn a decision policy, but rather to learn a risk model that will inform human decisions. That is, the risk assessment task is to accurately and fairly estimate the probability of an outcome under a given intervention.
Prior work has considered counterfactual RAIs in a temporal setting (Schulam and Saria, 2017). In this work, the trained model is evaluated on real data using the observed outcomes, and on simulated data. Evaluating against the observed outcomes can be misleading in settings in which treatment was not assigned randomly (see § 3.3.3). In our work we propose instead to adapt DR techniques, as have been used in the bandit literature for evaluating policies, to provide evaluations of counterfactual RAIs.
Counterfactual learning in the causal inference literature uses model selection based on DR estimation of counterfactual loss (Van der Laan et al., 2003). Whereas this approach evaluates counterfactual metrics implicitly, our approach does so explicitly, providing the estimators for standard classification metrics in § 3.3.3.
There is also a line of work focused on counterfactual learning in the presence of hidden confounders. (Kallus and Zhou, 2018a) propose policy learning via minimax regret learning over uncertainty sets. Their method is not immediately applicable to decision-support settings where RAIs are more informative to decision-makers than a policy recommendation. (Madras et al., 2019) propose using deep latent variable models to model hidden confounders via proxies in the data and evaluate how well this model learns an optimal policy. While their model may be used for learning a risk assessment model, they do not address how to evaluate the model in such a setting, which is the focus of our work. § 3 of our paper assumes no hidden confounders, and future work could attempt to incorporate these techniques for handling hidden confounders. We note that our theoretical analysis in § 4 holds even in the presence of hidden confounding.
2. Fairness and causality
A growing literature on counterfactual fairness has offered notions of fairness based on the counterfactual of the protected attribute (or its proxy) (Kusner et al., 2017; Wang et al., 2019; Kilbertus et al., 2017). In this work, a policy is considered fair if it would have made the same decision had the individual had a different value of the protected attribute (and hence, potentially different values of features affected by the attribute). In this setting, the treatment decision is the outcome, and the protected attribute is the ‘treatment’. By contrast, we consider counterfactual treatment decisions and consider a future observation to be the outcome.This distinction is also made in a survey of fairness literature (Mitchell et al., 2018).
Another line of work considers unfair causal pathways between the protected attribute (or its proxy) and the outcome variable or target of prediction (Nabi and Shpitser, 2018; Zhang and Bareinboim, 2018b). These papers characterize or explain discrimination via path-specific effects, which are defined by interventions on the protected attribute. We do not consider interventions on (i.e. counterfactuals of) the protected attribute; rather, we propose methods that account for interventions on treatment decisions in training and evaluation.
Fairness definitions based on the counterfactual of the protected attribute are not widely used in RAI settings for two reasons: one technical and one practical. The technical challenge is that the assumptions required to estimate these counterfactual metrics prohibit the use of important features, such as prior history, or require full specification of the structural causal model (SCM) (Zhang and Bareinboim, 2018a; Kusner et al., 2017, 2019) These requirements are too restrictive for our settings of interest where we have insufficient domain knowledge to construct the SCM and where we are unable to disregard important predictors like prior history. More significantly, the practical concern is that these definitions are ill-suited for risk assessment settings like child welfare screening. As we discuss in § 4, decisions made based on the counterfactual protected attribute may cause further harm to the protected groups.
Our work bears conceptual similarity to the analysis of residual unfairness when there is selection bias in the training data that induces covariate shift at test time as discussed in (Kallus and Zhou, 2018b). In settings where cases are systematically screened out from the training set, such as loan approvals in which we do not get to see whether someone who was denied a loan would have repaid, they find that applying fairness-corrective methods is insufficient to achieve parity. We consider a different but related setting in which we observe outcomes for all cases, but these outcomes are under different treatments. We propose fairness definitions that account for the effect of these treatments on the observed outcomes, and analyze the conditions under which existing methods can achieve this notion of counterfactual fairness.
Counterfactual Modeling and Evaluation
Before proceeding to introduce the learning approaches and evaluation methods considered in this work, we pause to clarify the types of risk-based decision policies to which our evaluation strategy as presented is tailored, and provide some background on algorithm-assisted decision making in child welfare hotline screening.
RAIs typically inform human decisions either by identifying cases that are the most (or least) risky, or by identifying cases that are the most (or least) responsive. The evaluation metrics we consider are most directly relevant in the paradigm where human decision-makers wish to intervene on the riskiest cases. However, our method can readily be adapted (as discussed in § 3.3) for paradigms in which interventions are being targeted based on responsiveness.
The motivating application for our work is child welfare screening. Child welfare service agencies across the nation field over 4.1 million child abuse and neglect calls each year (U.S. Department of Health & Human Services, 2019). Call workers must decide whether to “screen in” a call, which refers to opening an investigation into the family. The child welfare system is responsible for responding to all cases where there is significant suspicion that the child is in present or impending danger. The standard of practice is therefore to identify the riskiest cases. Jurisdictions in California, Colorado, Oregon and Pennsylvania are all in various stages of developing and integrating RAIs into their call screening processes. The RAIs are trained on historical data to predict adverse child welfare outcomes, such as re-referral to the hotline or out-of-home foster care placement (Chouldechova et al., 2018). The decision to investigate a call can affect the likelihood of the target outcomes.
2. Learning models of risk
In this section we introduce “observational” (standard practice) and “counterfactual” forms of model training.
2.2. Counterfactual
Consistency: . This assumes there is no interference between treated and control units. This is a reasonable assumption in the child welfare setting since opening an investigation into one case will not likely affect another case’s observed outcome.We set the treatment to be the same value for all children in a family.
Exchangeability: . This assumes that we measured all variables that jointly influence the intervention decision and the potential outcome . This is an untestable assumption but it may be reasonable in the child welfare setting where the measured variables capture most of the information the call screeners use to make their decision (see Section 3.4.2 for more details).
3. Evaluation
To evaluate how well our models of risk might inform decision-making in the paradigm where interventions should be targeted at the riskiest cases, we assess performance metrics such as precision, true positive rate (TPR), false positive rate (FPR), and calibration.In the paradigm where interventions are to be targeted at the most responsive cases, performance metrics such as discounted cumulative gain (DCG) or Spearman’s rank correlation coefficients are more natural choices for evaluation. DR estimates can be constructed for these metrics as well. Since the task is to evaluate how well the model predicts risk under a baseline intervention, we specify the performance metrics in terms of . The target counterfactual TPR is
A model is well-calibrated in the counterfactual sense when
where define a bin of predictions. We describe two standard practice approaches for evaluation, noting why these approaches do not adequately estimate the counterfactual targets. We introduce our proposed approach that uses doubly robust (DR) estimation.All evaluations are computed on a test partition that is separate from the train partition
3.2. Evaluation on the Control Population
3.3. Doubly-robust (DR) Counterfactual Evaluation
where denotes the score of our counterfactual model. The IPW estimate uses the observed outcome on the control population and reweighs the control population to resemble the full population:
DR estimatorsIn survey inference, this is known as the generalized regression estimator (Särndal et al., 1989). combine the plug-in estimate with an IPW-residual bias-correction term for the control cases:
Next we consider the counterfactual targets in Equations 1- 4. We identify the target under our causal assumptions and then state the DR estimator. We emphasize the distinction that is the score of any model we wish to evaluate whereas is the score of our counterfactual model in § 3.2.2.
The DR estimate for the denominator is in Equation 5.
The target counterfactual precision is identified as
The target in Equation 4 is identified as
Then we use the normal approximation to compute the interval: where for a 95% confidence interval.
The target counterfactual FPR is identified as
For the denominator we use where is in Eq 5.
4. Results
Figure 1 displays PR, ROC, and calibration curves.The code for this experiment is given in https://github.com/mandycoston/counterfactual DR evaluation most closely aligns with the true counterfactual evaluation. Notably, the observational evaluation suggests that the observational model outperforms the counterfactual model whereas the true counterfactual evaluation shows the counterfactual model performs better.
4.2. Child Welfare
We also apply counterfactual learning and evaluation to the problem of child welfare screening. The baseline intervention is screen-out (which means no investigation occurs). The data consists of over 30,000 calls to the hotline in Allegheny County, Pennsylvania, each containing more than 1000 features describing the call information as well as county records for all individuals associated with the call. The call features are categorical variables describing the allegation types and worker-assessed risk and danger ratings. The county records include demographic information such as age, race and gender as well as criminal justice, child welfare, and behavioral health history. The outcome is re-referral within a six month period. Our approach contrasts to prior work which used placement out-of-home as the outcome (Chouldechova et al., 2018; De-Arteaga et al., 2018). This outcome is only observed for cases under investigation; therefore it cannot be used to identify , the risk under no investigation.
We use random forests to train the observational and counterfactual risk assessments as well as the propensity score model. We used reweighing to correct for covariate shift but did not observe a boost in performance, likely because we have sufficient data and we used a non-parametric model.
We present the PR, ROC and calibration curves in Figure 2. The observational evaluation suggests that the observational model performs better. The control evaluation suggests that the counterfactual and observational models of risk perform equally well. Our DR evaluation suggests the counterfactual model has both better discrimination and calibration in estimating the probability of re-referral under screen-out. In Figure 2(c), the observational evaluation suggests that the observational model is well-calibrated whereas the counterfactual model is overestimating risk; this is expected because the counterfactual model assesses risk under no investigation whereas the observed outcomes include cases whose risk was mitigated by child welfare services. The control evaluation suggests that the two models are similarly calibrated. The DR evaluation shows that the counterfactual model is well-calibrated and the observational model underestimates risk. This makes intuitive sense because the observational model is not accounting for that fact that treatment reduced risk for the screened-in cases.
We see further evidence that the observational model performs poorly on the treated population in the drop in ROC curves between the control evaluation and DR evaluation in Figure 2(b). Deploying such a model would mean failing to identify the people who need and would benefit from treatment. The observational and control evaluations do not show this significant limitation; DR evaluation is the only evaluation that illustrates the poor performance of the observational model on the treated population.
We also evaluate the different models according to whether they are equally predictive, in the sense of being equally well calibrated, across racial groups. Research suggests child welfare processes may disproportionately involve black families (Dettlaff et al., 2011). Here we ask whether the observational or counterfactual model is more equitable. We compare calibration rates by race in Figure 3. The observational evaluation suggests that the counterfactual model of risk is poorly calibrated by race. The DR evaluation shows that the counterfactual model is well-calibrated by race and indicates that the observational model underestimates risk on both black and white cases.
Overall the observational evaluation suggests that the observational model performs better whereas the DR evaluation suggests the counterfactual model performs better. Since we do not have access to the true counterfactual to validate these results, we further consider how well the models align with expert assessment of risk.
4.3. Expert Evaluation
At various stages in the child welfare process, social workers assign treatment based on their assessment of risk. Social workers sequentially make three treatment decisions:
Whether to screen in a case for investigation
Whether to offer services for a case under investigation
Whether to place a child out-of-home after an investigation
Assuming that social workers are competent at assessing risk, we expect the group placed out-of-home (3) to have the highest risk distribution, followed by the group offered services (2), followed by those screened in, and finally we expect the screened out group to have the lowest risk. Figure 4 shows that the counterfactual model exhibits this expected behavior whereas the observational model does not. The observational model assesses the screened out population to have more high risk cases than any other treatment group. This indicates that the observational model is underestimating risk on the treated groups (investigated, services, and placed) since it fails to account for the risk-mitigating effects of these treatments. The observational model underestimates risk on those who were assigned effective treatments. These cases should be assigned treatment, but the observational model would suggest that they are low risk and should be screened out.
Such a mistake can have cascading effects downstream. We are particularly concerned about screening out cases that, had they been screened in, would have been accepted for services or placed out-of-home. Figure 5 shows the recall for placed cases and serviced cases as we vary the proportion of cases classified as high-risk. This plot shows that at any proportion the counterfactual model has significantly higher recall for both services and placement cases.
4.4. Task adaptation: Predicting Placement
Table 1 shows the area under the ROC and PR curves for the placement task. The observational model performs worse than a random classifier, whereas the counterfactual model shows some degree of discrimination. This suggests that the counterfactual model is learning a risk model that is useful in related risk tasks whereas the observational model is not.
The comparison to expert assessment of risk and the performance on a downstream risk task support the conclusions of our DR evaluation: the counterfactual model outperforms the observational model. In decision-making contexts, failure to account for treatment effects can lead one to the wrong conclusions about model performance, even potentially leading to the deployment of a model that underestimates risk for those who stand to gain most from treatment. In the next section, we consider how failure to account for treatment effects can impact fairness.
Counterfactual Fairness
Standard observational notions of algorithmic fairness are subject to the same pitfalls as observational model evaluation. In this section we propose counterfactual formulations of several fairness metrics and analyze the conditions under which the standard (observational) metric implies the counterfactual one.
We motivate the importance of defining these metrics counterfactually with an example. Suppose teachers are assessing the effectiveness and fairness of a model that predicts who is likely to fail an exam which they intend to use to assign tutoring resources. Suppose anyone tutored will pass. The tutoring session conflicts with girls’ sports practice so only male students are tutored. A model that perfectly predicts who will fail without the help of a tutor will have a higher observational FPR for men than women because some male students were tutored, which enabled them to pass. It would be wrong to conclude that this model is unfair with regards to FPR. Someone who would have been high-risk had they not been treated but whose risk was mitigated under treatment should not be considered a false positive. Failure to make this distinction could lead to unfairness, not only in settings where the treatment assignment varies according to the protected attribute but also in settings where the risk under treatment varies according to the protected attribute, as we can see in the next example.
Suppose that the classroom next door is also evaluating the model. This classroom offers tutoring during lunch so girls and boys both can attend; however they hired a tutor who happens to only be effective in preparing male students to pass. The teachers don’t know this and randomly assign this tutor to students regardless of gender. The model that perfectly predicts who will fail without a tutor has a higher observational FPR for men, but as before, it is wrong to conclude that the model is unfair with regards to FPR.
We distinguish our notion of counterfactual fairness from prior work which considered counterfactuals of the protected attribute (Kusner et al., 2017; Kilbertus et al., 2017; Wang et al., 2019), an approach which is counterproductive in our settings of interest. Consider a female student who is at high risk of failing because of gender discrimination at home or in the classroom e.g. parents or previous teachers have not given her the support they would have had she been male. Treating this student ”counterfactually as if she had been male all along” may suggest that we should not assign this student a tutor. In fact we must assign her a tutor in order to correct historical discrimination. Similar arguments can be made in settings like child welfare screening and loan approvals.
For three definitions of fairness (parity), we show that observational parity implies counterfactual parity if and only if a balance condition holds. We further show that an independence condition is sufficient for observational parity to imply counterfactual parity. We discuss why it is generally unlikely that the independence condition holds and even more unlikely that the finer balance condition holds when the independence condition fails. All proofs are provided in Appendix B.
BalBP holds under the following independence conditions, which provide sufficient conditions for oBP to imply cBP.
1.2. Predictive parity
Base parity and demographic parity may be ill-suited for settings where base rates differ by protected attribute due to disparate needs. Here we may instead desire parity in an error metric, such as precision. Positive predictive parity requires the precision (also known as positive predictive value) to be independent of the protected attribute, and negative predictive parity requires the negative predictive value to be independent of the protected attribute (Chouldechova, 2017; Kleinberg et al., 2016). We define observational Predictive Parity (oPP) as and counterfactual Predictive Parity (cPP) as where corresponds to negative predictive parity and corresponds to positive predictive parity.
BalPP is satisfied under the following independence conditions, which provide sufficient conditions for oPP to imply cPP.
IndPP will not hold in many settings. Note that and . Conditions require to contain all the information that tells us about treatment assignment that is not contained in . Since is typically trained to predict and not , it is quite unlikely that these conditions will hold in settings where there is bias in treatment assignment even when controlling for true risk. Condition allows differences in the risk distribution under treatment if we can fully explain these differences with . In the best case , but it is unlikely that the observed outcome, which is not causally well-defined, would explain differences in the risk distribution under treatment. As above, even if indPP does not hold, balPP may hold but it is difficult to reason why this should hold in any setting. Like Theorem 1, Theorem 2 also assumes a mild positivity-like assumption that is reasonable in risk assessment settings.
1.3. Equalized odds
In settings where TPR and FPR are more important than predictive value, we may desire parity in TPR and FPR, a fairness notion known as Equalized Odds (Hardt et al., 2016). Let observational Equalized Odds (oEO) require that and counterfactual Equalized Odds (cEO) require that .
The balance condition is satisfied under the following independence conditions, which comprise sufficient conditions for oEO to imply cEO.
The first two conditions of indEO require oBP and cBP, so indEO requires balBP to hold. In settings where there is discrimination in treatment assignment even when controlling for true risk, indEO is unlikely to hold. Even if there is no such discrimination, indEO will not hold if there are differences in the risk distributions under treatment since the last condition of 18 requires . indEO requires further conditions such as parity in the TPR/FPR against the outcome under treatment. If these conditions are not met, oEO could imply cEO if balEO holds, but it is difficult to reason about why this would hold for a setting when the independencies do not. Theorem 3 assumes two mild assumptions: the positivity-like assumption of Theorem 2 and .
Our theoretical analysis suggests that in many settings equalizing the observational fairness metric will not equalize the counterfactual fairness metric. We conclude by noting that the theorems hold when conditioning on any feature(s) , and in this context, these theorems are relevant to individual notions of fairness.
2. Experiments on synthetic data
We empirically demonstrate that equalizing the observational metric via fairness-corrective methods can increase disparity in the counterfactual metric on the synthetic data described in § 3.4.1.We do not perform the experiments on the child welfare data since it is balanced in terms of base rates and FPR/TPR with respect to race.
One approach to encourage demographic parity reweighs the training data to achieve base rate parity (Kamiran and Calders, 2012). Figure 6 shows that without any processing (“Original”), the counterfactual base rates are equal while the observational base rates show increasing disparity with . Reweighing applied to the observational outcome achieves oBP but induces disparity in the counterfactual base rate. Theorem 1 suggested this result: For , ; then it is unlikely that oBP implies cBP.
2.2. Post-processing for equalized odds
Table 2 shows that post-processing to equalize oGFPR and oGFNR induces imbalance in cGFPR and cGFNR.We use and . We report results for other values in Appendix E. In Figure 7 we see that the original model achieved cEO, but post-processing induced disparity to the detriment of the group that was less likely to be treated. Since treatment is beneficial, this “fairness” adjustment actually compounded the discrimination in the treatment assignment.
Conclusion
This paper demonstrates that training and evaluating models using observed outcomes can lead to the misallocation of resources due to the misestimation of risk for those most receptive to treatment. Furthermore, fairness-correcting methods that seek to achieve observational parity can lead to disparities on the relevant counterfactual metrics, and may further compound inequities in intial treatment assignment. The counterfactual approaches to learning, evaluation and predictive fairness assessment introduced in this paper provide more accurate and relevant indications of model performance.
References
Appendix A Identifications
In this section we give the identifications referenced in § 3.2.2 and § 3.3.3. Identification is the process of writing a counterfactual quantity in terms of observable quantities, based on causal assumptions. Our identifications rely on the causal assumptions of § 3.2 and assume that the model is learned and evaluated on separate train/test partitions (so the predictions are just a function of features ).
where the first line used the exchangeability assumption in § 3.2.2 and the second line used the consistency assumption in § 3.2.2.
A.2. Identification of Counterfactual TPR
In § 3.3.3, we identify the counterfactual TPR (or recall) as
The derivation is as follows. By definition of conditional expectation we have
We separately identify the numerator and denominator. Since we are evaluating on a test partition, is a function of . Then for the numerator we have
where the first line used iterated expectation, the second line used the definition of an indicator function, the third line use the fact that , the fourth line used exchangeability, and the fifth line used consistency.
To identify the denominator, we use iterated expectation and then apply exchangeability and consistency as we did for the counterfactual target:
A.3. Identification of Counterfactual Precision
In § 3.3.3, we identify the counterfactual precision as
where the first line uses iterated expected, the second line applies the fact that is just a function of , the third line uses exchangeability (from § 3.2.2) and the last lines uses consistency (from § 3.2.2).
A.4. Identification of Counterfactual Calibration
The derivation is the same as for precision since is just a function of .
A.5. Identification of Counterfactual FPR
In § 3.3.3, we identified counterfactual FPR as
The below derivation is similar to that of TPR. We can rewrite the target as
We separately identify the numerator and denominator. For the numerator we have
where the first line used iterated expectation, the second line used the definition of indicator function, the third line used the fact that is a binary random variable, the fourth line used exchangeability, and the last line used consistency.
where the second line used the derivation for the denominator in TPR.
Appendix B Proofs
In this section we give the proofs for theorems in § 4.1. These proofs assume consistency (defined in § 3.2.2).
Sufficient: In addition to oBP, we assume balBP holds. The left-hand sides of balBP (Equation 20) and oBP (Equation 19) are the same. Then applying the transitive property,
The following conditions are sufficient for oBP to imply cBP: and
If indBP holds, then balBP holds: Since , the LHS of Equation 20 is zero. By contraction is equivalent to and ; then the RHS of Equation 20 is also zero, so balBP holds. Since balBP is sufficient, indBP is sufficient for oBP to imply cBP. ∎
The proofs use the same techniques as for base rate parity.
In addition to oEO, we assume balEO holds. From oEO we have equation 26 and balEO is equation 27. The left-hand sides of equations 26 and 27 are the same so by the transitive property,
The following conditions are sufficient for oEO to imply cEO:
By contraction, the indEO conditions are equivalently written as ; ; ; ; ; . Under these assumptions, both sides of Equation 27 are 0, so the balEO condition holds under these independencies. Since balEO is sufficient, then indEO is sufficient for oEO to imply cEO. ∎
Appendix C Child Welfare Evaluation via Task adaptation on Services
In § 3.4.4, we assessed how well our models of risk performed on the related risk task of predicting out-of-home placement. Another related risk task considers the decision to accept a case for services. A case can only be accepted for services if it is under investigation, and we assume that the child welfare process assigns services based on the assessment of risk of child harm. Table 3 shows the AuROC and AuPR for the services task. As for the placement task, the observational model performs worse than random, and the counterfactual model performs better than random.
Appendix D Further Synthetic Evaluation
We supplement the empirical analysis in § 3.4.1 by presenting the PR, ROC, and calibration curves for several values of , the parameter describing treatment effect, and , the parameter describing treatment assignment bias. Each column denotes an evaluation method as described in Section 3.3.
In § 3.2, we presented the observational and counterfactual models of risk and we motivated why the observational model does not appropriately account for treatment effects. One might wonder if including treatment in the regression could control for treatment effects. In Figures 11 and 12, we present the PR, ROC, and calibration curves when including as a feature in the observational model for two values of , the parameter describing treatment effect. The true counterfactual evaluation shows that the observational model underperforms relative to the counterfactual model. This indicates that including as a feature does not appropriately control for treatment effects for the purposes of building RAIs.
Evaluation based only on the control or observational evaluation methods would lead to the wrong conclusions about model performance. Comparing the control and true counterfactual evaluations, we see that the observational model severely underperforms on the treated population. The control evaluation is misleading because it does not account for the poor performance on the treated population; only our DR evaluation (and the true counterfactual evaluation which is not feasible in practice) highlight this significant limitation of the observational model.
Appendix E Fairness-Corrective Methods on Synthetic Data
We supplement the empirical analysis of § 4.2 by displaying the results as we vary , the parameter describing treatment effect, and , the parameter describing treatment assignment bias. We show the counterfactual ROC curves as well as the table showing the generalized observational and counterfactual FPR and FNR using as the threshold.