Fair prediction with disparate impact: A study of bias in recidivism prediction instruments

Alexandra Chouldechova

Introduction

Risk assessment instruments are gaining increasing popularity within the criminal justice system, with versions of such instruments being used or considered for use in pre-trial decision-making, parole decisions, and in some states even sentencing . In each of these cases, a high-risk classification—particularly a high-risk misclassification—may have a direct adverse impact on a criminal defendant’s outcome. If RPI’s are to continue to be used, it is important to ensure that they do not result in unethical practices that disparately affect different groups.

Within the psychometrics literature, there exist widely accepted and adopted standards for assessing whether an instrument is fair in the sense of being free of predictive bias. These standards have recently been applied to the COMPAS and PCRA instruments, with initial findings suggesting that there is evidence of predictive bias when it comes to gender, but not when it comes to race .

In a recent widely popularized investigation of the COMPAS RPI conducted by a team at ProPublica, a different approach to assessing instrument bias told what appears to be a contradictory story . The authors found that the likelihood of a non-recidivating Black defendant being assessed as high-risk is nearly twice that of White defendants. While this analysis has met with much criticism, it has also made headlines. There is no doubt that it is now embedded in the national conversation on the use of RPI’s.

In this paper we show that the differences in false positive and false negative rates cited as evidence of racial bias in the ProPublica article are a direct consequence of applying an instrument that is free from predictive biasin the psychometric sense to a population in which recidivism prevalence differs across groups. Our main contribution is twofold. (1) First, we make precise the connection between the psychometric notion of test fairness and error rates in classification. (2) Next, we demonstrate how using an RPI that has different false postive and false negative rates between groups can lead to disparate impact when individuals assessed as high risk receive stricter penalties. Throughout our discussion we use the term disparate impact to refer to settings where a penalty policy has unintended disproportionate adverse impact on a particular group.

It is important to bear in mind that fairness itself—along with the notion of disparate impact—is a social and ethical concept, not a statistical one. An instrument that is free from predictive bias may nevertheless result in disparate impact depending on how and where it is used. In this paper we consider hypothetical use cases in which we are able to directly connect statistically quantifiable features of RPI’s to a measure of disparate impact.

The empirical results in this paper are based on the Broward County data made publicly available by ProPublica . This data set contains COMPAS recidivism risk decile scores, 2-year recidivism outcomes, and a number of demographic and crime-related variables. We restrict our attention to the subset of defendants whose race is recorded as African-American (bb) or Caucasian (ww).

Assessing fairness

We begin by with some notation. Let S=S(x)S=S(x) denote the risk score based on covariates X=xX=x, with higher values of SS corresponding to higher levels of assessed risk. Let R∈{b,w}R\in\{b,w\} denote the group that the individual belongs to, which may be one of the components of XX. Lastly, let Y∈{0,1}Y\in\{0,1\} be the outcome indicator, with 11 denoting that the given individual recidivates. In this notation, we can think of the psychometric test fairness condition roughly as follows.

A score S=S(x)S=S(x) is test-fair (well-calibrated)Depending on the context, we may further desire that this criterion is satisfied when we condition on some of the covariates. Our analysis extends to this case as well. if it reflects the same likelihood of recidivism irrespective of the individual’s group membership, RR. That is, if for all values of ss,

Figure 1 shows a plot of the observed recidivism rates across all possible values of the COMPAS score. We can see that the COMPAS RPI appears to adhere well to the test fairness condition. In their response to the ProPublica investigation, Flores et al. further verify this adherence using logistic regression.

To facilitate a simpler discussion of error rates, we introduce the coarsened score ScS_{c}, which is obtained by thresholding SS at some cutoff sHRs_{HR}.

The coarsened score simply assesses each defendant as being at high-risk or low-risk of recidivism. For the purpose of our discussion, we will think of ScS_{c} as a classifier used to predict the binary outcome YY. This allows us to summarize ScS_{c} in terms of a confusion matrix, as shown below.

It is easily verified that test fairness of SS implies that the positive predictive value of the coarsened score ScS_{c} does not depend on RR. More precisely, it implies that that the quantity

A direct implication of this simple expression is that when the recidivism prevalence differs between two groups, a test-fair score ScS_{c} cannot have equal false positive and negative rates across those groups.This observation is also made in independent concurrent work by Kleinberg et al. .

This observation enables us to better understand why the ProPublica authors observed large discrepancies in FPR and FNR between Black and White defendants.Black: FPR = 45%, FNR = 28%. White: FPR = 23%23\%, FNR = 48%48\% The recidivism rate among back defendants in the data is 51%, compared to 39% for White defendants. Since the COMPAS RPI approximately satisfies test fairness, we know that some level of imbalance in the error rates must exist.

Assessing impact

In this section we show how differences in false positive and false negative rates can result in disparate impact under policies where a high-risk assessment results in a stricter penalty for the defendant. Such situations may arise when risk assessments are used to inform bail, parole, or sentencing decisions. In the state of Pennsylvania, for instance, statutes permit the use of RPI’s in sentencing, provided that the sentence ultimately falls within accepted guidelines. We use the term “penalty” somewhat loosely in this discussion to refer to outcomes both in the pre-trial and post-conviction phase of legal proceedings. Even though pre-trial outcomes such as the amount at which bail is set are not punitive in a legal sense, we nevertheless refer to bail amount as a “penalty” for the purpose of our discussion.

There are notable cases where RPI’s are used for the express purpose of informing risk reduction efforts. In such settings, individuals assessed as high risk receive what may be viewed as a benefit rather than a penalty. The PCRA score, for instance, is intended to support precisely this type of decision-making at the federal courts level. Our analysis in this section specifically addresses use cases where high-risk individuals receive stricter penalties.

To begin, consider a setting in which guidelines indicate that a defendant is to receive a penalty tL≤T≤tHt_{L}\leq T\leq t_{H}. A very simple risk-based approach, which we will refer to as the MinMax policy, would be to assign penalties as follows:

The expected difference in penalty under the MinMax policy is given by

We will discuss two immediate Corollaries of this result.

Among individuals who do not recidivate, the difference in average penalty under the MinMax policy is

Among individuals who recidivate, the difference in average penalty under the MinMax policy is

When using a test-fair RPI in populations where recidivism prevalence differs across groups, it will generally be the case that the higher recidivism prevalence group will have a higher FPR and lower FNR. From equations (3.2) and (3.3), we can see that this would result in greater penalties for defendants in the higher prevalence group, both among recidivating and non-recidivating offenders.

One might expect that differences in false positive rates are largely attributable to the subset of defendants who are charged with more serious offenses and who have a larger number of prior arrests/convictions. While it is true that the false positive rates within both racial groups are higher for defendants with worse criminal histories, considerable between-group differences in these error rates persist across low prior count subgroups. Figure 2 shows a plot of false positive rates across different ranges of prior count for defendants charged with a misdemeanor offense, which is the lowest severity criminal offense category. As one can see, differences in false positive rates between Black defendants and White defendants persist across prior record subgroups.

A natural question to ask is whether the level of disparate impact, Δ\Delta, is related to some measures of effect size commonly used in scientific reporting. With a small generalization of the % non-overlap measure, we can answer this question in the affirmative.

The % non-overlap of two distributions is generally calculated assuming both distributions are normal, and thus has a one-to-one correspondence to Cohen’s dd .d=Sˉb−SˉwSDd=\frac{\bar{S}_{b}-\bar{S}_{w}}{SD}, where SDSD is a pooled estimate of standard deviation. Figure 3 shows that the COMPAS decile score is far from being normally distributed. A more reasonable way to calculate % non-overlap is to note that in the Gaussian case % non-overlap is equivalent to the total variation distance. Letting fr,y(s)f_{r,y}(s) denote the score distribution for race rr and recidivism outcome yy, one can establish the following sharp bound on Δ\Delta.

Discussion

The primary contribution of this paper was to show how disparate impact can result from the use of a recidivism prediction instrument that is known to be free from predictive bias. Our analysis focussed on the simple setting where a binary risk assessment was used to inform a binary penalty policy. While all of the formulas have natural analogs in the non-binary score and penalty setting, we find that all of the salient features are already present in the analysis of the simpler binary-binary problem.

In closing, we would like to note that there is a large body of literature showing that data-driven risk assessment instruments tend to be more accurate than professional human judgements , and investigating whether human-driven decisions are themselves prone to exhibiting racial bias . We should not abandon the data-driven approach on the basis of negative headlines. Rather, we need to work to ensure that the instruments we use are demonstrably free from the kinds of quantifiable biases that could lead to disparate impact in the specific contexts in which they are to be applied.

References