Fairness Under Unawareness: Assessing Disparity When Protected Class Is Unobserved
Jiahao Chen, Nathan Kallus, Xiaojie Mao, Geoffry Svacha, Madeleine Udell
Introduction
Models for high stakes decision making have ethical and legal needs to demonstrate a lack of discrimination with respect to protected classes (Munoz et al., 2016; Barocas and Selbst, 2016). Examples of such decisions include employment and compensation (Conway and Roberts, 1983; Greene, 1984), university admissions (Bickel et al., 1975), and sentence and bail setting (Berk et al., 2018; Chouldechova, 2017; Dressel and Farid, 2018). Another example relevant to the financial services industry is credit decisioning (Chen, 2018), which is a classification problem where these ethical concerns are enshrined in concrete regulatory compliance requirements. Credit decisions must be shown to comply with a myriad of federal and state fair lending laws, some of which are summarized in (Chen, 2018)In this paper, we restrict our discussion and citation of applicable laws to those of the United States of America.. Some of these laws define protected classes, such as race and gender, where discrimination on the basis of a customer’s membership in these classes is prohibited. Table 1 summarizes the protected classes defined by the Fair Housing Act (FHA) (US Congress, 1968) and Equal Credit Opportunity Act (ECOA) (US Congress, 4 10).
When demonstrating that credit decisions comply with these fair lending laws, we sometimes run into situations where fairness and bias assessments must be done on populations without knowing their memberships in protected classes, because it is illegal or operationally difficult to do so. For example, credit card and auto loan companies must demonstrate that the way they extend credit is not racially discriminatory, yet are not allowed to ask applicants what race they are when they apply for credit.Lenders may ask applicants to self-identify in a voluntary basis, with the understanding that the answer will not affect the outcome of the application and that the information is collected for compliance assessment only (Division of Consumer and Community Affairs, 1 07, 12 CFR §1002.5(b)). Similarly, health plans can only solicit race and ethnicity information for new members but cannot obtain the same information for existing members (Elliott et al., 2008). Given the lack of secure protocols that permit disparity evaluation with encrypted protected classes (Kilbertus et al., 2018), disparate impact assessments for these situations have to impute the mostly (or entirely) missing labels corresponding to the protected class, usually by relying on observed proxy variables that can predict class memberships. The imputed protected classes are then used by regulators in assessing disparate impact (but they are not allowed to be used in decision making). Generally, any model that imputes the missing protected attribute value based on other, observed variables is known as a proxy model, and such a model that is based on predicting conditional class membership probabilities is known as a probabilistic proxy model.
For example, for assessing adverse action with regard to race in credit decisions, regulators like the Bureau of Consumer Financial Protection (BCFP)Formerly the Consumer Financial Protection Bureau (CFPB). Citations and references reflect the name at time of publication. have been known in the past to use a probabilistic proxy model to impute the customers’ unknown race labels (Consumer Financial Protection Bureau, 2014). They used a naïve Bayes classifier, the Bayesian Improved Surname Geocoding (BISG) method, to predict the probability of race membership given the customer’s surname and address of residence (Elliott et al., 9 04). Specifically, assuming that surname and location are statistically independent given race, BISG uses Bayes’s rule to compute race membership probabilities from the conditional distributions of surname given race and of location given race as inferred by census data. This methodology (Consumer Financial Protection Bureau, 2014) notably supported a $98 million fine against a major auto loan lender (Consumer Financial Protection Bureau, 3 12). This case generated some controversy (Koren, 6 08; Committee on Financial Services, 5 11, 6 01, 7 01; Consumer Financial Protection Bureau, 8 05), in part due to empirical findings that the amount of disparate impact estimated by BISG appears to overestimate true disparities (Baines and Courchane, 4 11; Zhang, 6 01). However, the cause for this overestimation phenomenon is unknown, as is whether overestimation is to be expected always, or whether or not underestimation of disparate impact is also possible. This observation forms the motivation for our current work, which is broadly applicable to any fairness assessment where an unobserved protected class must be imputed using a proxy model. The aim is not to criticize the use of proxy models in general, but rather to provide a more informed analysis of the statistical biases inherent in any assessment where membership in protected classes must be imputed.
This paper investigates the bias in estimating demographic disparity (Definition 2.2) when a proxy model is employed to impute a protected class. We present the first theoretical results describing when the use of proxy models can lead to biased estimates of outcome disparity, which explains the overestimates observed in the past, and also offers insights on the practical use of proxy models. More specifically, our key contributions are:
We derive the (statistical) bias for the commonly used thresholded estimator, where a label is assigned only if the proxy model predicts a label with probability exceeding a predefined threshold (Consumer Financial Protection Bureau, 2014; Baines and Courchane, 4 11; Zhang, 6 01) (Definition 2.4, Theorem 3.3). We decompose its bias into multiple sources, which gives a set of interpretable conditions under which the thresholded estimator can over- or underestimate the outcome disparity.
We present a new weighted estimator for demographic disparity (Definition 2.5) that uses soft classification based on proxy model outputs as opposed to hard imputation. We derive its bias (Theorem 3.1) and find that the weighted estimator has only one bias source.
We validate our results on a public mortgage data set, using geolocation as the sole variable in a proxy model for race. We identify the specific source of bias that can account for the overestimation of the thresholded estimator, which can explain the overestimation bias of using proxy methods observed in previous literature.
We discover that the estimation bias is sensitive to the threshold used in class imputation based on the proxy model. This shows the intrinsic limitation of the thresholded estimator.
Evaluating the fairness of a binary decision
We have three main variables of interest:
with representing a favorable outcome, such as the approval of loan application or college admission offer, and representing an unfavorable outcome.
such as gender or race; often, we will write for the advantaged group and for the disadvantaged group.
a set of covariates taking values used to predict in a probabilistic proxy model.
We present only the binary case for simplicity. Unless otherwise stated, for multiclass , our results generalize straightforwardly to the pairwise outcome disparity between any advantaged group and any disadvantaged group. (Additional details are given in Appendices B and C.)
The mean group outcome for the group is
The demographic disparity, or Calders-Verwer gap (Calders and Verwer, 2010; Kamishima et al., 2012), , is the difference in mean group outcomes between the advantaged and disadvantaged groups:
A positive demographic disparity means that a higher proportion of the advantaged group receive a favorable outcome than the disadvantaged group , i.e. the disadvantaged group experiences an outcome disparity. Demographic disparity is simple to understand and is widely used (Lipton et al., 2017; Žliobaitė, 2015; Zafar et al., 2017), despite its flaws (Dwork et al., 2012; Hardt et al., 2016).
The ordinary approach is to predict a single value for class membership, (Baines and Courchane, 4 11; Zhang, 6 01):
Let . Then the thresholded estimated membership for the th unit is
where NA stand for an unclassified observation that is excluded from the subsequent outcome disparity evaluation.
Considering a unit with , this estimation rule can be summarized pictorially as follows
where dashed boxes represent correct (, green) and incorrect classifications (, red), and the middle is unclassified.
We can use these predicted labels to estimate the mean group outcomes and demographic disparity by mimicking the simple estimator for group sample means difference, but imputing for the unknown . This leads to the following estimator:
2. Weighted estimator
The form of the thresholded estimators above shows clearly the use of a hard classification rule. Under a probabilistic proxy model, this hard classification rule inevitably misclassifies some individuals, given the intrinsic uncertainty in classifying the protected class. Moreover, the threshold rule results in a group of unclassified individuals who are removed from the outcome disparity evaluation. To avoid these problems, we propose a new estimator that accounts for the soft classification generated by the proxy.
Let be defined as above. Then, the weighted estimators for mean group outcomes and demographic disparity are
3. Toy examples
In this part, we use two hypothetical toy examples to demonstrate the intuition regarding when the thresholded estimator (2) can overestimate the demographic disparity. In both examples, we consider evaluating the disparity of loan approval with respect to two races and . However, the true race is unknown, so the geolocation is used to estimate race. Suppose that there are only two neighborhoods where all people live in one neighborhood have high income and all people live in the other neighborhood have low income. We further assume that the high-income neighborhood is primarily occupied by the advantaged group and the low-income neighborhood is primarily occupied by the disadvantaged group . Therefore, the thresholding rule with classifies all people living in the high-income neighborhood as the advantaged group, i.e., , and all people living in the low-income neighborhood as the disadvantaged group, i.e., .
Consider an extreme lending policy: all people with high income get their loans approved, while all people with low income are rejected, no matter what their races are. See Table 2 for the illustration. Simple calculations show that the true demographic disparity . while the thresholded estimator .. Thus the thresholded estimator overestimates the demographic disparity. The reason for the overestimation is that the race proxy is correlated with the loan approval outcome because of the dependence between geolocation and socioeconomic status: people who live in the neighborhood primarily occupied by the advantaged group are also more likely to get loan approval because they have relatively high socioeconomic status, e.g. high income in this example. These people are classified as advantaged group because their neighborhoods are associated with high probability of belonging to the advantaged group. As a result, the thresholding rule misclassifies people who are from the disadvantaged group but get loan approval as the advantaged group. In this way, the thresholding rule leads to underestimates of the loan acceptance rate of the disadvantaged group. Analogously, the thresholding rule leads to overestimates of the loan acceptance rate of the advantaged group because it misclassifies people from the advantaged group but not likely to get loan approval as the disadvantaged group. Consequently, the misclassification of the protected attribute is dependent with the outcome, which ultimately leads to the overestimation of demographic disparity. The dependence between the misclassification and the outcome results from the inter-geolocation outcome variation: the socioeconomic status, race proxy probabilities (i.e., race ratios), and loan acceptance rates vary across different neighborhoods such that the favorable outcome is positively correlated with the probability of belonging to the advantaged group.
Now consider a hypothetical lending policy with affirmative action that approves the disadvantaged group with higher rate than the advantaged group with the same income level. But overall people with high income are still more likely to be accepted. See Table 3 for a concrete example. Simple calculations show that the true demographic disparity ., which means that the disadvantaged group overall has lower chance to get loan approval due to their population concentration in the low-income neighborhood. However, the thresholded estimator gives . that overestimates the demographic disparity. The reason for the overestimation is that the the disadvantaged group have higher approval rate than their neighbors from the advantaged group. As a result, misclassifying part of the disadvantaged group as the advantaged group raises the average approval rate for the advantaged group. Similarly, misclassifying part of the advantaged group as the disadvantaged group lowers down the average approval rate for the disadvantaged group. Consequently, the misclassification of the protected attribute is also dependent with the outcome and ultimately leads to overestimation of demographic disparity. In contrast to Example 2.6, the dependence between the misclassification and the outcome results from the intra-geolocation outcome variation: people from different protected groups living in the same locations have different chance of getting the favorable outcome.
In both examples, the protected attribute misclassification made by the thresholding rule is not uniformly at random. Instead, the misclassification shows systematic pattern with respect to the outcome either due to the inter-geolocation outcome variation (Example 2.6) or intra-geolocation outcome variation (Example 2.7). In Section 3, we will formalize the main intuition in these two examples and show that the interplay of the two bias sources captured by these two examples determines the overestimation or underestimation of the thresholded estimator. In Section 4, we will further identify the specific bias source that accounts for the estimation bias of the thresholded estimator in a mortgage data set.
Bias in thresholded and weighted estimators
In this section, we derive the asymptotic biases for the thresholded estimator (2) and weighted estimator (3) for demographic disparity. We also provide some interpretable sufficient conditions under which these two estimators overestimate or underestimate the demographic disparity.
Let be a binary protected class with values and . The bias of the weighted estimator in Definition 2.5 for demographic disparity in (1) is
where as , the biases in the weighted estimators for the mean group outcomes , for , converge almost surely to
We omit the (almost sure) notation later for brevity.
If is independent of conditionally on , then the weighted estimator for demographic disparity, , is asymptotically unbiased.
If is independent of conditionally on , then the advantaged group and disadvantaged group with the same values are equally treated in terms of the average outcome. Nevertheless, this does not contradict the existence of overall disparity against the disadvantaged group in terms of the unconditioned average outcome, as an example of Simpson’s paradox (Simpson, 1951). The conditional independence assumption required by Corollary 3.2 is trivially satisfied if is the output of some function , and includes the input features , since is now determined entirely by . Such a situation arises naturally when is the output of machine learning algorithms. In the Appendix C.4, we show a semi-synthetic example based on the mortgage dataset where the weighted estimator is asymptotically unbiased when the proxy model is well constructed so that the conditional independence assumption is satisfied.
2. Thresholded estimator
Before stating the asymptotic bias for the estimator (2), we define the following terms for :
where is the threshold used for estimating race, and is the class opposite to , i.e., if and if . measures the outcome mean discrepancy for two different protected groups within the same proxy probability range, and measures the outcome mean discrepancy across different proxy probability ranges for the same protected group. Consider the example of loan application with race proxy based on geolocation. In this example, measures the loan approval disparity between two race groups who live in locations that are primarily occupied by one of these two races (in terms of the threshold ). In contrast measures the loan approval rate disparity between people belonging to a race who live in locations that are primarily occupied by this race group and people belonging to this race who live in locations that are less occupied by this race. Therefore, roughly characterizes the intra-geolocation variation of loan outcome and roughly characterizes the inter-geolocation variation of loan outcome.
Let be a binary protected class with values and . The bias for the thresholded estimator in Definition 2.4 is:
and as , for and as the class opposite to ,
Here means that or .
The generalization to multiclass is straightforward and in Appendix C.1 we show that the bias formula in Theorem 3.3 for binary protected class capture the main effects of the bias for multiclass .
Theorem 3.3 shows that the occurrence of overestimation or underestimation depends on complex interplay of the terms. In the following corollary, we provide the simplest set of sufficient conditions for the overestimation and underestimation to demonstrate the main intuition.
and at least one inequality holds strictly. Then in the limit , , , and , i.e. the thresholded estimator overestimates the demographic disparity.
Conversely, let (i′) , (ii′) , (iii′) , and (iv′) , and at least one inequality holds strictly. Then in the limit , , , and , i.e. the thresholded estimator underestimates the demographic disparity.
The conditions (5i)—(5iv) are demonstrated in Figure 1.
Conditions (5i)—(5ii) exactly formalize the intuition captured by Example 2.7. In our loan application example, using a geolocation-based race proxy, these two conditions capture the intra-geolocation outcome variation: on average higher proportion of disadvantaged group receive the favorable outcome than the advantaged group among all the geolocations that are primarily occupied by one race (in terms of the threshold ).
The conditions (5iii)—(5iv) exactly formalize the intuition captured by Example 2.6. They characterize the dependence between the decision outcome and the proxy probability conditionally on the protected attribute. Specifically, condition (5iii) holds when the decision outcome is positively correlated with the probability of belonging to the advantaged group, while condition (5iv) holds when the decision outcome is negatively correlated with the probability of belonging to the disadvantaged group conditionally on the true protected attribute. In the example of loan application with race proxy based on geolocation, geolocation is correlated with both race ratio and socioeconomic status (e.g., income, FICO score, etc.): locations that are primarily occupied by the advantaged racial groups tend to be associated with relatively higher socioeconomic status while locations that are primarily occupied by the disadvantaged racial groups tend to be associated with relatively lower socioeconomic status. As a result, people who live in locations dominant by the advantaged racial group are assigned with high probability of belonging to the advantaged group, and they are more likely to get approved in loan application because of relatively higher socioeconomic status. Conversely, the loan approval is negatively correlated with the probability of belonging to disadvantaged groups. This example demonstrates that (5iii)—(5iv) are likely to hold when the predictors are also strongly correlated with the decision outcome, e.g., the geolocation is correlated with the loan approval because of the socioeconomic status disparities across different geolocations.
The following corollary shows that (5i)—(5iv) have effects with different magnitudes on the overestimation bias:
Let be a binary protected class with values and . For and as the class opposite to , the quantities in Theorem 3.3 are related in the following way:
Thus if the conditions (6) and (7) both hold, then and .
Numerical results
In this part, we simulate an example where geolocation is used to construct proxy for race. We use this example to demonstrate different sources of overestimation and underestimation bias of the weighted estimator and the thresholded estimator.
In each experiment, we set the total population size of neighborhoods to be 3000, 4000, and 5000 respectively. Each experiment is repeated times and the average estimated demographic disparity from the thresholded estimator, the weighted estimator, and using the true race is shown in Figure 2.
The income is normally distributed according to with mean value depending on both geolocation and race:
where controls the discrepancy of income between the two race groups within the same geolocation. We fix and vary from to , which indirectly varies the magnitude of intra-geolocation decision outcome variation between the two races. By construction, should affect only the in (5i)—(5ii).
The income is normally distributed according to with the mean value depending on only geolocation:
In other words, people with different races within the same geolocation have the same income distribution and thus the same decision outcome distribution. Therefore, there is no intra-geolocation decision outcome variation between the two races. As a result, when we vary from to , only inter-geolocation outcome variation changes, which affects in conditions (5iii)—(5iv).
While presented only for the particular choices of for Experiment 4.1.1 and for Experiment 4.1.2, the observation of which terms vary strongly with and hold true for quite a few different choices.
2. Estimation biases in the HMDA mortgage data set with geolocation proxy for race
In this section, we use the public HMDA (Home Mortgage Disclosure Act) data setData link: https://www.consumerfinance.gov/data-research/hmda/explore. to demonstrate demographic disparity estimation bias when using a probabilistic proxy for race. This data set contains mortgage loan application records in the U.S. for which the geolocation (state, county, and census tract), self-reported race/ethnicity, and loan origination outcome were reported, and has been used in the literature to evaluate the BISG race proxy (Baines and Courchane, 4 11; Consumer Financial Protection Bureau, 2014; Zhang, 6 01). We use the loan data for the years 2011—2012, consistent with (Consumer Financial Protection Bureau, 2014). We denote if a loan application was approved or originated, and if it was denied. The final sample contains around 17 million observations with non-missing geolocation, race, and loan origination outcome information.
We consider a race proxy based only on geolocation, as the public data set is anonymized and omits surnames. This proxy is derived from the racial and ethnic composition of the U.S. population that is over 18 years of age, using the census tracts of the 2010 decennial censusWe use the census tract level geolocation-only proxy constructed by the CFPB, as described in https://github.com/cfpb/proxy-methodology.. For example, the proportion of Hispanic, white, black, and API (Asian or Pacific Islander) in census tract 020100, Autauga County, Alabama are 0.02, 0.86, 0.10, 0.01 respectively, according to the 2010 census. This quadruple is assigned to all applicants from this census tract as their race proxy. This probabilistic proxy, while different from BISG, is nevertheless sufficient to demonstrate the general result regarding disparity estimation bias in Section 3. We present results for Hispanic, white, and black subpopulations in this section, and defer the result for API to the Appendix.
In Figure 3, we show the estimation bias of the thresholded estimator with different thresholds and the weighted estimator. Clearly the thresholded estimator underestimates the loan acceptance rate of black and Hispanic groups but it estimates the loan acceptance rate for White group accurately. As a result, the thresholded estimator overestimates the demographic disparity. Moreover, the overestimation bias tends to decrease as the threshold decreases. In contrast, the weighted estimator displays the opposite performance and underestimates the demographic disparity.
Bias source of thresholded estimator
Figure 4 shows the conditions (5i)—(5iv). Only (5iv) strictly holds, meaning that people from the disadvantaged group (Hispanic or black) living in census tracts where their race ratio exceeds the corresponding classification threshold have lower average loan acceptance rate than people from the disadvantaged group living in census tracts where their race ratio is not as high as the classification threshold. This captures the inter-geolocation outcome variation: the loan approval is negatively correlated with the disadvantaged group prevalence across the geolocations. In Example 2.6, we give a reason why loan acceptance is correlated with the race proxy probabilities: geolocation is correlated with both race proportions (i.e., race proxy probabilities) and socioeconomic status (e.g., income, FICO score, etc.) that affects loan approval. In Appendix C.3, we validate that such correlations are indeed obvious in the mortgage data set.
In contrast to conditions (5iv), conditions (5i)—(5ii) are strictly violated, and condition (5iii) is slightly violated. However, the thresholded estimators still have overestimation bias because (5iii)—(5iv) have dominant effects, as shown in Figure 5. Moreover, as the threshold increases, conditions (5iii)—(5iv) start to dominate, increasing the overestimation bias. Nevertheless, the apparent reduction of overestimation bias when using lower threshold does not mean that using lower threshold is better in practice. Instead, the reduction of overestimation bias is the consequence of delicate counterbalance between the two opposing bias sources: the violation of conditions (5i)—(5ii) contributes to underestimation bias and the conditions (5iii)—(5iv) contribute to overestimation bias. Thus, the smaller bias from using a lower threshold is not a robust finding. For example, for the experiments in Section 4, changing the threshold from to does not affect either the race imputation result or the demographic disparity estimation bias. This reflects the intrinsic complexity of the thresholded estimator.
In summary, according to Corollary 3.5, the fact that the thresholded estimation rule for race excludes many unclassified examples makes the overestimation predominantly be determined by the inter-geolocation outcome variation captured by conditions (5iii)—(5iv), especially when the threshold is very high. Moreover, strong racial segregation and socioeconomic status disparities across different geolocations make condition (5iii) or (5iv) (if not both) very likely to hold. As a result, the thresholded estimator tends to overestimate the demographic disparity, especially when high threshold is used. Since BISG also uses geolocation for race estimation, the reported overestimation bias of using thresholded estimator with BISG should be at least largely (if not all) due to the same reason. Moreover, the estimation bias of thresholded can be very sensitive to the threshold value because the threshold value influences the interplay of the bias sources in (5i)—(5iv). This reflects the intrinsic limitation of the thresholded estimator.
Bias source of the weighted estimator
Summary
The empirical results show that our theoretical analysis provides convincing explanations for the observed estimation bias of the thresholded estimator and the weighted estimator. We show that the bias sources of the weighted estimator and the thresholded estimator are very different: the bias of the weighted estimator is solely determined by the intra-geolocation variation, and the bias of the thresholded estimator is determined by both inter-geolocation variation and intra-geolocation variation. When high threshold is used, the inter-geolocation variation dominates. In the mortgage dataset, we show that the racial segregation, socioeconomic status, outcome disparity pattern with respect to geolocation make the thresholded estimator tend to overestimate the demographic disparity, and the weighted estimator tend to underestimate the demographic disparity. This explains the overestimation bias of the thresholded estimator with BISG reported by previous literature. Moreover, our results show that the estimation bias of thresholded estimator is sensitive to the threshold because of complex interplay of different bias sources. As a result, the observed estimation bias pattern in one setting may hardly generalize to other settings. In contrast, the weighted estimator only has one bias source, and is thus easier to reason about.
Conclusions
This paper presents the first theoretical analysis of bias in outcome disparity assessments using a probabilistic proxy model for the unobserved protected class. In Theorem 3.3, we derived the bias of a thresholded estimator (Definition 2.4) that has been described in the literature. We also gave sufficient conditions in Corollary 3.4 to understand when this methodology is biased, and to what extent. Our theoretical analysis is valid whenever a proxy model is used with a thresholded estimator to impute protected class membership, and is thus consistent with previous studies that had observed an overestimation bias of the thresholded estimator based on BISG. When applied to the public HMDA mortgage dataset with geolocation as the sole proxy for race membership (Section 4.2), we found that the estimation bias of the thresholded estimator depends on the complex interaction of multiple different biases, producing a strong sensitivity to the precise value of the threshold used. Further studies will be needed to demonstrate the robustness of this numerical finding, particularly when using other proxy models and other measures of fairness. Nevertheless, our work signals caution for choosing the ad hoc thresholds which are used in practice, as without the ground truth labels, we are unable to determine the optimal choice of threshold.
To alleviate the theoretical challenges of the thresholded estimator, we also proposed a weighted estimator (Definition 2.5), which propagated the uncertainty resulting from the probabilistic proxy onto the final estimand. We found that the estimation bias of this weighted estimator has only one bias source, arising from the loan approval discrepancy between difference races within the same geolocation, which led to an overall underestimate. As the behavior of this estimator’s bias is simpler to reason about, we believe that the weighted estimator may be a useful new method to incorporate into outcome disparity evaluations when proxy models are used, especially if the sign of estimation bias can be determined with external knowledge.
References
Appendix A Proofs for Section 3
By the strong law of large numbers, we have the almost sure convergence for the estimator
For the true mean group outcome, introduce the trivial conditional expectation with respect to :
Thus, we can regroup the terms in the difference to recognize the conditional covariance
The analogous results hold for everywhere. ∎
When is independent of conditionally on , it immediately follows that
Define the following events for :
where is the estimated protected attribute according to the thresholding rule (Definition 2.3).
where , , , , are defined analogously.
In summary, as , for and as the class opposite to , ,
where and are defined in Section 3.2, and
The conclusions follow immediately from Theorem 3.3. ∎
First, (i) is obvious according to the formulas of in Theorem 3.3.
Appendix B Multilevel unknown protected class
In this section, we suppose that the protected class can have more than two values and the value space of is denoted as . Denote , i.e., is the set of all values of other than and .
The quantities for in Theorem 3.3 are related in the following way:
Thus if the conditions in (i) and (ii) both hold, then and .
Appendix C Additional results on the mortgage dataset
C.2. Results for API
In Figure 8, we can observe that the estimation bias is quite different when estimating the demographic disparity for White and API. The thresholded estimator could overestimate or underestimate the demographic disparity while the weighted estimator only very slightly underestimate the demographic disparity. This is not totally surprising because the API population have very different geolocation distribution and socioeconomic status distribution than Hispanic and Black. API population tend to mix with other races in terms of living locations and the average socioeconomic status disparity between API and White is much smaller than the average socioeconomic status disparity between other minority groups and White. Figure 9 shows that the proportion of people who live in census tracts where API are accepted at lower average rate while White are accepted at higher average rate is much smaller than the proportion of people who live in census tracts where Black or Hispanic are accepted at lower average rate while White are accepted at higher average rate. This explains why the underestimation bias of the weighted estimator is very small.
C.3. Correlation between the race probability and socioeconomic status
Figure 10 shows the correlation between the the race proxy probabilities and yearly income or average loan acceptance rate. We can clearly observe that census tracts with high White probability overall have more mass on higher income and higher loan acceptance rate while census tracts with high Hispanic or Black probability have more mass on lower income and lower loan acceptance rate. This figure validates the fact that the geolocation encodes socioeconomic status disparities such that the race proxy probabilities are correlated with socioeconomic status variables (e.g., income in this example) and the loan approval outcome. These correlations account for the conditions (5iii)—(5iv) in the Corollary 3.4.