De-biasing "bias" measurement
Kristian Lum, Yunfeng Zhang, Amanda Bower
Introduction
In the interest of developing models that behave equitably for different swaths of the population, identifying and mitigating group-wise disparities in model performance–typically measured in terms of accuracy, false positive rates, and so on–have become a central theme of creating “fair” models. Models that exhibit group-wise performance disparities are often termed “biased”To be clear, we will refer to this type of “bias” as “group-wise model performance disparities” or “performance disparities” for short throughout our paper.–an overloaded term, which in this case, simply means that that model performance is substantively different for some groups than others. Canonical examples of models expressing between-group performance variability include the identification of racial disparities in false positive rates of recidivism prediction models (Angwin et al., 2016), race-gender disparities in accuracy rates in gender classification models (Buolamwini and Gebru, 2018), and racial disparities in police allocation in predictive policing (Lum and Isaac, 2016). In each of these cases, model performance disparities were readily apparent simply by inspecting the model performance for each group separately and judging that the magnitude of difference between the groups was unacceptably large. When there are a small number of groups, each composed of many observations–as was the case in all of the above examples–this method of measuring and evaluating group-wise model performance disparities is possible.
When the number of groups across which performance disparities is being evaluated is large, it is difficult for a human to make judgments about the degree of disparities present simply by inspecting the model performance for all groups. Human-understandable measurement of model performance disparities and uncertainty quantification thereof are important if decision-makers are to engage in a well-informed consideration of the trade-offs between the relative merits of a variety of candidate models, determine whether a particular model should be mitigated for performance disparities, and decide whether any model should be deployed at all. To illustrate, suppose a machine learning practitioner needs to select from among three candidate binary classification models that use text as an input. Figure 1 shows model accuracy broken down by the language of the text and age of the text’s author for three hypothetical models.Code to produce this figure, all simulations, and analysis in this paper is available at https://github.com/twitter-research/double-corrected-variance-estimator The height of each bar represents each model’s accuracy within the corresponding group. Which model exhibits the lowest group-wise performance disparities by age? Do any of them exhibit acceptable levels of model performance disparities with respect to age? Given a model, are there more disparities with respect to age or language? How does the group-wise performance disparities between models compare? While it is difficult to answer these questions just with three candidate models, in real-world settings, the situation can be drastically more difficult. Machine learning practitioners may be considering hundreds or thousands of models given the plethora of hyperparameter choices when training neural networks and other models. Answering these questions requires well-designed, summary metrics that transform high-dimensional outputs of model performance on each group into lower-dimensional summaries while retaining the salient information. We call these “meta-metrics” since they are a metric on the vector whose coordinates contain model performance metrics on each group. We use this term to differentiate them from the “base metric”, which is the performance metric applied to each group separately .
Meta-metrics have received little attention in the algorithmic fairness literature despite their ubiquity and importance. Specifically, there have been no investigations of whether the metrics that are used to measure group-wise model performance disparities are themselves biased in the statistical estimation sense of the word. Uninterpretable, statistically biased meta-metrics can lead to costly outcomes. For example, suppose we obtain measurements that incorrectly suggest a model has a large degree of group-wise model performance disparities and thus requires mitigation. If we retrain the model with additional “fairness” constraints, we will likely end up lowering the utility of the model for every group involved unnecessarily. On the other hand, if the measurement is not easily understood, e.g. on a small scale that does not lend itself to distinguishing different scenarios, performance disparities may go undetected. Unfortunately, as we show in Section 4, many of the meta-metrics currently in use are statistically biased representations of the underlying quantity they purport to represent. The statistical bias arises because they fail to account for statistical uncertainty/sampling variance in the base metrics across which they are summarizing information. In particular, they are statistically biased upwards, exaggerating the extent of performance disparities, potentially leading to mitigation in cases where model performance is negligible or even identical across groups. The statistical bias is particularly large when some sub-groups have large sampling variance, e.g. in situations where some groups consist of few individuals. This situation is common: when evaluating model performance disparities across the intersection of many group variables, as in (Kearns et al., 2018), even large data sets can have intersectional groups that consist of very few observations. For example, in the Adult income datasetDownloaded from the UCI Machine Learning Repository at https://archive.ics.uci.edu/ml/datasets/adult.– a commonly used benchmark dataset for evaluating model performance disparities– some intersectional groups defined by race and gender contain fewer than 50 individuals. Many intersectional groups defined by race and 10 year age bins contain only one individual per group.
Furthermore, typical approaches to quantifying statistical uncertainty of existing meta-metrics, e.g. bootstrapping, can fail to cover the true degree of between-group disparities at a much higher than expected rate because they capture uncertainty about a quantity that is itself a statistically biased estimator of the true underlying value. This failure is more pronounced when the statistical bias is large. Therefore, using existing tooling on problems with many groups, some of which are small, results in either measurements that are so statistically biased as to paint a misleading picture of between-group disparities, or using heuristics to remove small-sized groups, thereby potentially further marginalizing already marginalized groups by excluding them from group-wise model performance analysis.
In this paper we make several contributions. First, we demonstrate that commonly used meta-metrics depict a statistically biased representation of the true degree of model performance disparities. As a remedy, we introduce a corrected estimator of between-group performance disparities. Second, we illustrate the dangers of a naïve approach to uncertainty quantification via bootstrapping for meta-metrics. In short, we show that bootstrapping induces additional sampling variability that is unaccounted for by a simple statistical bias-corrected estimator. This leads to a situation where even the original corrected estimator (when applied to bootstrap samples) is statistically upward biased and bootstrap intervals have very poor coverage. Third, using the case of binary classification as an example– the most common paradigm in algorithmic fairness– we derive a novel, conceptually simple, and easily implementable double-corrected estimator. This estimator accounts not only for the sampling variance inherent to the original base metric but also the sampling variance induced by the bootstrap procedure itself. We demonstrate that this estimator has coverage near the nominal level in simulated examples. Finally, we show using a real data example that using naïve methods for quantifying model performance disparities can indicate substantial disparities in model performance whereas using our estimator that appropriately accounts for statistical uncertainty in the base metrics leads to the opposite conclusion. Taken in whole, this work illustrates the delicate and difficult nature of rigorously measuring and assessing uncertainty about the between-group model performance disparities that have come to characterize the field of algorithmic fairness.
The paper proceeds as follows. In Section 2, we survey related work on measuring model performance disparities. In Section 3, we introduce notation and in Section 4 we show mathematically and through simulation that all meta-metrics we have identified suffer from statistical bias. In Section 5, we develop a statistically unbiased estimator for summarizing group performance disparities with associated uncertainty quantification, along the way illustrating the pitfalls with naïve application of meta-metrics and bootstrapping. In Section 6, we present our example using the Adult Income dataset. In Section 7, we close with recommendations and pointers to corners of the statistics literature that address very similar problems that will likely prove useful to this task.
Related Work
Many group fairness papers assume that there are only two demographic groups of interest–a privileged group and a non-privileged group–in order to make theory feasible or presentation digestible (Liu et al., 2018; Hu et al., 2019; Bakalar et al., 2021; Zehlike et al., 2017; DiCiccio et al., 2020; Friedler et al., 2019; Yan et al., 2020). Although assessing binary group disparities in this case is straight forward, e.g. using the absolute value of the difference in false positive rates between two groups as a measure of unfairness (Bechavod and Ligett, 2017), this assumption is not appropriate in many real-world settings. While there are some papers that introduce theory or methods that can handle arbitrarily many demographic groups, many of these papers do not synthesize group disparities, and some even turn to binary demographic groups when validating their approaches (Hashimoto et al., 2018; Zhang et al., 2018; Madras et al., 2018; Dwork et al., 2018; Gorantla et al., 2021). Furthermore, empirical work typically shows bar graphs or tables to illustrate model performance disparities over multiple demographic groups. Although this granularity of information is undoubtedly useful and important on its own, these works do not provide a way to quickly synthesize group disparities of model performance across many potential models (Buolamwini and Gebru, 2018; Zehlike et al., 2022; Huang et al., 2020; Adragna et al., 2020).
Outside of the typical group fairness literature, there have been several other threads that address model performance disparities. While individual fairness approaches (Dwork et al., 2012) provide a way to measure disparities in a group-free way–seemingly alleviating the problem we set out to solve–oftentimes to practically validate these individual fairness approaches, group-wise disparities are measured (Yurochkin et al., 2019; Bower et al., 2021) leading us back to our initial problem. Second, statistical dependence between demographic groups and model predictions is a common way to measure group-wise model performance disparities (Komiyama et al., 2018; Mary et al., 2019; Tramer et al., 2017), e.g. correlation. However, the degree of dependence does not tell us the degree of group disparities in model performance, so these approaches in general are inappropriate for our purposes. In addition, we consider categorical demographic groups, not continuous demographic groups, so many of these approaches are not even applicable for the setting we study. Third, economic inequality metrics have been proposed for summarizing model disparities at the individual level without demographic information (Lazovich et al., 2021; Speicher et al., 2018; Saint-Jacques et al., 2020; McCurley, 2008; Bandy and Diakopoulos, 2021). While undoubtedly important, these demographic-free measures of inequality in the allocation of benefits from a model address a different problem. Demographic disparities are in and of themselves an important component of understanding the differential impacts of a model.
There are a small number of papers that propose approaches for practically measuring disparities when there are many demographic groups (Agarwal et al., 2018; Kearns et al., 2018; Ghosh et al., 2021; Geyik et al., 2019; Johndrow and Lum, 2019). All of these papers synthesize group disparities by either reporting the the maximal deviation from average performance (sometimes normalized by subgroup size (Kearns et al., 2018)), the mean absolute deviation from average performance (Johndrow and Lum, 2019), or some measure of the disparity between the maximum and minimum of the base metrics.
Furthermore, there is a growing collection of open source software tools for measuring and mitigating group-wise model performance disparities of machine learning models such as IBM’s AI Fairness 360 (Bellamy et al., 2018), Microsoft’s FairLearn (Bird et al., 2020), Google’s TensorFlow Fairness Indicators (Google, [n.d.]), LinkedIn’s LiFT (Vasudevan and Kenthapadi, 2020), FairTest (Tramer et al., 2017), Aequitas (Saleiro et al., 2018), FAT Forensics (Sokol et al., 2019), FairVis (Cabrera et al., 2019), FairSight (Ahn and Lin, 2019), Themis (Angell et al., 2018), Silva (Yan et al., 2020), and Responsibly (Hod, 18). Most of these tools can only handle binary demographic groups. For those that can handle arbitrarily many groups, they tend to use simple meta-metrics without consideration of statistical bias or cannot aggregate group disparities. IBM’s AI Fairness 360 tool (Bellamy et al., 2018) synthesizes between-group disparities via the generalized entropy index, although the majority of their group fairness metrics can only handle two demographic groups. Although Google TensorFlow Fairness Indicators tool can handle arbitrarily many demographic groups, users are required to determine the disparities through bar charts or tabular data that contains the measurements on each of the groups. Their approach is very similar to the hypothetical example depicted in Figure 1. Microsoft’s FairLearn tool (Bird et al., 2020) can handle arbitrarily many groups where group disparities are measured as a function of the maximum and minimum of model performance over all the groups. LinkedIn’s LiFT tool (Vasudevan and Kenthapadi, 2020) measures variability in group disparities through economic inequality metrics like the generalized entropy index. Neither FairLearn nor LiFT quantifies the uncertainty of the point estimates of these group disparity measures.
Table 1 summarizes the meta-metrics we have identified in both the academic literature and in open-source tools for group-wise model performance disparities measurement and mitigation. We include one additional meta-metric that we have not found used to summarize disparities in the literature or open source tools: the variance. Our reason for including the variance will quickly become apparent.
Notation
As in much of the algorithmic fairness literature, our examples focus on binary classification. In this setting, common choices of the base metric, , are the accuracy, false positive rate, positive predictive value, etc. In all examples presented, we simulate from the following model
This simulation framework is flexible enough to accommodate many of the base metrics that typically would be used for binary classification, though the interpretation of and is metric-dependent. For example, if is accuracy, then represents the number of correct classifications and represents the total number of individuals in the group. If is the false positive rate, then is the number of negative classifications among those observations that were truly positive and is the number of positives in group .
Statistical bias in group-wise model performance disparities metrics
Whereas the metric is applied to the data and predictions within a group to calculate model performance for that group specifically, meta-metrics summarize the degree to which the model’s performance varies across groups. These summary measures, , take in the vector of observed model performance, , and output a number that measures the degree of group-wise model performance disparities, i.e. a distance from the ideal state in which the model performs equally well for all groups. Here we consider statistical bias (the expected difference between the estimator and the truth) in several metrics that have been used in the algorithmic fairness literature, have been implemented in open source software for measuring and remediating disparities, and/or are common statistical measures of variability. Table 1 defines several such metrics.
Suppose we are interested in measuring the “true” degree to which model performance varies across groups, . We don’t observe , but we do observe – a vector of statistically unbiased estimates of . What happens if we use as a measurement of ? For several of the considered metrics, it is easy to see mathematically that is a statistically biased estimate of , though an expression for the expected upward-deviation is not available. For example, for the mean absolute deviation,
where the inequality comes by applying Jensen’s inequality to each term in the sum since the absolute value is a convex function. Similar arguments relying on Jensen’s inequality, can be made to show that is also a positively statistically biased estimator of .
For the variance, we can get an analytical expression for the expectation of in terms of and the s:
We have shown mathematically that several of the meta-metrics are statistically biased. Others are not so mathematically tractable. We investigate all meta-metrics in Table 1 empirically through simulation. In our simulation, five thousand individuals are evenly divided into groups for ranging from five to 150. That is, . In cases where does not evenly divide , the total population size deviates slightly from 5,000 due to rounding. We set to be the corresponding vector of “true” model performance values equally spaced on (i.e. ) and consider several different values of , the lower bound. Given and , we simulate from the model described in (1). Figure 2 shows a Monte Carlo estimate of the statistical bias, , averaged over 1,000 simulations for several values of and , for each of the considered meta-metrics. A different meta-metric is given in each panel of the figure.
For all considered meta-metrics, over-estimates . In practice, this is problematic for two reasons. In an absolute sense, this can paint a misleading picture of the extent to which a model performs differently across groups. For decision-makers who need this information to determine how and whether a model should be used or modified, exaggerating the degree to which the model is “unfair” may cause unnecessary model adjustments. In the context of the well-known fairness-accuracy trade-off (Dutta et al., 2020), having misleading measures of group-wise model performance disparities may cause model builders to trade off more accuracy than necessary in the name of eliminating group-wise model performance disparities, resulting in a less performant model for all individuals. In a relative sense, this can lead to misleading conclusions about how group-wise model performance disparities compares along different axes. For example, suppose we wanted to evaluate the degree of disparities for two “sensitive variables”: age and language, as in the example in Figure 1. These variables have different numbers of groups, with for age and for language. In the case where individuals are evenly distributed across groups as in this simulation, we can see that all of our metrics would, on expectation, return a higher estimate of group-wise model performance disparities for the variable with than the variable with , all else held constant. That is, we would erroneously conclude that the model had more group-wise model performance disparities with respect to language than age, even when the true between-group variability in model performance was identical for both variables. This remains true even when the model is “fair” for both axes (i.e. when the lower bound is 0.9 and all groups have equal true performance).
In this simulated example, the statistical bias increases as a function of for all metrics. This is an artifact of the fact that the observations are equally distributed across the groups. Comparing the statistical bias of a “sensitive variable” with individuals evenly distributed across many groups and another “sensitive variable” with individuals unevenly distributed across fewer groups may not exhibit the same pattern of a larger statistical bias for the many-group variable than the few-group variable.
Correcting for statistical bias in group-wise model performance disparities measurement
Given the statistical bias in all of the existing meta-metrics, what is to be done? The expression for the statistical bias in the variance offers one possibility. Because we know the statistical bias in as a function of the s, we can correct for it and use the estimator
where is the standard error of the estimate of model performance for group . As it turns out, (2) was first introduced by Cochran in the random effects ANOVA context in 1954 (Cochran, 1954) and was popularized for meta-analysis of between-study variability by Hedges and Olkin in 1985 (Hedges and Olkin, 1985). Often, in these other contexts, this estimator is truncated to preclude the possibility of a negative variance estimate, i.e. . We also adopt this convention. The estimator given in Equation (2) offers a conceptually clean and easily calculable estimate of between-group model performance variance. The untruncated version is statistically unbiased when statistically unbiased estimates of the standard errors are also available (Viechtbauer, 2005). We find this estimator appealing because it is easy to explain and easy to calculate without access to statistics-specific software packages. Going forward, due to the mathematical tractability of the variance as a meta-metric, we focus our investigation on measuring and quantifying uncertainty about the variance of model performance across groups.
We extend the simulation framework to focus in on four scenarios. These scenarios are summarized in Table 2. In all scenarios presented, we consider a training set of size of with . In the Equal Group-size condition, all groups have observations. In the Unequal Group-size condition, we create group sizes by linearly interpolating between 10 and 90 and rounding. The minimum group size is 10 and the maximum 90. Every integer between 10 and 90 has at least one group of that size; occasionally due to rounding there are two groups with the same number of observations. In the Equal Performance condition, . In these scenarios, the true variance across groups is zero. In the Unequal Performance condition, group performance is equally spaced between 0.1 and 0.9, as in the simulation in Figure 2. Here, the true variance in performance across groups is 0.055. For each scenario, we again use the simulation model given in (1).
For each scenario, we simulate data and calculate the variance of and the corrected variance of as in (2). To estimate , we use the plug-in estimator of the sampling variance of a binomial proportion, . We repeat this 1000 times. Figure 3 shows the distribution of these estimates. While we have already seen (both mathematically and through simulation) that the uncorrected variance (red) is upwardly statistically biased, this shows that correcting for the statistical bias by substracting off an estimate of the average standard error of the groups successfully results in between-group variance estimates that are centered at the truth (vertical black line), i.e. statistically unbiased. Next we explore what happens if we bootstrap this corrected variance estimator to obtain bootstrap intervals for uncertainty quantification.
2. Bootstrapping the (statistical) bias-corrected variance estimate
The corrected variance estimator defines an easy and explainable way to obtain a statistically unbiased point estimate of the true between-group performance variance. Bootstrapping offers a similarly conceptually simple and easily implementable method for calculating uncertainty intervals. To create non-parametric bootstrap intervals, one samples with replacement from the observed data. If we consider the number of individuals in each group to be fixed (as we do in the following examples), then for each group , we sample individuals with replacement from the collection of all individuals in group . This makes one bootstrap sample, . For each bootstrap sample, a statistic– such as the corrected variance– is calculated. This process is repeated times, resulting in samples of the target statistic. Uncertainty intervals for the statistic are calculated as the empirical quantiles of the estimates. For example, a standard 95% interval would be calculated by setting the lower and upper bounds of the interval to be the 0.025 and 0.975 quantiles of the samples of the statistic, respectively. There are many techniques for bootstrapping, such as versions that use the bootstrap samples only to calculate standard errors around point estimate obtained from the original sample (Efron and Tibshirani, 1994). Here, we use the “percentile” bootstrap. Future work should consider other techniques for bootstrapping in this context, such as using single-corrected estimator as a point estimate and the double-corrected estimator to obtain intervals about that point estimate.
Figure 4 shows the distribution of bootstrap samples of the corrected and uncorrected variance for one draw from the generative model. The horizontal bars show 95% bootstrap intervals for the corrected variance. It is important to differentiate that whereas in Figure 3, the histogram is made up of single estimates of the variance each corresponding to independent draws from the generative model, in Figure 4 the histogram is generated by bootstrap re-sampling from one draw from the generative model to obtain uncertainty estimates for the variance estimate associated with that single draw. Here, we see that the naïve application of the (corrected) variance estimator to each bootstrap sample results in distributions of estimates that are systematically shifted upwards with intervals (horizontal lines) that do not even come close to covering the true value (shown by the vertical black line). When bootstrapping, even our statistically unbiased estimator of the variance is statistically biased.
3. What happened?
Let’s go back to our original model given in (1). Under this model, the variance of our estimator is given by , which is what we’ve used as our estimate of the sampling variance of when applying the correction. However, when we generate bootstrap samples (by re-sampling the observations within groups), our generative model for each bootstrap sample becomes the following:
Applying iterated expectations, we find that the sampling variance of each bootstrap sample is given by
Again, using a plug-in estimate of the sampling variance of , we get . This implies that for each bootstrap sample, a “double-corrected” estimate of the variance of is given by
In summary, to obtain bootstrap intervals for the variance of , we must account not only for the variance in the generative model, but also the additional variance that occurs due to the bootstrapping procedure itself. Figure 5 shows bootstrap intervals and a histogram of bootstrap samples for the uncorrected, corrected, and double-corrected estimates of the between-group variance for one draw from the generative. Here we see that– at least for this replicate– by explicitly accounting for the additional variance induced by the bootstrap sampling process itself, we have once again successfully created an estimator that accurately captures between-group variance.
Extending beyond a single replicate, Table 3 shows the empirical coverage of 95% bootstrap intervals across 1000 replicates. We see that applying the double-corrected variance estimator to create bootstrap intervals of the between-group variance has coverage similar to the nominal 95% level.
Real data example
We compare the variance and double-corrected variance estimators of group-wise model performance disparities in a model built from the Adult Income datasetDownloaded from the UCI Machine Learning Repository at https://archive.ics.uci.edu/ml/datasets/adult., which is frequently used in algorithmic fairness research. It contains records of demographic and income data for 48,842 individuals from the 1994 census database. There are a total of 14 features, a weight column (fnlwgt), and a label column that indicates whether each person makes over $50K a year. We split the data randomly into train and test sets by 70%-30%, and trained histogram-based gradient boosting classification treesWe used the scikit-learn implementation https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.HistGradientBoostingClassifier.html. using all 14 features of the dataset. The weight column is dropped as per prior research (Zhang et al., 2018; Yurochkin et al., 2019). The resulting classifier has a 87.3% accuracy on the test set.
We are interested in how the model performs with respect to two demographic variables race and age. We divide age into 8 groups from 15 to 95, each spanning 10 years. The distribution of these two features are highly skewed. In the test set, across race, white accounts for 86% of the data, while American Indian and Eskimo account for only 0.9%. Across the 8 age groups, the smallest group, (85, 95], accounts for just 0.1%, while the largest group, (25, 35], accounts for 27%. We expect the presence of these small groups leads to increased statistical bias in estimated variance.
Figure 6 shows 95% bootstrap intervals for the uncorrected and double corrected variance estimates. We calculated the intervals over three base model performance metrics, selection rate (SR), false positive rate (FPR), and true positive rate (TPR), since they are the base metrics that popular fairness criteria such as demographic parity and equalized odds depend on. The figure shows that the discrepancy between the uncorrected and double-corrected intervals is particularly large for TPR: the double-corrected intervals cover zero, while the uncorrected intervals do not. In other words, had we just used the uncorrected intervals, we would have concluded with high statistical certainty that there was a large disparity in model’s TPR over both racial and age groups. However, after accounting for uncertainty about TPR within each group, we conclude that the disparities are much lower, and we cannot statistically rule out zero disparities in false positives as a possibility.
Conclusion
In this paper, we have demonstrated both theoretically and empirically that meta-metrics commonly used in the algorithmic fairness literature and in open-source tools designed for measuring group-wise model performance disparities are themselves statistically biased measurements. This occurs because existing meta-metrics fail to account for statistical uncertainty in the underlying base metrics, thus exaggerating between-group model performance disparities. In practice, this can cause misleading inferences about the extent of model performance disparities and erroneous conclusions about the relative degree of disparities between different sensitive variables. After identifying this problem, we propose a new estimator for evaluating between-group model performance disparities. This estimator builds upon the work of Cochran (Cochran, 1954) and Hedges and Olkin (Hedges and Olkin, 1985) by accounting for variability induced by bootstrap procedures. Our resulting double-corrected estimator offers a simple, conceptually clean, and easily implementable way to measure model performance disparities and associated uncertainty.
The method we have developed offers one approach to quantifying group-wise model performance disparities in the presence of many groups. To our knowledge, this is the first time this question has been studied in the context of algorithmic fairness. However, between-group variance estimation is well-studied in the statistics literature more generally. Common approaches include restricted maximum likelihood estimation of the variance parameter in random effects models (Harville, 1977). The literature on meta-analysis also contains many analogous methods. In that domain, the goal is to estimate variability in effect size across studies. There, represents the effect size of a particular study, and the standard error of that estimate. (See (Veroniki et al., 2016) for an excellent overview of methods available to estimate between-study effect size variance and its associated uncertainty in the context of meta-analysis and (Langan et al., 2019) for a recent simulation study comparing methods.) Many of these approaches also offer solutions for quantifying model performance disparities across many groups. Future work should perform simulation studies comparing different methods for measuring between-group variability with parameters tailored to the specifics of typical problems encountered when evaluating model “fairness.”
Finally, model performance disparities as measured by meta-metrics are not dispositive of the presence or absence of algorithmic harms. While large disparities typically indicate that an issue needs further investigation, small measured disparities do not guarantee that the system is fair or free from adverse impacts. Just as it would be foolish to claim a computer system is completely secure or a data set is completely private, it is always possible that there are undiscovered vulnerabilities that our measurements have not uncovered. While we have presented an approach to more accurately measure model performance disparities, these measurements cannot tell us whether the observed disparities are acceptable or whether we have calculated disparities with respect to appropriate grouping variables. These judgments are subjective and require understanding of the context in which a model will be used. Like all metrics, meta-metrics simply cannot capture the entirety of the impact of machine learning systems on people.
References
Appendix A Derivations
Derivation of the single-corrected variance