Fair Preprocessing: Towards Understanding Compositional Fairness of Data Transformers in Machine Learning Pipeline
Sumon Biswas, Hridesh Rajan
Introduction
Fairness of machine learning (ML) predictions is becoming more important with the rapid increase of ML software usage in important decision making (Dixon et al., 2018; Olson, 2011; Angwin et al., 2016; Goodall, 2016), and the black-box nature of ML algorithms (Galhotra et al., 2017; Aggarwal et al., 2019). There is a rich body of work on measuring fairness of ML models (Zafar et al., 2015; Dwork et al., 2012; Feldman et al., 2015; Hardt et al., 2016; Calders and Verwer, 2010; Chouldechova, 2017; Zemel et al., 2013; Speicher et al., 2018) and mitigate the bias (Zhang et al., 2018; Calders and Verwer, 2010; Zafar et al., 2015; Chouldechova, 2017; Goh et al., 2016; Hardt et al., 2016; Pleiss et al., 2017; Kamiran et al., 2012). Recent work (Brun and Meliou, 2018; Holstein et al., 2019; Friedler et al., 2019; Biswas and Rajan, 2020; Harrison et al., 2020; Chakraborty et al., 2019) has shown that more software engineering effort is required towards detecting bias in complex environment and support developers in building fairer models.
The majority of work on ML fairness has focused on classification task with single classifier (Galhotra et al., 2017; Feldman et al., 2015; Aggarwal et al., 2019; Bower et al., 2017). However, real-world machine learning software operate in a complex environment (Bower et al., 2017; D’Amour et al., 2020). In an ML task, the prediction is made after going through a series of stages such as data cleaning, feature engineering, etc., which build the machine learning pipeline (Amershi et al., 2019; Yang et al., 2020). Studying only the fairness of the classifiers (e.g., Decision Tree, Logistic Regression) fails to capture the fairness impact made by other stages in ML pipeline. In this paper, we conducted a detailed analysis on how the data preprocessing stages affect fairness in ML pipelines.
Prior research observed that bias can be encoded in the data itself and missing the opportunity to detect bias in earlier stage of ML pipeline can make it difficult to achieve fairness algorithmically (Kirkpatrick, 2017; Holstein et al., 2019; Grgic-Hlaca et al., 2018; Dixon et al., 2018). Additionally, bias mitigation algorithms operating in the preprocessing stage were shown to be successful (Kamiran and Calders, 2012; Feldman et al., 2015). Therefore, it is evident that the preprocessing stages of ML pipeline can introduce bias. However, no study has been conducted to measure the fairness of the preprocessing stages and show how it impacts the overall fairness of the pipeline. In this paper, we used the causal method of fairness to reason about the fairness impact of preprocessing stages in ML pipeline. Then, we leveraged existing fairness metrics to measure fairness of the preprocessing stages. Using the measures, we conducted a thorough analysis on a benchmark of 37 real-world ML pipelines collected from three different sources, which operate on five datasets. These ML pipelines allowed us to evaluate fairness of a wide selection of preprocessing stages from different categories such as data standardization, feature selection, encoding, over/under-sampling, imputation, etc. For comparative analysis, we also collected data transformers e.g., StandardScaler, MinMaxScaler, PCA, l1-normalizer, QuantileTransformer, etc., from the pipelines as well as corresponding ML libraries, and evaluated fairness. Finally, we investigated how fairness of these preprocessing techniques (local fairness) composes with other preprocessing stages, and the whole pipeline (global fairness). Specifically, we answered the following three research questions.
RQ1 (fairness of preprocessing stages): What are the fairness measures of each preprocessing stage in ML pipeline? RQ2 (fair transformers): What are the fair (and biased) data transformers among the commonly used ones? RQ3 (fairness composition): How fairness of data preprocessing stages composes in ML pipeline?
How local fairness compose into global fairness?
Does choosing a downstream transformer depend on the fairness of an upstream transformer?
To the best of our knowledge, we are the first to evaluate the fairness of preprocessing stages in ML pipeline. Our results show that by measuring the fairness impact of the stages, the developers would be able to build fairer predictions effectively. Furthermore, the libraries can provide fairness monitoring into the data transformers, similar to the performance monitoring for the classifiers. Our evaluation on real-world ML pipelines also suggests opportunities to build automated tool to detect unfairness in the preprocessing stages, and instrument those stages to mitigate bias. We have made the following contributions in this paper:
We created a fairness benchmark of ML pipelines with several stages. The benchmark, code and results are shared in our replication packagehttps://github.com/sumonbis/FairPreprocessing in GitHub repository, that can be leveraged in further research on building fair ML pipeline.
We introduced the notion of causality in ML pipeline and leveraged existing metrics to measure the fairness of preprocessing stages in ML pipeline.
Unfairness patterns have been identified for a number of stages.
We identified alternative data transformers which can mitigate bias in the pipeline.
Finally, we showed the composition of stage-specific fairness into overall fairness, which is used to choose appropriate downstream transformer that mitigates bias.
The paper is organized as follows: §2 describes the motivating examples, §3 describes the existing metrics and our approach. In §4, we described the benchmark and experiments. §5 explores the results, §6 provides a comparative study among transformers, and §7 evaluates the fairness composition. Finally, §9 describes the threats to validity, §10 discusses related work, and §11 concludes.
Motivation
In this section, we present two ML pipelines which show that the preprocessing stage affects the fairness of the model and it is important to study the bias induced by certain data transformers.
Yang et al. (Yang et al., 2020) studied the following ML pipeline which was originally outlined by Propublica for recidivism prediction on Compas dataset (Angwin et al., 2016). The goal is predict future crimes based on the data of defendants. The fairness values, in terms of statistical parity difference (SPD: -0.102) and equal opportunity difference (EOD: -0.027), suggest that the prediction is biased towardsBias towards a group connotes that the prediction favours that group. Caucasian defendants when race is considered as sensitive attribute. The pipeline consists of several preprocessing stages before applying LogisticRegression classifier. Data preprocessing includes cleaning, encoding categorical features, and missing value imputation. Recent research (Yang et al., 2020) showed that the data transformation in this pipeline is not symmetric across gender groups i.e., male defendants are filtered more than the female. Do these data transformations introduce unfairness in the prediction? If yes, what are the unfairness measures of these transformers? Is it possible to leverage existing metrics to measure the unfairness of each component? If we can understand the effect of each data transformer, it would be possible to choose data preprocessing technique wisely to avoid introducing bias as well as mitigate the inherent bias in data or classifier.
2. Motivating Example 2
The following ML pipeline is collected from the benchmark used by Biswas and Rajan (Biswas and Rajan, 2020) for studying fairness of ML models. This pipeline operates on German Credit dataset. Here, the goal is to predict the credit risk (good/bad) of individuals based on their personal data such as age, sex, income, etc. In this pipeline, before training the classifier, data has been processed using two transformers: PCA for principal component analysis, and SelectKBest for selecting high-scoring features. The fairness value (SPD: 0.005) shows that prediction is slightly biased towards female candidates. However, if the transformers are not applied, then prediction becomes biased towards male (SPD: -0.117). By applying one transformer at a time, we observed that PCA alone is not causing the change of fairness. In this case, SelectBest is causing bias towards female, which in turn mitigating the overall fairness of the pipeline. Therefore, in addition to study the fairness of transformers in isolation, it is important to understand how fairness of components composes in the pipeline.
Our key idea is to leverage causal reasoning and observe fairness impact of a stage on prediction. To do that we create alternative pipeline by removing a stage. For example, from the above pipeline, we remove the SelectKBest and compare the predictions with original pipeline. We observe that SelectKBest is causing 1.1% of the female and 3.6% of the male participants to change predictions from favorable (good credit) to unfavorable (bad credit). Since the stage is causing more unfavorable decisions to male, the stage is biased towards female. Thus, we used existing fairness criteria to measure fairness impact of a stage and propose novel metrics.
Methodology
In this section, first, we describe the background of ML pipeline, focussing on the data preprocessing stages. Second, we formulate the method and metrics to measure fairness of a certain preprocessing stage with respect to the pipeline it is used within.
Amershi et al. proposed a nine-stage machine learning pipeline with data-oriented (collection, cleaning, and labeling) and model-oriented (model requirements, feature engineering, training, evaluation, deployment, and monitoring) stages (Amershi et al., 2019). Other research (Abadi et al., 2016; Baylor et al., 2017) also described data preprocessing as an integral part of the ML pipeline. The pipelines in the motivating examples are depicted in Figure 1, which follows the representation provided by Yang et al. (Yang et al., 2020). In this paper, we adapted the canonical definition of pipeline from Scikit-Learn pipeline specification (Buitinck et al., 2013; Scikit-Learn Pipeline, 2020), which is aligned with the ML models studied in the literature for fair classification tasks (Bellamy et al., 2018; Yang et al., 2020; Biswas and Rajan, 2020; Friedler et al., 2019; Aggarwal et al., 2019; Galhotra et al., 2017). We are interested in investigating the fairness of the data preprocessing stages in the pipeline, which is depicted with grey boxes in Figure 1.
To summarize, a canonical ML pipeline is an ordered set of stages with a set of preprocessing stages () and a final classifier (). Each preprocessing stage, operates on the data already processed by preceding stages . A data preprocessing stage can be a data transformer or a set of custom operations. A data transformer is a well-known algorithm or method to perform a specific operation such as variable encoding, feature selection, feature extraction, dimensionality reduction, etc. on the data (Buitinck et al., 2013). For example, in the second motivating example, two transformers (PCA and SelectKBest) have been used. Custom transformation includes data/task-specific contextual operations on the dataset. For example, in §2.1 (line 2-3), the data instances that do not contain a value in the range for the feature days_b_screening_arrest, have been filtered. This means the pipeline ignored the data of the defendants with more than 30 days between their screening and arrest. This formulation of ML pipeline allowed us to evaluate fairness of the preprocessing stages in real-world ML tasks.
2. Existing Fairness Metrics
We have leveraged existing fairness metrics to measure the fairness of the whole pipeline. Many fairness metrics have been proposed in the literature for measuring fairness of classification tasks (Bellamy et al., 2018; Binns, 2017; Dixon et al., 2018). In general, the fairness metrics compute group-specific classification rates (e.g., true positives, false positives), and calculates the difference between groups to measure the fairness. In this paper, we adopted the representative group fairness metrics used by (Friedler et al., 2019; Biswas and Rajan, 2020). Specifically, we leveraged the following metrics: statistical parity difference (SPD) (Kamiran and Calders, 2012; Zafar et al., 2015; Feldman et al., 2015), equal opportunity difference (EOD) (Hardt et al., 2016), average odds difference (AOD) (Hardt et al., 2016), and error rate difference (ERD) (Chouldechova, 2017). Given a dataset with instances, let, actual classification label be , predicted classification label be , and sensitive attribute be . Here, if the label is favorable to the individuals, otherwise . For example, classification task on German Credit dataset predicts the credit risk (good/bad credit) of individuals. In this case, if the prediction is good credit, otherwise . Suppose, for privileged group (e.g., White), and for unprivileged group (e.g., non-White), . SPD is computed by observing the probability of giving favorable label to each group and taking the difference. EOD measures the true-positive rate difference between groups. AOD calculates both true positive rate and false positive rate difference and then takes the average. ERD calculates the sum of false positive rate difference and false negative rate difference between groups. The definitions of these metrics are as follows:
Disparate impact (DI) and statistical parity difference (SPD) both measure the same rate i.e., probability of classifying data instance as favorable, but DI computes the ratio of privileged and unprivileged groups’ rate, whereas SPD computes the difference. Therefore, from DI and SPD, we only used SPD in our evaluation.
3. Fairness of Preprocessing Stages
Suppose, is a pipeline with stages and our goal is to evaluate the fairness of the stage , where . In other words, we want to measure the fairness impact of on the prediction made by . To achieve that we applied the causal reasoning for evaluating fairness. The causality theorem was proposed by Pearl (Pearl, 2000, 2009) and further studied extensively to reason about fairness in many scenarios (Galhotra et al., 2017; Kusner et al., 2017; Zhang and Bareinboim, 2018; Salimi et al., 2019; Russell et al., 2017). Causality notion of fairness captures that everything else being equal, the prediction would not be changed in the counterfactual world where only an intervention happens on a variable (Galhotra et al., 2017; Kusner et al., 2017; Russell et al., 2017). For example, Galhotra et al. proposed causal discrimination score for fairness testing (Galhotra et al., 2017). The authors created test inputs by altering original protected attribute values of each data instance, and observed whether prediction is changed for those test inputs. If the intervention causes the prediction to be changed, we call the software causally unfair with respect to that intervention. In our case, if a preprocessing stage be the intervention, to measure the fairness of , we have to capture the prediction disparity caused by the intervention . This causal reasoning of fairness is a stronger notion since it provides causality in software by observing changes in the outcome made by a specific stage in the pipeline (Galhotra et al., 2017; Pearl, 2009).
From pipeline , we construct another pipeline by only excluding the stage from . After applying the stage in , to what extent the prediction of changes, and whether the change is favorable to any group? Broadly, this can be measured by observing the prediction difference between and and computing the fairness of these changes using the fairness metrics from (1).
Suppose, the predictions made by the two pipelines are and . Let, be the impact set for , which denotes the prediction parity between and such that for data instance, if , then , otherwise . By causality, the fairness of preprocessing stage (denoted by ) is calculated based on , with respect to a fairness metric , which is shown in (2a). We noticed that a few preprocessing stages, specifically the encoders can not be removed without replacing with an alternative stage. For such situations, we have defined the fairness of with reference to another stage , denoted by in (2b).
Zelaya also used the similar method for quantifying the effect of a preprocessing stage with a goal of computing volatility of a stage (Zelaya, 2019). Volatility quantifies how much impact a preprocessing stage has on the outcome by computing the probability of prediction changes. However, it does not capture the fairness of the stage, since a stage can cause high change in the prediction by maintaining the predictions fair. Next in §3.3.2, we have extended our causality based formulation of (2a) for each fairness metric in (1) to capture the fairness impact of each preprocessing stage. Similar to (Galhotra et al., 2017), the benefit of this formulation is, the measures do not require an oracle, since the prediction equivalence of pipelines and serves the goal of evaluating fairness of the stage. Note that the rest of the definitions in §3.3.2 are independent of (2a) and (2b).
3.2. Fairness Metircs for Preprocessing Stage
We have leveraged the definition of metrics SPD, EOD, AOD, and ERD from (1) to capture the stage-specific fairness of . Essentially, the new metrics will identify the disparities between and and use corresponding fairness criteria to measure how much favors a specific group with respect to other group(s).
Suppose, among data instances, are from the unprivileged group and from the privileged group. computes how many of the data instances have been changed from unfavorable to favorable after applying the stage . To do that we count changes in both directions (unfavorable to favorable and favorable to unfavorable), and take the difference. The sign of preserves the direction of changes. Finally, the metric is computed by taking the difference of rates () between unprivileged and privileged groups. Note that the metric captures fairness by measuring the difference of favorable change rates between groups. Simply counting the mismatches between and could provide degree of changes in but would not capture fairness. Furthermore, computing favorable changes to both groups separately and evaluating the disparity between them captures fairness according to the original definition of .
Similarly, is defined using the following equation. In this case, only the true-positive changes are considered as suggested by the definition of EOD from (1).
Since AOD computes the average of true positive (TP) rate and false positive (FP) rate, first the change set for TP and FP predictions is computed. Then averaging the probability of changes for TP and FP, the change rates are computed for both groups. Finally, is calculated by taking the difference of rates between privileged and unprivileged groups.
Finally, is computed using the change of count in both false positives (FP) and false negatives (FN) as mentioned in the definition of ERD in (1).
Thus far, we have four fairness metrics (, , , and ) to measure the fairness of the stage. In general, the rates computed by each metric () follow the same range of the original metrics . Therefore, the above metrics have a range . Positive values indicate bias towards unprivileged group, negative values indicate bias towards privileged group, and values very close to 0 indicate fair preprocessing stage.
Evaluation
In this section, we describe the benchmark dataset and pipelines that we used for evaluation. Then we present the experiment design and results for answering the research questions.
We collected ML pipelines used in prior studies for fairness evaluation. First, Biswas and Rajan collected a benchmark of 40 ML models collected from Kaggle that operate on 5 different datasets e.g., German Credit (Hofmann, 1994), Adult Census (Kohavi, 1996), Bank Marketing (Moro et al., 2014), Home Credit (Kaggle, 2017a) and Titanic (Kaggle, 2017b). However, the authors did not study the fairness at the component level, rather the ultimate fairness of the classifiers e.g., RandomForest, DecisionTree, etc. We revisited these Kaggle kernels and collected the preprocessing stages used in the pipelines. We noticed that Home Credit dataset (Kaggle, 2017a) in this benchmark is not unified like the other datasets, distributed over multiples CSV files, and the models under this dataset do not operate on the same data files. Hence, these models (8 out of 40) are not suitable for comparing fairness of data preprocessing stages.
Second, we collected the pipelines provided by Yang et al. (Yang et al., 2020). The authors released 3 pipelines on two different datasets - Adult Census and Compas. Third, Zelaya (Zelaya, 2019) studied the volatility of the preprocessing stages using two pipelines on a fairness dataset i.e., German Credit. We included these pipelines in our benchmark. Thus, we created a benchmark of 37 ML pipelines that operate on five datasets. The pipelines with the stages in each dataset category and their performances are shown in Table 1. Below we present a brief description of the datasets and associated tasks.
German Credit. The dataset contains 1000 data instances and 20 features of individuals who take credit from a bank (Hofmann, 1994). The target is to classify whether the person has a good/bad credit risk.
Adult Census. The dataset is extracted by Becker (Kohavi, 1996) from 1994 census of United States. It contains 32,561 data instances and 12 features including demographic data of individuals. The task is to predict whether the person earns over 50K in a year.
Bank Marketing. This dataset contains a bank’s marketing campaign data of 41,188 individuals with 20 features (Moro et al., 2014). The goal is to classify whether a client will subscribe to a term deposit.
Titanic. The dataset contains information about 891 passengers of Titanic (Kaggle, 2017b). The task is to predict the survival of the individuals on Titanic. The sensitive attribute of this dataset is sex.
Compas. The dataset contains data of 6,889 criminal defendants in Florida. Propublica used this dataset and showed that the recidivism prediction software used in US courts discriminates between White and non-White (Angwin et al., 2016). The task is to classify whether the defendants will re-offend where race is considered as the sensitive attribute.
2. Experiment Design
Each pipeline in the benchmark consists of one or more preprocessing stages followed by the classifier. In this paper, our main goal is to evaluate the fairness of different preprocessing stages using the fairness metrics described in §3. The benchmark, code and results are released in the replication package (Biswas and Rajan, 2021).
The experiment design for evaluating the pipelines is shown in Figure 2. First, for each pipeline, we identified the preprocessing stages. For example, the pipeline in §2.1 contains six preprocessing stages. To evaluate the fairness of a stage in a pipeline , we create an alternative pipeline by removing the stage . For stages that can not be removed, we replaced with a reference stage . Among the preprocessing stages shown in Table 1, we found only the encoders can not be removed. We experimented with all the encoders in Scikit-Learn library (Scikit Learn, 2019b), i.e., OneHotEncoder, LabelEncoder, OrdinalEncoder, and found that OneHotEncoder does not exhibit any bias. Therefore, we used OneHotEncoder as the reference stage for the encoders in our experiment.
Second, the original dataset is split into training (70%) and test set (30%). Then two copies of training data are used to train pipeline and . After training the classifiers, two models predict the label for the same set of test data instances. Then, similar to the experimentation of (Zelaya, 2019), for each prediction label we compare the two predictions and with the true prediction label . This comparison provides the necessary data to compute the four fairness metrics. Similar to (Friedler et al., 2019; Biswas and Rajan, 2020), for each stage in a pipeline, we run this experiment ten times, and then report the mean and standard deviation of the metrics, to avoid inconsistency of the randomness in the ML classifiers. Finally, we followed the ML best practices so that noise is not introduced evaluating the fairness of preprocessing stages. For example, while applying some transformation, lack of data isolation might introduce noise in the evaluation, e.g., when applying PCA on dataset, it is important to train the PCA only using the training data. If we use the whole dataset to train the PCA and transform data, then information from the test-set might leak. Third, since a stage operates on the data processed by the preceding stage(s), there are interdependencies between them. We always maintained the order of the stages while removing or replacing a stage (§5). To observe fairness of data transformers without interdependencies, we applied them on vanilla pipelines (§6).
Fairness of Preprocessing Stages
In this paper, we used a diverse set of metrics, to evaluate fairness of preprocessing stages. While developing an ML pipelines, if the developer has a comprehensive idea of the fairness of preprocessing stages, it would be convenient to build a fair pipeline. The evaluation has been done on 69 preprocessing stages in 37 ML pipelines from 5 dataset categories. For Compas dataset, we found one pipeline (§2.1). Five out of six stages in this pipeline exhibit no bias, which has been discussed later. For other 4 dataset categories, the evaluation has been shown in Figure 3. In this section, first, we discuss how we can interpret the metrics. Second, we answer the first research question and discuss the findings from our evaluation.
We investigated the fairness of the preprocessing stages using four metrics: , , , and . These metrics measure the fairness of the stages by using the existing fairness criteria, e.g., measures the fairness of a stage with respect to statistical parity difference () criteria. These fairness criteria evaluate algorithmic fairness of ML pipelines (Feldman et al., 2015; Zafar et al., 2015; Calders and Verwer, 2010). The unfairness characterized by these criteria is measured based on the prediction disparities, although the root cause can be the training data or the algorithm (e.g., data preprocessing, classifier) itself. Therefore, when an ML model is identified as unfair, it implies that in the given predictive scenario, the outcome is biased. Similarly, the metrics proposed in this paper measure algorithmic unfairness caused by a specific preprocessing stage with respect to its pipeline. For instance, in Figure 3, pipeline GC4 has two stages: LabelEncoder and StandardScaler. The fairness metrics suggest that LabelEncoder is biased towards unprivileged group (positive value), and StandardScaler is biased towards privileged group (negative value). The stages for which the measures are very close to zero, can be considered as fair preprocessing.
The metrics can provide different fairness signals for a certain stage. For example, in AC4, shows positive fairness, whereas the other metrics suggest negative fairness for both the stages PCA and StandardScaler. This disparity occurs because different metrics accounts for different fairness criteria. In this case, , depends on the false positive and false negative rate difference. No other metric is concerned about the false negative rate difference, and hence provides a different fairness signal than other metrics. In practice, appropriate fairness criteria can vary depending on the task, usage scenario, or involved stakeholders. Study suggests that developers need to be aware of different fairness indicators to build fairer pipelines (Friedler et al., 2019). Therefore, we defined and evaluated fairness of stages with respect to multiple metrics.
2. Fairness Analysis of Stages
The pipelines used both built-in algorithm imported from libraries i.e., data transformers (Buitinck et al., 2013), as well as custom preprocessing stages. The stages found in each pipeline are shown in Table 1, and the fairness measures of those stages are plotted in Figure 3. Although the unfairness exhibited by a stage is with respect to the pipeline, we found fairness patterns of some stages and investigated them further. In general, our findings show that the stages which change the underlying data distribution significantly, or modify minority data are responsible for increasing bias in the pipelines.
Most of the real-world datasets contain missing values (MV) for several reasons such as data creation errors, not-applicable (N/A) attributes, incomplete data collection, etc. In our benchmark, Adult Census and Titanic contain MV that required further processing in the pipeline. 7.4% rows in Adult Census and 20.2% rows in Titanic have at least one missing feature in the dataset. The pipelines either remove the rows with MV or apply certain imputation (Scikit Learn, 2019c) technique that replaces the MV with mean, median or most frequently occurred values. Removal of rows with MV can significantly change the data distribution, which introduces bias in the pipeline. For example, both TT1 and TT2 removed data items with MV by applying df.dropna() method, which introduces bias in the prediction (Figure 3). Research has shown that MV are not uniformly distributed over all groups and data items from minority groups often contain more MV (Dixon et al., 2018). If those data items are entirely removed, the representation of minority groups in the dataset becomes scarce. On the other hand, TT3 applied mean-imputation and TT4 applied both median- and mode-imputation using df.fillna(), which exhibits fairness compared to data removal. While our findings suggest that removing data items with MV introduces bias, the most popular fairness tools AIF 360 (Bellamy et al., 2018), Aequitas (Saleiro et al., 2018), Themis-ML (Bantilan, 2018) ignore these data items and remove entire row/column. Our evaluation strategy confirms that the tools can integrate existing imputation methods (Scikit Learn, 2019c) in the pipeline and allow users to choose appropriate ones. Additionally, more research is needed to understand and develop imputation techniques that are fairness aware.
We found that most of the feature engineering stages, especially the custom transformations exhibit bias in the pipeline. For example, the pipelines in Titanic dataset used custom feature engineering, since the dataset contains composite features which may provide additional information about the individuals. For instance, TT8 operates on the feature name to create a new feature title e.g., Mr, Mrs, Dr, etc. This transformation aims at better prediction of the survival of passengers by extracting the social status, but creates high bias between male and female.
In Adult Census, the feature eduction of individuals contain values such as preschool, 10th, 1st-4th, prof-school, etc., which have been replaced by broad categories such Dropout, HighGrad, Masters in the pipeline AC7. In addition, instead of using age as continuous value, the feature has been discretized into number of bins. In both cases, the original data values have been modified, which has caused unfairness in the pipeline. Nevertheless, some pipelines (AC3, AC8) have custom feature transformations that are fair. Previous studies showed that certain features contribute more to the predictive quality of the model (Garreau and Luxburg, 2020; Ribeiro et al., 2018). Feature importance in prediction and corelation of features with the sensitive attribute also led to bias detection (Chakraborty et al., 2020; Grgic-Hlaca et al., 2018) in ML models. However, does creating new features (by removing certain semantics) from a potentially biased feature increase the fairness, is an open question. Our method to quantify the fairness of such changes can guide further research in this direction.
Two most used encoding techniques for converting categorical feature to numerical feature are OneHotEncoder and LabelEncoder. OneHotEncoder creates new columns by replacing one column for each of the categories. LabelEncoder does not increase the number of the columns, and gives each category an integer label between and (). In our evaluation, we found that LabelEncoder introduces bias in German Credit and Titanic dataset but OneHotEncoder does not change fairness. Since LabelEncoder imposes a sequential order between the categories, it might create a linear relation with the target value, and hence have an impact on the classifier to change fairness. For example, pipelines TT7 creates a new feature called Family based on the surname of the person. This feature has a large number of unique categories (667 unique ones in 891 data instances). Therefore, the non-sparse representation in LabelEncoder adds additional weight to the feature, which is causing unfairness in TT7. Developers might avoid OneHotEncoder because it suffers from the curse of dimensionality and the ordinal relation of data is lost. In that case, developers should be aware of the fairness impact of the encoder. One solution might be using PCA for dimensionality reduction, which has been done in GC7.
We have plotted the standard error of the metrics as error bars in Figure 3. Firstly, it shows that the metrics in German Credit and Titanic dataset are more unstable. The reason is that the size of these two datasets is less than the other three datasets. German Credit dataset has 1000 instances, and Titanic has 891 instances. Adult Census and Bank Marketing dataset have more than 30K instances. If the sample size is large, data distribution tends to be similar even after taking a random train-test split (Friedler et al., 2019). However, when the dataset size is smaller, the distribution is changed among different train-test splitting. Furthermore, we have found that is more unstable than other metrics. depends on the change of false positive and false negative rates. However, in most cases, the pipelines are optimized for accuracy and precision, since these are some best performing ones collected from Kaggle. Therefore, before deploying preprocessing stages, it would be desirable to test the stability of over multiples executions.
For the Compas dataset, we evaluated the six stages shown in §2.1. All the stages exhibited data filtering show bias. The data filtering also showed bias close to zero (less than .005) with respect to all the metrics. Although Yang et al. (Yang et al., 2020) argued that this pipeline filters data in different proportions from male and female group, our evaluation confirms it does not cause unfairness. This pipeline has been used by Propublica (Angwin et al., 2016) to show the bias in the prediction. Therefore, it is understandable that they did not employ any preprocessing that introduces bias in the pipelines. Other than that, almost all the preprocessing stages in Bank Marketing pipelines also exhibit very little unfairness, which suggest that the preprocessing on this dataset are fair in general.
A few stages show different behavior when they are used in composition with different classifiers. For example, StandardScaler has been applied on both GC6 and GC8. While GC6 employs a RandomForest classifier, GC8 uses K-Neighbors classifier. We have observed the opposite fairness measures for StandardScaler in these two pipelines. Therefore, fairness can be dominated by the underlying properties of data or the pipeline where it is applied. We have further investigated this phenomenon by applying transformers on different classifiers in the next section.
3. Fairness-Performance Tradeoff
In this section, we investigated the fairness-performance tradeoff for the preprocessing stages. The original performances of each pipeline have been reported in Table 1. To investigate fairness of a stage, we created pipeline by removing the stage from original pipeline (§2). To understand fairness-performance tradeoff, we evaluated performance (both accuracy and f1 score) of and in the same experimental setup. Then we computed the performance difference to observe the impact of the stage on performance. For example, gives the accuracy increase (or decrease, if negative) after applying a stage. We plotted the performance impacts of the stages with their fairness measures in Figure 4.
First, many preprocessing stages have negligible performance impact. In Figure 4, 19 out 63 stages exhibits accuracy and f1 score change in the range [-0.005, 0.005], which indicates performance change 0.05%. We found that in all of these cases, except AC9 and AC10, the preprocessing stages are fair with a very small degree of bias. Second, tradeoff between performance and fairness is observed for the stages which improve performance. 17 stages improve accuracy or f1 score more than 0.05%, which further exhibits moderate to high degree of bias. Overall, the most biased stages - TT7(LE), TT8(CT), TT4(CT), TT1(MV), GC8(SS), are improving performance. This stage-specific tradeoff is aligned with the overall performance-fairness tradeoff discussed in prior work (Biswas and Rajan, 2020; Friedler et al., 2019; Chakraborty et al., 2019), which can be compared quantitatively by the work of Hort et al. (Hort et al., 2021). Third, we found that some stages decrease the performance, either accuracy or f1 score. Surprisingly, most of these stages also exhibit high degree of bias. For instance, the most performance-decreasing stages - BM4(SS), AC7(PCA), GC10 (Undersampling), are showing more bias. Our fairness evaluation would facilitate developers to identify and remove such stages in the pipeline.
Fair Data Transformers
In §5, we found that many data preprocessing stages are biased. Many bias mitigation techniques applied in preprocessing stage have been shown successful (Feldman et al., 2015; Kamiran and Calders, 2012). If we process data with appropriate transformer, then it might be possible to avoid bias and mitigate inherent bias in data or classifier. Even if a data transformer is biased towards a specific group, it could be useful to mitigate bias if original data or model exhibits bias towards the opposite group. To that end, we want to investigate the fairness pattern of the data transformers. However, in our evaluation (Figure 3), some transformers have been used only in specific situations e.g., SMOTE has been only applied on German Credit dataset. What is the fairness of this transformer when used on other datasets and classifier? In this section, we setup experiments to evaluate the fairness of commonly used data transformers on different datasets and classifiers.
First, we collected the classifiers used in each dataset category from the benchmark. Then, for each dataset, we created a set of vanilla pipelines. A vanilla pipeline is a classification pipeline which contain only one classifier. Second, we found a few categories of preprocessing stages from our benchmark shown in Table 2. For each transformer used in each stage, we collected the alternative transformers from corresponding library. For example, in our benchmark, StandardScaler from Scikit-Learn library has been used for scaling data distribution in many pipelines. We collected other standardizing algorithms available in Scikit-Learn. We found that besides StandardScaler, Scikit-Learn also provides MaxAbsScaler, MinMaxScaler, and Normalizer standardize data (Scikit Learn, 2019b). Similarly, a data oversampling technique SMOTE has been used in the benchmark, we collected another undersampling technique ALLKNN and a combination of over- and undersampling sampling technique SMOTENN from IMBLearn library (Lemaître et al., 2017). Third, in each of the vanilla pipelines, we applied the transformers and evaluated fairness using the method used in §4.2 with respect to four metrics. We found that pipelines under Titanic uses custom transformation, and most of the built-in transformers are not appropriate for this dataset. So, to be able to make the comparison consistent, we conducted this evaluation on four datasets: German Credit, Adult Census, Bank Marketing, Compas. Finally, we did not use transformers for imputation and encoder stages. Encoding transformers (LabelEncoder, OneHotEncoder), have been applied on most of the pipelines and their behavior has been understood. The fairness measures of each transformer on different classifiers have been plotted in Figure 5.
Fairness among the datasets follows a similar pattern. This further confirms that the unfairness is rooted in data. The Compas dataset shows the least bias. Although racial discrimination has been reported for this dataset (Angwin et al., 2016), this is a more curated dataset than the other three. By looking at the overall trend of fairness, we observe that sampling techniques have the most biased impact on prediction. Other than that feature selection transformers have more impact than other ones.
Sampling techniques are often used in ML tasks when dataset is class-imbalanced. Unlike the other transformers, sampling techniques make horizontal transformation to the training data. The oversampling technique SMOTE creates new data instances for the minority class by choosing the nearest data points in the feature space. Undersampling techniques balance dataset by removing data items from majority class. Although balancing dataset has been shown to increase fairness (Dixon et al., 2018), our evaluation suggest that in three out of four datasets, it increases bias.
From Figure 5, we can see that sampling techniques exhibit the most unfairness. In German Credit dataset, different classifier reacts differently when sampling is done. DecisionTree classifier exhibits most unfairness for both oversampling and undersampling towards privileged group i.e., male. Interestingly, the combination of over- and undersampling also fails to show fairness. Furthermore, both German Credit and Bank Marketing pipelines exhibit bias towards unprivileged group, which might be desired when compared to bias towards privileged.
Selecting the best performing feature can give performance improvement of the pipeline. However, unfairness can be encoded in specific features (Grgic-Hlaca et al., 2018). While selecting best features, some features which encodes unfairness, can dominate the outcome. Thus, many classifiers in German Credit, Adult Census, and Bank Marketing show unfairness because of reduced number of features, which has been also observed by Zhang and Harman (Zhang and Bareinboim, 2018). Surprisingly, SelectFpr exhibited very little or no bias compared to the other feature selection methods. A detailed investigation suggests that SelectBest and SelectPercentile select only the most contributing features. However, SelectFpr performs false positive rate test on each feature, and if it falls below a threshold, the feature is removed (Scikit Learn, 2019a). Therefore, it does not apply harsh pruning, which contributes to the fairness of the prediction.
These transformers modify the mean and variance of the data by applying linear or non-linear transformation. However, they do not change the feature importance on the classifiers. Therefore, in most of the cases, these transformers (especially, StandardScaler and RobustScaler) are fair. Some classifiers show bias after applying these transformers such as, KNC in Compas. The unfairness exhibited by those pipelines are introduced by the classifiers, since these classifiers show similar bias pattern for other transformers as well. The scalers can impact the fairness significantly if there are many outliers in data. That is why we see more bias for the scalers in German Credit dataset. Therefore, although standardizing transformers are fair in general, they can be biased in composition with specific classifier or data property.
Fairness Composition of Stages
From our evaluation, we found that many data transformers have fairness impact on ML pipeline. In this section, we compare the local fairness (fairness measures of preprocessing stages) with the global fairness (fairness measures of whole pipeline). First, we answer whether the local fairness composes in the global fairness. Second, we investigate if we can leverage the composition to mitigate bias by choosing appropriate transformers.
We evaluated the global fairness of Adult Census pipelines (Table 1) using the four existing metrics from (1). We calculated the fairness difference of these pipelines before and after applying the preprocessing stages. Additionally, we have evaluated the stage-specific fairness metrics. Both the local fairness and difference in global fairness of those pipelines have been plotted in Figure 6.
We can see that local and global fairness follow the same trend in most of the pipelines. This confirms that local fairness is directly contributing to the global fairness. However, the global fairness is computed based on the overall change in the prediction, whereas the local fairness considers the predictions for only those data instances which have been altered after applying a transformer (3). For example, in Figure 6, for some pipelines (e.g., AC9, AC10), global and local fairness exhibit different trends. In these cases, the overall classification rate difference is not similar to the rate difference of altered labels. This means that the stages changed the labels such that it shows bias towards privileged. But when those changes in the labels are considered in addition to all the labels (global fairness), the bias difference could not capture the actual impact of that stage. We have verified this observation by manually inspecting the altered prediction labels. Thus, we can conclude that the local fairness composes to the global fairness. Specifically, if a preprocessing stage shows bias for privileged group, it pulls the global fairness towards the fairness direction of privileged group. However, only observing the global fairness difference, we can not measure the fairness of a given stage or transformer.
2. Bias Mitigation Using Appropriate Transformers
For a given transformer in an ML pipeline, a downstream transformer operates on data already processed by the given transformer and an upstream transformer is applied before the given one. Since the fairness of a preprocessing stage composes to the global fairness, can we choose a downstream transformer to mitigate bias in ML pipeline? In this section, we empirically show that the global unfairness can be mitigated by choosing the appropriate downstream transformer.
Consider the use-case of classification task on German Credit dataset with different classifiers similar to Figure 5. Suppose, the original pipeline is constructed using undersampling technique. Since this pipeline exhibits bias, as shown in Figure 5, can we choose a downstream data standardizing transformer that mitigates that bias? In this use case, undersampling is the upstream transformer, and any standardizing transformer is the downstream transformer.
We showed the evaluation for XGB classifier and KNC classifier, since these two exhibits most bias when the upstream transformer was applied in §6. We plotted the global fairness after applying only the upstream transformer in the left of Figure 7. We also reported the local fairness of the standardizing transformers in Table 3. Now, since undersampling method exhibits bias towards privileged group for XGB, we look for the transformer that is biased towards privileged group. In Figure 7, among other transformers, Normalizer is the most successful to mitigate bias of the upstream transformer. Similarly, for KNC, the upstream operator exhibits bias towards privileged group. From Table 3, we can see that MinMaxScaler is the most biased transformer towards the opposite direction. As a result, applying MinMaxScaler mitigates bias the most. Note that the other downstream transformers also follow the fairness composition with its upstream transformer. Therefore, by measuring fairness of the preprocessing stages, developers would be able to instrument the biased transformers and build fair ML pipelines.
Discussion
We took the first step to understand the fairness of components in ML pipelines. Our method helps to provide causality in software and reason about behavior of components based on the impact on outcome. This method can be extended further to evaluate the fairness of other software modules (Pan and Rajan, 2020) in ML pipeline and localize faults (Wardat et al., 2021). Moreover, we found most of the stages exhibited bias, to a low or higher degree. The fairness measures of different components can be leveraged towards fairness-aware pipeline optimization to satisfy fairness constraints. For example, US Equal Employment Commission suggests selection-rate difference between groups less than 20% (US Equal Employment Opportunity Commission, 1979). Also, pipeline optimization techniques, e.g., TPOT (Olson and Moore, 2016), Lara (Kunft et al., 2019) can be potentially utilized for pipeline optimization.
Furthermore, research has been conducted to understand the impact of preprocessing stages with respect to performance improvement (Crone et al., 2006; Uysal and Gunal, 2014; Chandrasekar and Qian, 2016). This paper will open research directions to develop preprocessing techniques that improve performance by keeping the fairness intact. We also reported a number of fairness patterns of preprocessing stages that inducing bias in the pipeline such as missing value processing, custom feature generation, feature selection. Moreover, instrumentation of the stages can mitigate the inherent bias of the classifiers. It shows opportunities to build automated tools for identifying fairness bugs in AI systems and recommending fixes (Islam et al., 2020, 2019). Finally, current fairness tools (e.g., AIF 360 (Bellamy et al., 2018), Aequitas (Saleiro et al., 2018)) can be augmented by incorporating data preprocessing stages into the pipelines and letting users have control over the data transformers and observe or mitigate bias. Similarly, the libraries can provide API support to monitor fairness of the transformers.
Threats to Validity
Internal validity refers to whether the fairness measures used in this paper actually captures the fairness of preprocessing stages. To mitigate this threat, we used existing concepts and metrics to build new set of metrics. Causality in software (Pearl, 2000, 2009) has been well-studied, and causal reasoning in fairness has also been popular (Kusner et al., 2017; Zhang and Bareinboim, 2018; Salimi et al., 2019; Russell et al., 2017), since it can provide explanation with respect to change in the outcome. Besides, this method do not require an oracle because the prediction equivalences provide necessary information to measure the impact of the intervention (Galhotra et al., 2017). Furthermore, in §7, we conducted experiments on local and global fairness to show how new metrics composes in the pipeline.
External validity is concerned about the extent the findings of this study can be generalized. To alleviate this threat, we conducted experiments on a large number of pipeline variations. We collected the pipelines from three different sources. Moreover, we collected alternative transformers from the ML libraries for comparative analysis. Finally, for the same dataset categories, we used multiple classifiers and fairness metrics so that the findings are persistent.
Related Works
The machine learning community has defined different fairness criteria and proposed metrics to measure the fairness of classification tasks (Zafar et al., 2015; Dixon et al., 2018; Pearl, 2009; Dwork et al., 2012; Feldman et al., 2015; Hardt et al., 2016; Calders and Verwer, 2010; Chouldechova, 2017; Zemel et al., 2013; Speicher et al., 2018). Following the measurement of fairness in ML models, many mitigation techniques have also been proposed to remove bias (Kamiran and Calders, 2012; Zhang et al., 2018; Feldman et al., 2015; Dixon et al., 2018; Calders and Verwer, 2010; Zafar et al., 2015; Chouldechova, 2017; Goh et al., 2016; Kamishima et al., 2012; Hardt et al., 2016; Pleiss et al., 2017; Kamiran et al., 2012). This body of work mostly concentrates on the theoretical aspect of fairness in a single classification task. Recently, software engineering community has also focused on the fairness in ML, mostly on fairness testing (Tramer et al., 2017; Galhotra et al., 2017; Udeshi et al., 2018; Aggarwal et al., 2019). These works propose methods to generate appropriate test data inputs for the model and prediction on those inputs characterizes fairness. Some research has been conducted to build automated tools (Adebayo and Kagal, 2016; Udeshi et al., 2018; Sokol et al., 2019) and libraries (Bellamy et al., 2018) for fairness. In addition, empirical studies have been conducted to compare, contrast between fairness aspects, interventions, tradeoffs, developers concerns, and human aspects of fairness (Friedler et al., 2019; Biswas and Rajan, 2020; Harrison et al., 2020; Holstein et al., 2019; Zhang and Harman, 2021).
Fairness in Composition. Dwork and Ilvento argued that fairness is dynamic in a multi-component environment (Dwork and Ilvento, 2018). They showed that when multiple classifiers work in composition, even if the classifiers are fair in isolation, the overall system is not necessarily fair. Bower et al. discussed fairness in ML pipeline, where they considered pipeline as sequence of multiple classification tasks (Bower et al., 2017). They also showed that when decisions of fair components are compounded, the final decision might not be fair. For example, while interviewing candidates in two stages, fair decision in each stage may not guarantee a fair selection. D’Amour et al. studied the dynamics of fairness in multi-classification environnement using simulation (D’Amour et al., 2020). In these research, fairness composition is shown over multiple tasks and the authors did not consider fairness of components in single ML pipeline. We position our paper here to study the impact of preprocessing stages in ML pipeline and evaluate the fairness composition.
Conclusion
Data preprocessing techniques are used in most of the machine learning pipelines in composition with the classifier. Studies showed that fairness of machine learning predictions depends largely on the data. In this paper, we investigated how the data preprocessing stages affect fairness of classification tasks. We proposed the causal method and leveraged existing metrics to measure the fairness of data preprocessing stages. The results showed that many stages induce bias in the prediction. By observing fairness of these data transformers, fairer ML pipelines can be built. In addition, we showed that existing bias can be mitigated by selecting appropriate transformers. We released the pipeline benchmark, code, and results to make our techniques available for further usages. Future research can be conducted towards developing automated tools to detect bias in ML pipeline stages and instrument that accordingly.