Attributing Fair Decisions with Attention Interventions

Ninareh Mehrabi, Umang Gupta, Fred Morstatter, Greg Ver Steeg, Aram Galstyan

Introduction

Machine learning algorithms that optimize for performance (e.g., accuracy) often result in unfair outcomes (Mehrabi et al. 2021). These algorithms capture biases present in the training datasets causing discrimination toward different groups. As machine learning continues to be adopted into fields where discriminatory treatments can lead to legal penalties, fairness and interpretability have become a necessity and a legal incentive in addition to an ethical responsibility (Barocas and Selbst 2016; Hacker et al. 2020). Existing methods for fair machine learning include applying complex transformations to the data so that resulting representations are fair (Gupta et al. 2021; Moyer et al. 2018; Roy and Boddeti 2019; Jaiswal et al. 2020; Song et al. 2019), adding regularizers to incorporate fairness (Zafar et al. 2017; Kamishima et al. 2012; Mehrabi, Huang, and Morstatter 2020), or modifying the outcomes of unfair machine learning algorithms to ensure fairness (Hardt et al. 2016), among others. Here we present an alternative approach Code can be found at: https://github.com/Ninarehm/Attribution, which works by identifying the significance of different features in causing unfairness and reducing their effect on the outcomes using an attention-based mechanism.

With the advancement of transformer models and the attention mechanism (Vaswani et al. 2017), recent research in Natural Language Processing (NLP) has tried to analyze the effects and the interpretability of the attention weights on the decision making process (Wiegreffe and Pinter 2019; Jain and Wallace 2019; Serrano and Smith 2019; Hao et al. 2021). Taking inspiration from these works, we propose to use an attention-based mechanism to study the fairness of a model. The attention mechanism provides an intuitive way to capture the effect of each attribute on the outcomes. Thus, by introducing the attention mechanism, we can analyze the effect of specific input features on the model’s fairness. We form visualizations that explain model outcomes and help us decide which attributes contribute to accuracy vs. fairness. We also show and confirm the observed effect of indirect discrimination in previous work (Zliobaite 2015; Hajian and Domingo-Ferrer 2013; Zhang, Wu, and Wu 2017) in which even with the absence of the sensitive attribute, we can still have an unfair model due to the existence of proxy attributes. Furthermore, we show that in certain scenarios those proxy attributes contribute more to the model unfairness than the sensitive attribute itself.

Based on the above observations, we propose a post-processing bias mitigation technique by diminishing the weights of features most responsible for causing unfairness. We perform studies on datasets with different modalities and show the flexibility of our framework on both tabular and large-scale text data, which is an advantage over existing interpretable non-neural and non-attention-based models. Furthermore, our approach provides a competitive and interpretable baseline compared to several recent fair learning techniques.

To summarize, the contributions of this work are as follows: (1) We propose a new framework for attention-based classification in tabular data, which is interpretable in the sense that it allows to quantify the effect of each attribute on the outcomes; (2) We then use these attributions to study the effect of different input features on the fairness and accuracy of the models; (3) Using this attribution framework, we propose a post-processing bias mitigation technique that can reduce unfairness and provide competitive accuracy vs. fairness trade-offs; (4) Lastly, we show the versatility of our framework by applying it to large-scale non-tabular data such as text.

Approach

In this section, we describe our classification model that incorporates the attention mechanism. It can be applied to both text and tabular data and is inspired by works in attention-based models in text-classification (Zhou et al. 2016). We incorporate attention over the input features. Next, we describe how this attention over features can attribute the model’s unfairness to certain features. Finally, using this attribution framework, we propose a post-processing approach for mitigating unfairness.

Equality of Opportunity Difference (EqOpp):

The resulting representation, rr, is passed to the feed-forward layers for classification. In this work, we have used two feed-forward layers (See Fig. 1 for overall architecture).

2 Fairness Attribution with Attention Weights

The aforementioned classification model with the attention mechanism combines input feature embeddings by taking a weighted combination. By manipulating the weights, we can intuitively capture the effects of specific features on the output. To this end, we observe the effect of each attribute on the fairness of outcomes by zeroing out or reducing its attention weights and recording the change. Other works have used similar ideas to understand the effect of attention weights on accuracy and evaluate interpretability of the attention weights by comparing the difference in outcomes in terms of measures such as Jensen-Shannon Divergence (Serrano and Smith 2019) but not for fairness. We are interested in the effect of features on fairness measures. Thus, we measure the difference in fairness of the outcomes based on the desired fairness measure. A large change in fairness measure and a small change in performance of the model would indicate that this feature is mostly responsible for unfairness, and it can be dropped without causing large impacts on performance. The overall framework is shown in Fig. 1. First, the outcomes are recorded with the original attention weights intact (Fig. 1(a)). Next, attention weights corresponding to a particular feature are zeroed out, and the difference in performance and fairness measures is recorded (Fig. 1(b)). Based on the observed differences, one may conclude how incorporating this feature contributes to fairness/unfairness.

To measure the effect of the kthk^{th} feature on different fairness measures, we consider the difference in the fairness of outcomes of the original model and model with kthk^{th} feature’s effect removed. For example, for statistical parity difference, we will consider SPD(y^o,a)−SPD(y^zk,a)\text{SPD}(\hat{\mathbf{y}}_{o},\mathbf{a})-\text{SPD}(\hat{\mathbf{y}}_{z}^{k},\mathbf{a}). A negative value will indicate that the kthk^{th} feature helps mitigate unfairness, and a positive value will indicate that the kthk^{th} feature contributes to unfairness. This is because y^zk\hat{y}_{z}^{k} captures the exclusion of the kthk^{th} feature (zeroed out attention weight for that feature) from the decision-making process. If the value is positive, it indicates that not having this feature makes the bias lower than when we include it. Notice here, we focus on global attribution, so we measure this over all the samples; however, this can also be turned into local attribution by focusing on individual sample ii only.

3 Mitigating Bias by Removing Unfair Features

As discussed in the previous section, we can identify features that contribute to unfair outcomes according to different fairness measures. A simple technique to mitigate or reduce bias is to reduce the attention weights of these features. This mitigation technique is outlined in Algorithm 1. In this algorithm, we first individually set attention weights for each of the features in all the samples to zero and monitor the effect on the desired fairness measure. We have demonstrated the algorithm for SPD, but other measures, such as EqOdd, EqOpp, and even accuracy can be used (in which case the “unfair_feature_set” can be re-named to feature set which harms accuracy instead of fairness). If the kthk^{th} feature contributes to unfairness, we reduce its attention weight using decay rate value. This is because y^zk\hat{\mathbf{y}}_{z}^{k} captures the exclusion of the kthk^{th} feature (zeroed attention weight for that feature) compared to the original outcome y^o\hat{\mathbf{y}}_{o} for when all the feature weights are intact; otherwise, we use the original attention weight. We can also control the fairness-accuracy trade-off by putting more attention weight on features that boost accuracy while keeping the fairness of the model the same and down-weighting features that hurt accuracy, fairness, or both.

This post-processing technique has a couple of advantages over previous works in bias mitigation or fair classification approaches. First, the post-processing approach is computationally efficient as it does not require model retraining to ensure fairness for each sensitive attribute separately. Instead, the model is trained once by incorporating all the attributes, and then one manipulates attention weights during test time according to particular needs and use-cases. Second, the proposed mitigation method provides an explanation and can control the fairness-accuracy trade-off. This is because manipulating the attention weights reveals which features are important for getting the desired outcome, and by how much. This provides an explanation for the outcome and also a mechanism to control the fairness-accuracy trade-off by the amount of the manipulation.

Experimental Setup

We perform a suite of experiments on synthetic and real-world datasets to evaluate our attention based interpretable fairness framework. The experiments on synthetic data are intended to elucidate interpretability in controlled settings, where we can manipulate the relations between input and output feature. The experiments on real-world data aim to validate the effectiveness of the proposed approach on both tabular and non-tabular (textual) data. Below we describe the experiments, datasets, and respective baselines.

We enumerate the experiments and their goals as follows: Experiment 1: Attributing Fairness with Attention The purpose of this experiment is to demonstrate that our attribution framework can capture correct attributions of features to fairness outcomes. We present our results for tabular data in Sec. 4.1. Experiment 2: Bias Mitigation via Attention Weight Manipulation In this experiment, we seek to validate the proposed post-processing bias mitigation framework and compare it with various recent mitigation approaches. The results for real-world tabular data are presented in Sec. 4.2. Experiment 3: Validation on Textual Data The goal of this experiment is to demonstrate the flexibility of the proposed attention-based method by conducting experiments on non-tabular, textual data. The results are presented in Sec. 4.3.

2 Datasets

To validate the attribution framework, we created two synthetic datasets in which we control how features interact with each other and contribute to the accuracy and fairness of the outcome variable. These datasets capture some of the common scenarios, namely the data imbalance (skewness) and indirect discrimination issues, arising in fair decision or classification problems.

First, we create a simple scenario to demonstrate that our framework identifies correct feature attributions for fairness and accuracy. We create a feature that is correlated with the outcome (responsible for accuracy), a discrete feature that causes the prediction outcomes to be biased (responsible for fairness), and a continuous feature that is independent of the label or the task (irrelevant for the task). For intuition, suppose the attention-based attribution framework works correctly. In this case, we expect to see a reduction in accuracy upon removing (i.e., making the attention weight zero) the feature responsible for the accuracy, reduction in bias upon removing the feature responsible for bias, and very little or no change upon removing the irrelevant feature. With this objective, we generated a synthetic dataset with three features, i.e., x=[f1,f2,f3]x=[f_{1},f_{2},f_{3}] as followsWe use x∼Ber(p)x\sim\text{Ber}(p) to denote that xx is a Bernoulli random variable with P(x=1)=pP(x=1)=p.:

Clearly, f2f_{2} has the most predictive information for the task and is responsible for accuracy. Here, we consider f1f_{1} as the sensitive attribute. f1f_{1} is an imbalanced feature that can bias the outcome and is generated such that there is no intentional correlation between f1f_{1} and the outcome, yy or f2f_{2}. f3f_{3} is sampled from a normal distribution independent of the outcome yy, or the other features, making it irrelevant for the task. Thus, an ideal classifier would be fair if it captures the correct outcome without being affected by the imbalance in f1f_{1}. However, due to limited data and skew in f1f_{1}, there will be some undesired bias — few errors when f1=0f_{1}=0 can lead to large statistical parity.

Using features that are not identified as sensitive attributes can result in unfair decisions due to their implicit relations or correlations with the sensitive attributes. This phenomenon is called indirect discrimination (Zliobaite 2015; Hajian and Domingo-Ferrer 2013; Zhang, Wu, and Wu 2017). We designed this synthetic dataset to demonstrate and characterize the behavior of our framework under indirect discrimination. Similar to the previous scenario, we consider three features. Here, f1f_{1} is considered as the sensitive attribute, and f2f_{2} is correlated with f1f_{1} and the outcome, yy. The generative process is as follows:

In this case f1f_{1} and yy are correlated with f2f_{2}. The model should mostly rely on f2f_{2} for its decisions. However, due to the correlation between f1f_{1} and f2f_{2}, we expect f2f_{2} to affect both the accuracy and fairness of the model. Thus, in this case, indirect discrimination is possible. Using such a synthetic dataset, we demonstrate a) indirect discrimination and b) the need to have an attribution framework to reason about unfairness and not blindly focus on the sensitive attributes for bias mitigation.

Real-world Datasets

We demonstrate our approach on the following real-world datasets: Tabular Datasets: We conduct our experiments on two real-world tabular datasets often used to benchmark fair classification techniques — UCI Adult (Dua and Graff 2017) and Heritage Healthhttps://www.kaggle.com/c/hhp datasets. The UCI Adult dataset contains census information about individuals, with the prediction task being whether the income of the individual is higher than $50k or not. The sensitive attribute, in this case, is gender (male/female). The Heritage Health dataset contains patient information, and the task is to predict the Charleson Index (comorbidity index, which is a patient survival indicator). Each patient is grouped into one of the 9 possible age groups, and we consider this as the sensitive attribute. We used the same pre-processing and train-test splits as in Gupta et al. (2021). Non-Tabular or Text Dataset: To demonstrate the flexibility of our approach, we also experiment with a non-tabular, text dataset. We used the biosbias dataset (De-Arteaga et al. 2019). The dataset contains short bios of individuals. The task is to predict the occupation of the individual from their bio. We utilized the bios from the year 2018 from the 2018_34 archive and considered two occupations for our experiments, namely, nurse and dentist. The dataset was split into 70-15-15 train, validation, and test splits. De-Arteaga et al. (2019) has demonstrated the existence of gender bias in this prediction task and showed that certain gender words are associated with certain job types (e.g., she to nurse and he to dentist).

3 Bias Mitigation Baselines

We compared our bias mitigation approach to a number of recent state of the art methods. For our experiments with tabular data, we focus on methods that are specifically optimized to achieve statistical parity. Results for other fairness notions can be found in the appendix. For our baselines, we consider methods that learn representations of data so that information about sensitive attributes is eliminated. CVIB (Moyer et al. 2018) realizes this objective through a conditional variational autoencoder, whereas MIFR (Song et al. 2019) uses a combination of information bottleneck term and adversarial learning to optimize the fairness objective. FCRL (Gupta et al. 2021) optimizes information theoretic objectives that can be used to achieve good trade-offs between fairness and accuracy by using specialized contrastive information estimators. In addition to information-theoretic approaches, we also considered baselines that use adversarial learning such as MaxEnt-ARL (Roy and Boddeti 2019), LAFTR (Madras et al. 2018), and Adversarial Forgetting (Jaiswal et al. 2020). Note that in contrast to our approach, the baselines described above are not interpretable as they are incapable of directly attributing features to fairness outcomes. For the textual data, we compare our approach with the debiasing technique proposed in De-Arteaga et al. (2019), which works by masking the gender-related words and then training the model on this masked data.

Results

First, we test our method’s ability to capture correct attributions in controlled experiments with synthetic data (described in Sec. 3.2). We also conduct a similar experiment with UCI Adult and Heritage Health datasets which can be found in the appendix. Fig. 2 summarizes our results by visualizing the attributions, which we now discuss.

In Scenario 1, as expected, f2f_{2} is correctly attributed to being responsible for the accuracy and removing it hurts the accuracy drastically. Similarly, f1f_{1} is correctly shown to be responsible for unfairness and removing it creates a fairer outcome. Ideally, the model should not be using any information about f1f_{1} as it is independent of the task, but it does. Therefore, by removing f1f_{1}, we can ensure that information is not used and hence outcomes are fair. Lastly, as expected, f3f_{3} was the irrelevant feature, and its effects on accuracy and fairness are negligible. Another interesting observation is that f2f_{2} is helping the model achieve fairness since its exclusion means that the model should rely on f1f_{1} for decision making, resulting in more bias; thus, removing f2f_{2} harms accuracy and fairness as expected.

In Scenario 2, our framework captures the effect of indirect discrimination. We can see that removing f2f_{2} reduces bias as well as accuracy drastically. This is because f2f_{2} is the predictive feature, but due to its correlation with f1f_{1}, it can also indirectly affect the model’s fairness. More interestingly, although f1f_{1} is the sensitive feature, removing it does not play a drastic role in fairness or the accuracy. This is an important finding as it shows why removing f1f_{1} on its own can not give us a fairer model due to the existence of correlations to other features and indirect discrimination. Overall, our results are intuitive and thus validate our assumption that attention-based framework can provide reliable feature attributions for the fairness and accuracy of the model.

2 Attention as a Mitigation Technique

As we have highlighted earlier, understanding how the information within features interact and contribute to the decision making can be used to design effective bias mitigation strategies. One such example was shown in Sec. 4.1. Often real-world datasets have features which cause indirect discrimination, due to which fairness can not be achieved by simply eliminating the sensitive feature from the decision process. Using the attributions derived from our attention-based attribution framework, we propose a post-processing mitigation strategy. Our strategy is to intervene on attention weights as discussed in Sec. 2.3. We first attribute and identify the features responsible for the unfairness of the outcomes, i.e., all the features whose exclusion will decrease the bias compared to the original model’s outcomes and gradually decrease their attention weights to zero as also outlined in Algorithm 1. We do this by first using the whole fraction of the attention weights learned and gradually use less fraction of the weights until the weights are completely zeroed out.

For all the baselines described in Sec. 3.3, we used the approach outlined in Gupta et al. (2021) for training a downstream classifier and evaluating the accuracy/fairness trade-offs. The downstream classifier was a 1-hidden-layer MLP with 50 neurons along with ReLU activation function. Our experiments were performed on Nvidia GeForce RTX 2080. Each method was trained with five different seeds, and we report the average accuracy and fairness measure as statistical parity difference (SPD). Results for other fairness notions can be found in the appendix. CVIB, MaxEnt-ARL, Adversarial Forgetting and FCRL are designed for statistical parity notion of fairness and are not applicable for other measures like Equalized Odds and Equality of Opportunity. LAFTR can only deal with binary sensitive attributes and thus not applicable for Heritage Health dataset. Notice that our approach does not have these limitations. For our approach, we vary the attention weights and report the resulting fairness-accuracy trade offs.

Fig. 3 compares fairness-accuracy trade-offs of different bias mitigation approaches. We desire outcomes to be fairer, i.e., lower values of SPD and to be more accurate, i.e., towards the right. The results show that using attention attributions can indeed be beneficial for reducing bias. Moreover, our mitigation framework based on the manipulation of the attention weights is competitive with state-of-the-art mitigation strategies. However, most of these approaches are specifically designed and optimized to achieve parity and do not provide any interpretability. Our model can not only achieve comparable and competitive results, but it is also able to provide explanation such that the users exactly know what feature and by how much it was manipulated to get the corresponding outcome. Another advantage of our model is that it needs only one round of training. The adjustments to attention weights are made post-training; thus, it is possible to achieve different trade-offs. Moreover, our approach does not need to know sensitive attributes while training; thus, it could work with other sensitive attributes not known beforehand or during training. Lastly, here we merely focused on mitigating bias (as our goal was to show that the attribution framework can identify problematic features and their removal would result in bias mitigation) and did not focus too much on improving accuracy and achieving the best trade-off curve which can be considered as the current limitation of our work. We manipulated attention weights of all the features that contributed to unfairness irrespective of if they helped maintaining high accuracy or not. However, the trade-off results can be improved by carefully considering the trade-off each feature contributes to with regards to both accuracy and fairness (e.g., using results from Fig. 2) to achieve better trade-off results which can be investigated as a future direction (e.g., removing problematic features that contribute to unfairness only if their contribution to accuracy is below a certain threshold value). The advantage of our work is that this trade-off curve can be controlled by controlling how many features and by how much to be manipulated which is not the case for most existing work.

3 Experiments with Non-Tabular Data

In addition to providing interpretability, our approach is flexible and useful for controlling fairness in modalities other than tabular datasets. To put this to the test, we applied our model to mitigate bias in text-based data. We consider the biosbias dataset (De-Arteaga et al. 2019), and use our mitigation technique to reduce observed biases in the classification task performed on this dataset. We compare our approach with the debiasing technique proposed in the original paper (De-Arteaga et al. 2019), which works by masking the gender-related words and then training the model on this masked data. As discussed earlier, such a method is computationally inefficient. It requires re-training the model or creating a new masked dataset, each time it is required to debias the model against different attributes, such as gender vs. race. For the baseline pre-processing method, we masked the gender-related words, such as names and gender words, as provided in the biosbias dataset and trained the model on the filtered dataset. On the other hand, we trained the model on the raw bios for our post-processing method and only manipulated attention weights of the gender words during the testing process as also provided in the biosbias dataset.

In order to measure the bias, we used the same measure as in (De-Arteaga et al. 2019) which is based on the equality of opportunity notion of fairness (Hardt et al. 2016) and reported the True Positive Rate Difference (TPRD) for each occupation amongst different genders. As shown in Table 1, our post-processing mitigation technique provides lower TRPD while being more accurate, followed by the technique that masks the gendered words before training. Although both methods reduce the bias compared to a model trained on raw bios without applying any mask or invariance to gendered words, our post-processing method is more effective. Fig. 4 also highlights qualitative differences between models in terms of their most attentive features for the prediction task. As shown in the results, our post-processing technique is able to use more meaningful words, such as R.N. (registered nurse) to predict the outcome label nurse compared to both baselines, while the non-debiased model focuses on gendered words.

Related Work

Fairness. The research in fairness concerns itself with various topics, such as defining fairness metrics, proposing solutions for bias mitigation, and analyzing existing harms in various systems (Mehrabi et al. 2021). In this work, we utilized different metrics that were introduced previously, such as statistical parity (Dwork et al. 2012), equality of opportunity and equalized odds (Hardt et al. 2016), to measure the amount of bias. We also used different bias mitigation strategies to compare against our mitigation strategy, such as FCRL (Gupta et al. 2021), CVIB (Moyer et al. 2018), MIFR (Song et al. 2019), adversarial forgetting (Jaiswal et al. 2020), MaxEnt-ARL (Roy and Boddeti 2019), and LAFTR (Madras et al. 2018). We also utilized concepts and datasets that were analyzing existing biases in NLP systems, such as (De-Arteaga et al. 2019) which studied the existing biases in NLP systems on the occupation classification task on the bios dataset.

Interpretability. In this work, we introduced an attribution framework based on the attention weights that can analyze fairness and accuracy of the models at the same time and reason about the importance of each feature on fairness and accuracy. There is a body of work in NLP literature that tried to analyze the effect of the attention weights on interpretability of the model (Wiegreffe and Pinter 2019; Jain and Wallace 2019; Serrano and Smith 2019). Other work also utilized attention weights to define an attribution score to be able to reason about how transformer models such as BERT work (Hao et al. 2021). Notice that although Jain and Wallace (2019) claim that attention might not be explanation, a body of work has proved otherwise including (Wiegreffe and Pinter 2019) in which authors directly target the work in Jain and Wallace (2019) and analyze in detail the problems associated with this study. In our work, we also find that attention can be useful and can extract meaningful information which can be beneficial in many aspects. In addition, Vig et al. (2020) analyze the effect of the attention weights in transformer models for bias analysis in language models. However, their approach is different and has a more causal take on investigating the bias. Their study is specific to language models and does not necessarily apply to broader tasks and existing fairness definitions. Aside from interpretability and fairness, we utilized concepts from the NLP literature for designing our attention-based model that can be applicable to tabular data (Vaswani et al. 2017; Zhou et al. 2016).

Discussion

In this work, we analyzed how attention weights contribute to fairness and accuracy of a predictive model. To do so, we proposed an attribution method that leverages the attention mechanism and showed the effectiveness of this approach on both tabular and text data. Using this interpretable attribution framework we then introduced a post-processing bias mitigation strategy based on attention weight manipulation. We validated the proposed framework by conducting experiments with different baselines, as well as fairness metrics, and different data modalities.

Although our work can have a positive impact in allowing to reason about fairness and accuracy of models and reduce their bias, it can also have negative societal consequences if used unethically. For instance, it has been previously shown that interpretability frameworks can be used as a means for fairwashing which is when malicious users generate fake explanations for their unfair decisions to justify them (Anders et al. 2020). In addition, previously it has been shown that interpratability frameworks are vulnerable against adversarial attacks (Slack et al. 2020). We acknowledge that our framework may also be targeted by malicious users for malicious intent that can manipulate attention weights to either generate fake explanations or unfair outcomes. An important future direction can be to analyze and improve robustness of our framework along with others.

References

Appendix A Appendix

We included additional bias mitigation results using other fairness metrics, such as equality of opportunity and equalized odds on both of the Adult and Heritage Health datasets in this supplementary material. We also included additional post-processing results along with additional qualitative results both for the tabular and non-tabular dataset experiments. More details can be found under each sub-section.

Here, we show the results of our mitigation framework considering equality of opportunity and equalized odds notions of fairness. We included baselines that were applicable for these notions. Notice not all the baselines we used in our previous analysis for statistical parity were applicable for equality of opportunity and equalized odds notions of fairness; thus, we only included the applicable ones. In addition, LAFTR is only applicable when the sensitive attribute is a binary variable, so it was not applicable to be included in the analysis for the heritage health data where the sensitive attribute is non-binary. Results of these analysis is shown in Figures 6 and 7. We once again show competitive and comparable results to other baseline methods, while having the advantage of being interpretable and not requiring multiple trainings to satisfy different fairness notions or fairness on different sensitive attributes. Our framework is also flexible for different fairness measures and can be applied to binary or non-binary sensitive features.

In addition, we show how different features contribute differently under different fairness notions. Fig. 8 demonstrates the top three features that contribute to unfairness the most along with the percentages of the fairness improvement upon their removal for each of the fairness notions. As observed from the results, while equality of opportunity and equalized odds are similar in terms of their problematic features, statistical parity has different trends. This is also expected as equality of opportunity and equalized odds are similar fairness notions in nature compared to statistical parity.

We also compared our mitigation strategy with the Hardt etl al. post-processing approach (Hardt et al. 2016). Using this post-processing implementation https://fairlearn.org, we obtained the optimal solution that tries to satisfy different fairness notions subject to accuracy constraints. For our results, we put the results from zeroing out all the attention weights corresponding to the problematic features that were detected from our interpretability framework. However, notice that since our mitigation strategy can control different trade-offs we can have different results depending on the scenario. Here, we reported the results from zeroing out the problematic attention weights that is targeting fairness mostly. From the results demonstrated in Tables 3 and 4, we can see comparable numbers to those obtained from (Hardt et al. 2016). This again shows that our interpretability framework yet again captures the correct responsible features and that the mitigation strategy works as expected.

A.2 Results on non-tabular Data

We also included some additional qualitative results from the experiments on non-tabular data in Fig. 5.

A.3 Interpreting Fairness with Attention

Fig. 9 shows results on a subset of the features from the UCI Adult and Heritage Health datasets (to keep the plots uncluttered and readable, we incorporated the most interesting features in the plot), and provide some intuition about how different features in these datasets contribute to the model fairness and accuracy. While features such as capital gain and capital loss in the UCI Adult dataset are responsible for improving accuracy and reducing bias, we can observe that features such as relationship or marital status, which can be indirectly correlated with the feature sex, have a negative impact on fairness. For the Heritage Health dataset, including the features drugCount ave and dsfs max provide accuracy gains but at the expense of fairness, while including no Claims and no Specialities negatively impact both accuracy and fairness.

A.4 Information on Datasets and Features

More details about each of the datasets along with the descriptions of each feature for the Adult dataset can be found athttps://archive.ics.uci.edu/ml/datasets/adult and for the Heritage Health dataset can be found at https://www.kaggle.com/c/hhp. In our qualitative results, we used the feature names as marked in these datasets. If the names or acronyms are unclear kindly reference to the references mentioned for more detailed description for each of the features. Although most of the features in the Adult datasets are self-descriptive, Heritage Health dataset includes some abbreviations that we list in Table 2 for the ease of interpreting each feature’s meaning.