Fairness Through Causal Awareness: Learning Latent-Variable Models for Biased Data

David Madras, Elliot Creager, Toniann Pitassi, Richard Zemel

Introduction

In this work, we consider the problem of fair decision-making from biased datasets. Much work has been done recently on the problem of fair classification (Zafar et al., 2017; Hardt et al., 2016; Bechavod and Ligett, 2017; Agarwal et al., 2018), yielding an abundant supply of definitions, models, and algorithms for the purposes of learning classifiers whose outputs satisfy distributional constraints. Some of the canonical problems for which these algorithms have been proposed are loan assignment (Hardt et al., 2016), criminal risk assessment (Chouldechova, 2017), and school admissions (Friedler et al., 2016). However, none of these problems are fully specified by the classification paradigm. Rather, they are decision-making problems: each problem requires an action (or “treatment”) to be taken in the world, which in turn yields an outcome. In other words, the central question is how to intervene in an ongoing and evolving process, rather than predict outcomes alone (Barabas et al., 2018).

Decision-making, i.e. learning to intervene, requires a fundamentally different approach from learning to classify: historical training data are the product of past interventions and thus provide an incomplete view of all possible outcomes. Only actions which were previously chosen yield observable outcomes in the training data, while the implicit counterfactual outcomes (the outcome that would have occurred had another action been taken) are never observed. The incompleteness of this data can have great impact on learning and inference (Rubin, 1976).

It has been widely argued that biased data yields unfair machine learning systems (Kallus and Zhou, 2018; Hashimoto et al., 2018; Ensign et al., 2018). In this work we examine dataset bias through the lens of causal inference. To understand how past decisions may bias a dataset, we first must understand how sensitive attributes may have affected the generative process which created the dataset, including the (historical) decision makers’ actions (treatments) and results (outcomes). Causal inference is well suited to this task: since we are interested in decision-making rather than classification, we should be interested in the causal effects of actions rather than correlations. Causal inference has the added benefit of answering counterfactual queries: What would this outcome have been under another treatment? How would the outcome change if the sensitive attribute were changed, all else being equal? These questions are core to the mission of learning fair systems which aim to inform decision-making (Kusner et al., 2017).

While there is much that causal inference can offer to the field of fair machine learning, it also poses several significant challenges. For example, the presence of hidden confounders—unobserved factors that effect both the historical choice of treatment and the outcome—often prohibits the exact inference of causal effects. Additionally, understanding effects at the individual level can be especially complex, particularly if the outcome is non-linear in the data and treatments. These technical difficulties are often amplified by the problem scope of modern machine learning, where large and high-dimensional datasets are commonplace.

To address these challenges, we propose a model for fairly estimating individual-level causal effects from biased data, which combines causal modeling (Pearl, 2009) with approximate inference in deep latent variable models (Kucukelbir et al., 2017; Louizos et al., 2017). Our focus on individual-level causal effects and counterfactuals provides a natural fit for application areas requiring fair policies and treatments for individuals, such as finance, medicine, and law. Specifically, we incorporate the sensitive attribute into our model as a confounding factor, which can possibly influence both the treatment and the outcome. This is a first step towards achieving “fairness through awareness” (Dwork et al., 2012) in the interventional setting.

Our model also leverages recent advances in deep latent-variable modeling to model potential hidden confounders as well as complex, non-linear functions between variables, which greatly increases the class of relationships which it can represent. Through experimental analysis, we show that our model can outperform non-causal models, as well as causal models which do not consider the sensitive attribute as a confounder. We further explore the performance of this model, showing that fair-aware causal modeling can lead to more accurate, fairer policies in decision-making systems.

Background

We employ Structural Causal Models (SCMs), which provide a general theory for modeling causal relationships between variables (Pearl, 2009). An SCM is defined by a directed graph, containing vertices and edges, which respectively represent variables in the world and their pairwise causal relationships. There are two types of vertices: exogenous variables U\mathcal{U} and endogenous variables V\mathcal{V}. Exogenous variables are unspecified by the model; we model them as unexplained noise distributions, and they have no parents. Endogenous variables are the objects we wish to understand; they are descendants of endogenous variables. The value of each endogenous variable is fully determined by its ancestors. Each V∈VV\in\mathcal{V} has some function fVf_{V} which maps the values of its immediate parents to its own. This function fVf_{V} is deterministic; any randomness in an SCM is due to its exogenous variables.

In this paper, we are primarily concerned with three endogenous variables in particular: XX, the observable features (or covariates) of some example; TT, a treatment which is applied to an example; and YY, the outcome of a treatment. Our decision problem is: given an example with particular values for its features, X=xX=x, what value should we assign to treatment TT in order to produce the best outcome YY? This is fundamentally different from a classification problem, since typically we observe the result of only one treatment per example Note that we use the terms treatment and outcome as general descriptors of a decision made/action taken and its result, respectively. These terms are associated with an alternative theory of causal inference (Rubin, 2005) which can also be used to describe the methods we propose, but which we will not discuss in this paper. .

To answer this decision problem, we need to understand the value YY will take if we intervene on TT and set it to value tt. Our first instinct may be to estimate P(Y∣T=t,X=x)P(Y|T=t,X=x). However, this is unsatisfactory in general. If we are estimating these probabilities from observational data, then the fact that xx received treatment tt in the past may have some correlation with the historical outcome YY. This “confounding” effect—the fact that XX has an effect on both TT and YY is depicted in Figure 1(a), by the arrows pointing out of XX into TT and YY. For instance, in an observational medical trial, it is possible that young people are more likely to choose a treatment, and also that young people are more likely to recover. A supervised learning model, given this data, may then overestimate the average effectiveness of the treatment on a test population. Broadly, to understand the effect of assigning treatment tt, supervised learning is not enough; we need to model the functions {fV}\{f_{V}\} of the SCM.

Once we have a fully defined SCM, we can use the dodo operation (Pearl, 2009) to simulate the distribution over YY given that we assign some treatment tt—we denote this as P(Y∣do(T=t),X=x)P(Y|do(T=t),X=x). We do the dodo through graph surgery: we assign the value tt to TT by removing all arrows going into TT from the SCM and setting the corresponding structural equation output to the desired value regardless of its input fT(⋅)=tf_{T}(\cdot)=t. We then set X=xX=x and continue with inference of YY as we normally would.

A common assumption in causal modelling is the “no hidden confounders” assumption, which states that there are no unobserved variables affecting both the treatment and outcome. We follow Louizos et al. (2017), and use variational inference to model confounders that are not directly observed but can be abstracted from proxies. In Sec. 4 we consider the implications of this approach and discuss alternative assumptions.

2. Approximate Inference

Individual and population-level causal effects can be estimated via the do operation when the values of all confounding variables are observed (Pearl, 2009), which motivates the common no-hidden-confounders assumption in causal inference. However this assumption is rather strong and precludes classical causal inference in many situations relevant to fair machine learning, e.g., where ill-quantified and hard-to-observe factors such socio-economic status (SES) may significantly confound the observable data. Therefore we follow Louizos et al. (2017) in modeling unobserved confounders using a high dimensional latent variable ZZ to be inferred for each observation (X,T,Y)(X,T,Y). They prove that if the full joint distribution is successfully recovered, individual treatment effects are identifiable, even in the presence of hidden confounders. In other words, causal effects are identifiable insofar as exact inference can be carried out, and the observed covariates are sufficiently informative.

Because exact inference of ZZ is intractable for many interesting models, we approximately infer ZZ by variational inference, specifying q(Z∣X,T,Y)q(Z|X,T,Y) using a parametric family of distributions and learning the parameters that best approximate the true posterior p(Z∣X,T,Y)p(Z|X,T,Y) by maximizing the evidence lower bound (ELBO) of the marginal data likelihood (Wainwright et al., 2008). In particular, we amortize inference by training a neural network (whose functional form is specified separately from the causal model) to predict the parameters of qq given (X,T,Y)(X,T,Y) (Kingma and Welling, 2014). Amortized inference is much faster but less optimal than local inference (Kim et al., 2018); alternate inference strategies could be explored for applications where the importance of accuracy in individual estimation justifies the additional computational cost.

3. TARNets

TARNets (Shalit et al., 2017) are a class of neural network architectures for estimating outcomes of a binary treatment. The network comprises two separate arms—each predicts the outcomes associated with a separate treatment—that share parameters in the lower layers. The entire network is trained end to end using gradient-based optimization, but with only one arm (the one with the treatment which was actually given) receiving error signal for any given example. The TARNet prediction of result RR and input variables VV and potential intervention II is expressed by combining the shared representation function Φ\Phi with the two functions h0,h1h_{0},h_{1} corresponding to the separate prediction arms. This yields two composed functions,

with h0,h1,Φh_{0},h_{1},\Phi realized as neural networks. Shalit et al. (2017) explore a group-wise MMD penalty on the outputs of Φ\Phi; we do not use this.

Fair Causal Inference

As stated in Sec. 2.1, we are interested in modeling the causal effects of treatments on outcomes. However, when attempting to learn fairly from a biased dataset, this problem takes on an extra dimension. In this context, we become concerned with understanding causal effects in the presence of a sensitive attribute (or protected attribute). Examples include race, gender, age, or SES. When learning from a historical data, we may believe that one of these attributes affected the observable treatments and outcomes, resulting in a biased dataset.

Lum and Isaac (2016) give an example in the domain of predictive policing of how a dataset of drug crimes may become biased with respect to race through unfair policing practices. They note that it is impossible to collect a dataset of all drug crimes in some area; rather, these datasets are really tracking drug arrests. Due to a higher level of police presence in heavily Black than heavily White communities, recorded drug arrests will by nature over-represent Black communities. Therefore, a predictive policing algorithm which attempts to fit this data will continue the pattern of over-policing Black communities. Lum and Isaac (2016) provide experimental validation of this hypothesis through simulation, contrasting the output of a common predictive policing algorithm with independent, demographic-based estimates of drug use by neighborhood. Their work shows that wrongly specifying a learning problem as one of supervised classification can lead to replicating past biases. In order to account for this in the learning process, we should be aware of the biases which shaped the data — which may include sensitive attributes that historically affected the treatment and/or outcome.

Using the above example for concreteness, we specify the variables at play. The decision-making problem is: should police be sent to neighborhood XX at a given time? The variables are:

A∈{0,1}A\in\{0,1\}: a sensitive attribute. For example the majority race of a neighborhood.

T∈{0,1}T\in\{0,1\}: a treatment. For example the presence or absence of police in a certain neighborhood on a particular day.

Y∈\mathdsRY\in\mathds{R}: an outcome. For example the number of arrests recorded in a given neighborhood on a particular day.

X∈\mathdsRDX\in\mathds{R}^{D}: DD-dimensional observed features. For example statistics about the neighborhood, which may change day-to-day

We will represent sensitive attributes and treatments as binary throughout this paper; we recognize this is not always an optimal modeling choice in practice. Note that the choice of treatment will causally alter the outcome—an arrest cannot occur if there are no police in the area. Furthermore, the sensitive attribute can causally effect the outcome as well; research has shown that policing can disparately effect various races, even controlling for police presence (Gelman et al., 2007) (the treatment in this case).

We note that in various domains, there may be more variables of interest than the ones we list here, and more appropriate causal models than those shown in Fig. 1. However, we believe that the setup we describe is widely applicable and contains the minimal set of variables to be useful for fairness-aware causal analysis. We are interested in calculating causal effects between the above variables. In particular, we seek answers to the following three questions:

This will help us to understand which TT is likely to produce a favorable outcome for a given XX. Let us denote yT=t(x,a)=\mathdsE[y∣do(T=t),X=x,A=a]y_{T=t}(x,a)=\mathds{E}[y|do(T=t),X=x,A=a] as the expected conditional outcome under T=tT=t, that is, the ground truth value taken by YY when the treatment TT is assigned the value tt, and conditioning on the values x,ax,a for the features and sensitive attribute respectively. Then, we can express the individual effect of TT on YY (IET→YIE_{T\rightarrow Y}) as

This allows us to understand how the treatment assignment was biased in the data. Similarly, we can define tA=a(x)=\mathdsE[t∣do(A=a),X=x]t_{A=a}(x)=\mathds{E}[t|do(A=a),X=x], which is the expected conditional treatment in the historical data when the value aa is assigned to the sensitive attribute. Then, the individual effect of AA on TT can be expressed as

This allows us to understand what bias is introduced into the historically observed outcome. We can also define yA=a(x)=\mathdsE[y∣do(A=a),X=x,T=tA=a(x)]y_{A=a}(x)=\mathds{E}[y|do(A=a),X=x,T=t_{A=a}(x)] as the expected conditional outcome under A=aA=a; the ground truth value of YY conditioned on the features being xx if the sensitive attribute were assigned the value aa, and the treatment TT were assigned the ground truth value tA=a(x)t_{A=a}(x). Then, we can express the individual effect of AA on YY as

1. Intervening on Sensitive Attributes

There has been some disagreement around the notion of intervening on an immutable (or effectively immutable) sensitive attribute. Holland (1986) argue that there is “no causation without manipulation” — i.e. an attribute can never be a cause; only an experience undergone can be. Briefly stated, they argue that if the factual and counterfactual versions cannot be “defined in principle, it is impossible to define the causal effect”. In a counterargument, Marini and Singer (1988) claim that a “synthesis of intrinsic and extrinsic determination [provides] a more adequate picture of causal relations” — meaning that both externally imposed experiences (extrinsic) and internally defined attributes (intrinsic) are valid conceptual components of a theory of causation. We agree with this view — that the notion of a causal effect of an immutable attribute is valid, and believe that it is particularly useful in a fairness context.

Specifically pertaining to race, some argue it is possible to understand the causal effect of an immutable attribute in terms of the effects of more manipulable attributes (proxies). VanderWeele and Robinson (2014) argue that, rather than interpreting a causal effect estimate of AA as a hypothetical randomized intervention on AA, one can interpret it as a particular type of intervention on some other set of manipulable variables related to AA (under certain graphical and distributional assumptions on those variables). Sen and Wasow (2016) take a constructivist approach, and consider race to be composed of constituent parts, some of which can be theoretically manipulated. They describe several experimental designs which could estimate the effects of immutable attributes.

Another issue with intervening on sensitive attributes is that, since many are “assigned at conception”, all observed covariates XX are post-treatment (Sen and Wasow, 2016) (as reflected in the design of our SCM in Fig. 1(d)). In statistical analysis, a frequent approach is to ignore all post-treatment variables to avoid introducing collider biases (Gelman et al., 2007; King et al., 1994). However, in our model, the purpose of the covariates is to deduce the true (unobserved) values of the latent ZZ for that individual. Therefore, when conditioning on the observed covariates, correlation of AA and ZZ is the objective, rather than an undesired side effect. This is the first step (“Abduction”) of computing counterfactuals (according to Pearl (2009)); we can think of this as adjusting for bias (of the sensitive attribute) in the XX-generating process.

Proposed Method

In this section we first conceptualize and describe our proposed causal model—depicted in Fig. 1(d)—then discuss the parameterization of the corresponding SCMs and learning procedure. A common causal modelling approach is to define a new SCM for each problem Pearl (2009), taking advantage of domain specific knowledge for that particular problem. This stands in contrast to a classic machine learning (ML) approach, which aims to process data and draw conclusions as generally as possible, by automatically discovering patterns of correlation in the data. While the causal modelling approach is capable of detecting effects the ML approach cannot, the ML approach is attractive since it provides modularity, generality and a more automated data processing pipeline. In this work, we aim to interpolate between the two approaches by considering a single, general causal model for observational data. Our model contains what we argue are a minimal set of fairly general causal variables for discovering treatment effects and biases in the data-generation process, allowing us to interface causally with arbitrary data that fits the proposed structure.

Two features of our causal model are noteworthy. First is the explicit consideration of the sensitive attribute—a potential source of dataset bias—as a confounder, which causally affects both the treatment TT and the outcome YY. This contrasts with approaches from outside the fairness literature (e.g. (Louizos et al., 2017), Fig. 1(b)), which in a fairness setting (Fig. 1(c)) would treat potential sensitive attributes as equivalent to other observed features. Our model accounts for the possibility that a sensitive attribute may have causal influence on the observed features, treatments and outcomes and the historical process which generated them. It makes the sensitive attribute distinct from the other attributes of XX, which we understand not as confounders but observed proxies. We can think of this as a causal modeling analogue of “fairness through awareness”. By actively adjusting for causal confounding effects of sensitive attributes, we can build a model which accounts for the interplay between the treatment and outcome for both values of the sensitive attribute.

The other noteworthy aspect of our model is the latent variable ZZ. Together, ZZ and AA make up all the confounding variables. We note two important points about these confounders. Firstly, we clarify that the model class we propose (a latent Gaussian and a deep neural network), is not necessarily the definitive model of the confounders of TT and YY; however, it is a flexible one, with numerous applications in machine learning (Rezende et al., 2014). Secondly, we note that causal inference and machine learning have different conventions around unobserved (i.e. latent) variables — in causal inference, these variables are generally considered to be nameable objects in the world (e.g. SES, historical predjudice), whereas in machine learning they represent some unspecified (and perhaps abstract) structure in the data. Our ZZ follows the machine learning convention.

As in Louizos et al. (2017), ZZ represents all the unobserved confounding variables which effect the outcomes or treatments (other than AA). The features XX can be seen as proxies (noisy observations) for the confounders (Z,AZ,A). Altogether, the endogenous variables in our model are XX, AA, ZZ, TT, and YY. We also have exogenous variables ϵX,ϵA,ϵZ,ϵT,ϵY\epsilon_{X},\epsilon_{A},\epsilon_{Z},\epsilon_{T},\epsilon_{Y} (not shown), each the immediate parent of (only) their respective endogenous variable. The structural equations are:

Since ZZ does not necessarily refer to tangible objects in the world, it is reasonable that Z⊥AZ\perp A in our model. This does not prevent a characteristic such as SES (which may be correlated with AA) from being a confounder — rather, ZZ could represent the component of SES which is not based on AA. Since both confounders are inputs to all other variables in the SCM, the model can learn to represent variables which are based on AA, (e.g. SES) as a joint distribution of ZZ and AA.

With this SCM in hand, we can estimate various interventional outcomes, if we know the values of fV ∀ V∈{Z,A,X,T,Y}f_{V}\ \forall\ V\in\{Z,A,X,T,Y\}. For instance, we might estimate:

which are the expected values over outcomes of interventions on TT, TT and AA, and just AA, respectively.

However, the problem with the calculations in Eq. 6 is that ZZ is unobserved, so we cannot simply condition on its value. Rather, we observe some proxies XX. Since the structural equations go the other direction — XX is a function of ZZ, not the other way around — inferring ZZ from a given XX is a non-trivial matter.

In summary, we need to learn two things: a generative model which can approximate the structural functions ff, and an inference model which can approximate the distribution of ZZ given XX. Following the lead of Louizos et al. (2017), we use variational inference parametrized by deep neural networks to learn the parameters of both of these models jointly. In variational inference, we aim to learn an approximate distribution over the joint variables P(Z,A,X,T,Y)P(Z,A,X,T,Y), by maximizing a variational lower bound on the log-probability of the observed data. As demonstrated in Louizos et al. (2017), the causal effects in the model become identifiable if we can learn this joint distribution. We extend their proof in Appendix B to show identifiability holds when including the sensitive attribute in the model (as in Fig. 1(d)).

We discuss here the identifiability condition from Louizos et al. (2017). Given some treatment TT and outcome YY, the classic “no hidden confounders” assumption asserts that the set of observed variables OO blocks all backdoor paths from TT to YY. Louizos et al. (2017) weaken this: they assume that there is a set of confounding variables Z=OZ∪UZZ=O_{Z}\cup U_{Z} such that ZZ blocks all backdoor paths from TT to YY, where OZO_{Z} are observed and UZU_{Z} are unobserved. They claim that if we recover the full joint distribtuion of p(Z,X,T,Y)p(Z,X,T,Y), then we can identify the causal effect T→YT\rightarrow Y. However, this is only possible if we have sufficiently informative proxies XX. While recovering the full joint distribution does not mean we have to measure every confounder, we do have to at least measure some proxy for each confounder.

This is a weaker assumption, but not fully general. There may be confounding factors which cannot be inferred from the proxies XX — in this case, our model will be unable to learn the joint distribution, and the causal effect will be unidentifiable. In this case, we are back to square one; our causal estimates may be inaccurate. Determining the exact fairness implications of this remains an open problem — it would depend on which confounders were missing, and which proxies were already collected. A complicating factor is that testing for unconfoundedness is difficult, and usually requires making further assumptions (Tran et al., 2016). Therefore we might unintentionally make unfair inferences if we are unaware that we cannot infer all confounders. If we think this is the case, one solution is to collect more proxies. This provides an alternative motivation for the idea of increasing fairness by measuring additional variables (Chen et al., 2018).

To learn a generative model of the data which is faithful to the structural model defined in Eq. 4, we define distributions pp which will approximate various conditional probabilities in our model. We model the joint probability assuming the following factorization:

Each of these pp corresponds to an ff in Eq. 4 — formally, p(V=v∣W=w)=PϵV[fV(W,ϵV)]p(V=v|W=w)=P_{\epsilon_{V}}[f_{V}(W,\epsilon_{V})] for an endogenous variable VV and subset of endogenous variables WW, where {V},W⊂{Z,A,X,T,Y}\{V\},W\subset\{Z,A,X,T,Y\}. For simplicity, we choose computationally tractable probability distributions for each conditional probability in Eq. 7:

where DZ,DXD_{Z},D_{X} are the dimensionalities of ZZ and XX respectively, and πA∈\pi_{A}\in is the empirical marginal probability of A=1A=1 across the dataset (if this is unknown, we could use a Beta prior over that distribution; in this paper we assume AA is observed for every example). For p(Y∣Z,A,T)p(Y|Z,A,T), we use either a Bernoulli or a Gaussian distribution, depending on if YY is binary or continuous:

To flexibly model the potentially complex and non-linear relationships in the true generative process, we specify several of the distribution parameters from Eqs. 8 and 9 as the output of a function gVg_{V}, which is realized by a neural network (or TARNet (Shalit et al., 2017)) with parameters θV\theta_{V}. We parametrize the model of XX with neural networks gXμ,gXσg_{X}^{\mu},g_{X}^{\sigma}:

We use TARNets (Shalit et al., 2017) (see Sec. 2.3) to parameterize the distributions over TT and YY. In our model, AA acts as the “treatment” for the TARNet that outputs TT. Likewise AA and TT are joint treatments affecting YY — our YY model can be seen as a hierarchical TARNet, with one TARNet for each value of AA, where each TARNet has an arm for each value of TT. In all, this yields the following parametrization:

and the same for μY(Z,A,T)\mu_{Y}(Z,A,T) and σY2(Z,A,T)\sigma_{Y}^{2}(Z,A,T); where σ\sigma is the sigmoid function σ(x)=11+exp⁡(−x)\sigma(x)=\frac{1}{1+\exp(-x)} and gRI=0(V,I)g^{I=0}_{R}(V,I) are defined as in Sec. 2.3.

We further define an inference model qq, to determine the values of the latent variables ZZ given observed X,AX,A. This takes the form:

where the normal distribution is reparametrized analogously to Eq. 10 with networks gZμ,gZσg_{Z}^{\mu},g_{Z}^{\sigma}. Since AA is always observed, we do not need to infer it, even though it is a confounder. We note that this is a different inference network from the one in Louizos et al. (2017) — we do not use the given treatments and outcomes in the inference model. We found it to be a simpler solution (no auxiliary networks necessary), and did not see a large change in performance. This is similar to the approach taken in Parbhoo et al. (2018).

To learn the parameters of this model, we can maximize the expected lower bound on the log probability of the data (the ELBO), which takes the form below, which we note is also a valid ELBO to optimize for lower-bounding the conditional log-probability of the treatments and outcomes given the data.

Related Work

Our work most closely relates to the Causal Effect Variational Autoencoder (Louizos et al., 2017). Some follow-up work is done by Parbhoo et al. (2018), who suggest a purely discriminative approach using the information bottleneck. Our model differs from this work in that they did not include a sensitive attribute in their model, and their model does not contain a “reconstruction fidelity” term, in this case log⁡p(xi∣zi,ai)\log{p(x_{i}|z_{i},a_{i})}. Previous papers which learn causal effects using deep learning (with all confounders observed) include Shalit et al. (2017) and Johansson et al. (2016), who propose TARNets as well as some form of balancing penalty.

The intersection of fairness and causality has been explored recently. Counterfactual fairness — the idea that a fair classifier is one which doesn’t change its prediction under the counterfactual value of XX when AA is flipped — is a major theme (Kusner et al., 2017). Criteria for fairness in treatments are proposed in Nabi and Shpitser (2018), and fair interventions are further explored in Kusner et al. (2018). Zhang and Bareinboim (2018) present a decomposition which provides a different way of understanding of unfairness in a causal inference model. Other work focuses on the causal relationship between sensitive attributes and proxies in fair classification (Kilbertus et al., 2017).

Kallus and Zhou (2018) explore the idea of learning from biased data, making the point that a “fair” predictor learned on biased data may not be fair under certain forms of distributional shift, while not touching on causal ideas. Some conceptually similar work has looked at the “selective labels” problem (Lakkaraju et al., 2018; De-Arteaga et al., 2018), where only a biased selection of the data has labels available. There has also been related work on feedback loops in fairness, and the idea that past decisions can affect future ones, in the predictive policing (Lum and Isaac, 2016; Ensign et al., 2018) and recommender systems (Hashimoto et al., 2018) contexts, for example. Barabas et al. (2018) advocate for understanding many problems of fair prediction as ones of intervention instead. Another variational autoencoder-based fairness model is proposed in Louizos et al. (2016), but with the goal of fair representation learning, rather than causal modelling. Dwork et al. (2012) originated the term “fairness through awareness”, and argued that the sensitive attribute needed to be given a place of privilege in modelling in order to reduce unfairness of outcomes.

Experiments

In this section we compare various methods for causal effect estimation. The three effects we are interested in are

A→TA\rightarrow T, the causal effect of AA on TT:

A→YA\rightarrow Y, the causal effect of AA on YY:

T→YT\rightarrow Y, the causal effect of TT on YY:

Note that all three effects are individual-level; that is, they are conditioned on some observed XX (and possibly AA), and then averaged across the dataset.

We evaluate our model using semi-synthetic data. The evaluation of causal models using non-synthetic data is challenging, since a random control trial on the intervention variable is required to validate correctness — this is doubly true in our case, where we are concerned with two different possible interventions. Additionally, while data from random control trials for treatment variables exists (albeit uncommon), conducting a random control trial for a sensitive attribute is usually impossible.

We have adapted the IHDP dataset (Multisite, 1990; Brooks-Gunn et al., 1994)—a standard semi-synthetic causal inference benchmark—for use in the setting of causal effect estimation under a sensitive attribute. The IHDP dataset is from a randomized experiment run by the Infant Health and Development Program (in the US), which ”targeted low-birth-weight, premature infants, and provided the treatment group with both intensive high-quality child care and home visits from a trained provider” (Hill, 2011). Pre-treatment variables were collected from both child (e.g. birth weight, sex) and the mother at time of birth (e.g. age, marital status) and behaviors engaged in during the pregnancy (e.g. smoked cigarettes, drank alcohol), as well as the site of the intervention (where the family resided). We choose our sensitive attribute to be mother’s race, binarized as White and non-White. We follow a similar method for generating outcomes to the Response B surface proposed in Hill (2011). However, our setup differs since we are interested in additionally modelling a sensitive attribute and hidden confounders, so there are three more steps which must be taken. First, we need to generate outcomes YY for each example for T∈{0,1}T\in\{0,1\} under the counterfactual sensitive attribute AA. Second, we need to generate a treatment assignment for each example for the counterfactual value of the sensitive attribute. Finally, we need to remove some data from the observable measurements to act as a hidden confounder, as in Louizos et al. (2017).

We detail our full data generation method in Appendix A. We denote the outcome YY under interventions do(T=t),do(A=ado(T=t),do(A=a) as yT=t,A=ay_{T=t,A=a}. The subroutines in Algorithms 2 and 3 generate all factual and counterfactual outcomes and treatments for each example, one for each possible setting of AA and/or TT. Values of the constants that we use for data generation can be found in Appendix A.

We choose our hidden confounding feature ZZ to be birth weight. In the second (optional) step of data generation, we choose to remove 0, 1, or 2 other features. Especially if we choose features which are highly correlated with the hidden confounder, this has the effect of making the estimation problem more difficult. When removing 0 features, we do nothing. When removing 1 feature, we remove the feature which is most highly correlated with ZZ (head size). When removing 2 features, we remove two features most highly correlated with ZZ (head size & weeks born preterm).

2. Experimental Setup

We run four different models for comparison, including the one we propose. Since we are interested in estimating three different causal effects simultaneously (A→T,A→Y,T→YA\rightarrow T,A\rightarrow Y,T\rightarrow Y), we cannot compare against most standard causal inference benchmark models for treatment effect estimation. The models we test are the following:

Counterfactual MLP (CFMLP): a multilayer perception (MLP) which takes the treatment and sensitive attribute as input, concatenated to XX, and aims to predict outcome. Counterfactual outcomes are calculated by simply flipping the relevant attributes and re-inputting the modified vector to the MLP. A similar auxiliary network learns to predict the treatment from a vector of XX concatenated to AA.

Counterfactual Multiple MLP (CF4MLP): a set of four MLPs — one for each combination of (A,T)∈{0,1}2(A,T)\in\{0,1\}^{2}. Examples are inputted into the appropriate MLP for the factual outcome, and simply inputted into another MLP for the appropriate counterfactual outcome. A similar pair of auxiliary networks predict treatment.

Causal Effect Variational Autoencoder with Sensitive Attribute (CVAE-A, Fig. 1(c)): a model similar to Louizos et al. (2017), but with the simpler inference model we propose. We incorporate a sensitive attribute AA by concatenating XX to AA as input; counterfactuals along AA are taken by flipping AA and re-inputting the modified vector. Counterfactuals along TT are taken as in Louizos et al. (2017).

Fair Causal Effect Variational Autoencoder (FCVAE, Fig. 1(d)): our proposed fair-aware causal model, with AA concatenated to ZZ as confounders. We run two versions: one where AA is used to help with reconstructing XX and inferring ZZ from XX (FCVAE-1), and one where it is not (FCVAE-2). Formally, the inference model and generative model of XX in FCVAE-1 are q(Z∣X,A)q(Z|X,A) and p(X∣Z,A)p(X|Z,A), and in FCVAE-2 are q(Z∣X)q(Z|X) and p(X∣Z)p(X|Z) respectively. In both versions, AA is a confounder of both the treatment and the outcome.

The CFMLP is purely a classification baseline. It learns a mapping from input to output, estimating the conditional distribution P(Y∣X,A,T)P(Y|X,A,T). The CF4MLP shares this goal, but has a more complex architecture—it learns a disjoint set of parameters for each setting of interventions, allowing it to model completely separate generative processes. However, it is still ultimately concerned with supervised prediction. Furthermore, neither of these models is built to consider the impact of hidden confounders.

The CVAE-A is a model for causal inference of outcomes from treatments. Therefore, we should expect it to perform well in estimating T→YT\rightarrow Y. It is also created to model these effects under hidden confounders. Therefore, the difference between CVAE-A and the MLPs will tell us the improvement which comes from appropriate causal modelling rather than classification.

However, the CVAE-A does not consider the sensitive attribute AA as a confounder; rather, it treats it simply as another covariate of XX. So in comparing the FCVAE to the CVAE-A, we observe the improvement that comes from causally modelling the dataset unfairness stemming from a sensitive attribute. In comparing the FCVAE to the MLPs, we observe the full impact of the FCVAE — joint causal modelling of treatments, outcomes, sensitive attributes, and hidden confounders. See Appendix C for experimental details.

3. Results

In this section, we evaluate how well the models from Sec. 6.2 can estimate the three causal effects A→T,A→Y,T→YA\rightarrow T,A\rightarrow Y,T\rightarrow Y. To avoid confusion with the words treatment and outcome, in each of these three causal interactions, we will refer to to the causing variable as the intervention variable, and the affected variable as the result variable. To evaluate how well our model can estimate causal effects, we use PEHE: Precision in Estimation of Heterogeneous Effects (Hill, 2011). This is calculated as: PEHE=\mathdsE[((r1−r0)−(r^1−r^0))2]PEHE=\sqrt[]{\mathds{E}[((r_{1}-r_{0})-(\hat{r}_{1}-\hat{r}_{0}))^{2}]}, where rir_{i} is the ground truth value of result from the intervention ii, and r^i\hat{r}_{i} is our model’s estimate of that quantity. PEHE measures our ability to model both the factual (ground truth) and the counterfactual results.

In Tables 1-3, we show the PEHE for each of the models described in Sec. 6.2, for each causal effect of interest. Each table shows results for a version of the dataset with 0-2 of the most informative features removed (as measured by correlation with the hidden confounder). Therefore, the easiest problem is with zero features removed, the hardest is with two. Note that in IHDP, Y∈\mathdsRY\in\mathds{R}.

Generally, as expected, we observe that the causal models achieve lower PEHE for most estimation problems. Also as expected, we observe that that the PEHE for the more complex estimation problems (A→Y,T→YA\rightarrow Y,T\rightarrow Y) increases as the most useful proxies are removed from the data. We suspect there is less variation in the results for A→TA\rightarrow T since it is a simpler problem: there are no extra confounders (other than ZZ) or mediating factors to consider.

We find that our model (the FCVAE) compares favorably to the other models in this experiment. We see that in general, the fair-aware models (FCVAE-1 and FCVAE-2) have lower PEHE than all other models when estimating the causal effects relating to the sensitive attribute (A→Y,A→TA\rightarrow Y,A\rightarrow T). Furthermore, the FCVAE also performs similarly to the CVAE-A at T→YT\rightarrow Y estimation as well, demonstrating a slight improvement (at least in the more difficult 1, 2 features removed cases).

One interesting note is that FCVAE-1 (where AA is used in reconstruction of XX and in inference of ZZ) and FCVAE-2 seem to perform similarly, with FCVAE-2 being slightly better, if anything. This may seem surprising at first, since one might imagine that using AA would allow the model to learn better representations of XX, particularly for the purpose of doing counterfactual inference across AA.

To explore this further, we examine in table 4 the latent representations ZZ learned by each model in terms of their encoder mutual information between ZZ and XX, which is calculated as KL(q(Z∣X)∣∣p(Z))KL(q(Z|X)||p(Z)), the KL-divergence from the encoder posterior to the prior. This quantity is roughly the same for both versions of the FCVAE, implying that the inference network q(Z∣⋅)q(Z|\cdot) does not leverage the additional information provided by AA in its latent code ZZ. This is in fact sensible because FCVAE has access to AA as an observed confounder in modeling the structural equations. We also noticed that CVAE contains about one bit of extra information in its latent code, implying some degree of success in capturing relevant information about AA in ZZ. But if CVAE models all confounders during inference, why does it underperform relative to FCVAE estimating the downstream causal effects, especially A→YA\rightarrow Y? By making explicit the role of AA as confounder, we hypothesize that FCVAE can learn the interventional distributions with respect to AA (e.g., p(Y∣T,do(A=a),Z))p(Y|T,do(A=a),Z)) rather than the conditional distributions of CVAE (e.g., p(Y∣T,Z(A)))p(Y|T,Z(A))); we suspect that the gating mechanism of the TARNet implementation of the structural equations to be important in this regard.

3.2. Learning a Treatment Policy

The next natural question is: how does estimating these causal effects contribute to a fair decision-making policy? We examine two dimensions of this. We define a policy π:X,A→T\pi:X,A\rightarrow T as a function which maps inputs (features and sensitive attribute) to treatments. We suppose the goal is to assign treatments TT using a policy T^=π(X,A)\hat{T}=\pi(X,A) that maximizes its expected value V(π)V(\pi), defined here as the expected outcome YY it achieves over the data, i.e. V(π)=\mathdsEx,a[Y∣do(T=π(x,a)),A=a,X=x]V(\pi)=\mathds{E}_{x,a}[Y|do(T=\pi(x,a)),A=a,X=x]. For example, we could imagine the treatments to be various medications, and the outcome to be some health indicator (e.g. number of months survived post-treatment).

We can derive a policy from an outcome prediction model simply by outputting the predicted argmax value over treatments, i.e. π(x,a)=argmax⁡t∈T\mathdsEY^[Y^∣do(T=t),A=a,X=x]\pi(x,a)=\operatorname*{argmax}_{t\in T}\mathds{E}_{\hat{Y}}[\hat{Y}|do(T=t),A=a,X=x], where Y^\hat{Y} is the model’s prediction of the true outcome YY. The optimal policy π⋆(x,a)=argmax⁡t∈T\mathdsEY[Y∣do(T=t),A=a,X=x]\pi^{\star}(x,a)=\operatorname*{argmax}_{t\in T}\mathds{E}_{Y}[Y|do(T=t),A=a,X=x] takes the argmax over ground truth outcomes every time.

First, we look at the mean regret of the policy π\pi, which is the difference between its achieved value and the the value of the optimal policy: R(π)=V(π⋆)−V(π)R(\pi)=V(\pi^{\star})-V(\pi). We note that in general, a policy’s regret is not easy to compute or bound without assumptions on the outcome distribution in the data. In Table 5, we display the expected regret values for the learned policies. We observe that the fair-aware model achieves lower regret than the unaware causal model, and much lower regret than the non-causal models, for both the easier and more difficult settings of the IHDP data.

Next, we attempt to measure the policy’s fairness. Most fairness metrics are designed for evaluating classification, not for intervention. However, Chen et al. (2018) explore an idea which is easily adjusted to the interventional setting: that an algorithm is unfair if it is much less accurate on one subgroup. Here, we adapt this notion to evaluate treatment policy fairness.

For any xx, let us say the policy π\pi is accurate if it chooses the treatment which in fact yields the best outcome for that individual; i.e. if π(x,a)=π⋆(x,a)\pi(x,a)=\pi^{\star}(x,a). We can define the accuracy of the policy Acc(π)=\mathdsEx,a[\mathds1(π(x,a)=π⋆(x,a))]Acc(\pi)=\mathds{E}_{x,a}[\mathds{1}(\pi(x,a)=\pi^{\star}(x,a))], where \mathds1\mathds{1} is an indicator function. We can define the subgroup accuracy AccαAcc_{\alpha} as accuracy calculated while conditioning (not intervening) on a particular value α\alpha of AA: Accα(π)=\mathdsEx∣A=α[\mathds1(π(x,α)=π⋆(x,α))]Acc_{\alpha}(\pi)=\mathds{E}_{x|A=\alpha}[\mathds{1}(\pi(x,\alpha)=\pi^{\star}(x,\alpha))]. We condition rather than intervene on AA here since we are interested in measuring the impact of the policy on real, existing populations, rather than hypothetical ones. Finally, to evaluate the fairness of the policy, we can look at the accuracy gap: ∣Acc1(π)−Acc0(π)∣|Acc_{1}(\pi)-Acc_{0}(\pi)|. If this is high, the model is more unfair, since the policy has been more successful at modelling one group than the other, and is much more consistently choosing the correct treatment for individuals in that group.

In Table 6 we display the accuracy gaps for our models and baselines on the IHDP dataset. We observe that the FCVAE achieves a smaller accuracy gap than those which do not consider the effect of the sensitive attribute. This is an encouraging sign that by understanding the confounding influence of sensitive attributes in biasing historical datasets, we can learn treatment policies which are more accurate for all subgroups of the data.

Discussion

In this paper, we proposed a causally-motivated model for learning from potentially biased data. We emphasize the importance of modeling the potential confounders of historical datasets: we model the sensitive attribute as an observed confounder contributing to dataset bias, and leverage deep latent variable models to approximately infer other hidden confounders.

In Sec. 6.3.2, we demonstrated how to use our model to learn a simple treatment policy from data which assigns treatments more accurately and fairly than several causal and non-causal baselines. Looking forward, the estimation of sensitive attribute causal effects suggests several compelling new research directions, which we non-exhaustively discuss here:

Counterfactual Fairness: Our model learns outcomes for counterfactual values of both TT and AA. This means we could choose to implement a policy where we assess everyone under the same value a′a^{\prime}, by assigning treatments to all individuals, no matter their original value aa of AA, based on the inferred outcome distribution P(Y∣do(A=a′),X,T)P(Y|do(A=a^{\prime}),X,T). Such a policy respects the definition of counterfactual fairness proposed by Kusner et al. (2017), which requires invariance to counterfactuals in AA at the individual level.

Path-Specific Effects: Our model allows us to decompose A→YA\rightarrow Y into direct and indirect effects through mediation analysis of TT (Robins and Greenland, 1992). By estimating this decomposition, we could learn a policy which respects path-specific fairness, as proposed by Nabi and Shpitser (2018).

Analyzing Historical Bias: Estimating causal effects between AA, TT, and YY permits for the analysis and comparison of bias in historical datasets. For instance, the effect A→TA\rightarrow T is a measure of bias in a historical policy, and the effect A→YA\rightarrow Y is a measure of bias in whatever system historically generated the outcome. This could serve as the basis of a bias auditing technique for data scientists.

Data Augmentation: The absence of data (especially not-at-random) has strong implications for downstream modeling in both fairness (Kallus and Zhou, 2018) and causal inference (Rubin, 1976). Our model outputs counterfactual outcomes for both AA and TT, which could be used for fair missing data imputation (Van Buuren, 2018; Sterne et al., 2009). This could in turn enable the application of simpler methods like supervised learning to interventional problems.

Fair Policies Under Constraints: In this paper, we consider an approach to fairness where understanding dataset bias is paramount, rather than the more common fairness-accuracy constraint-based tradeoff (Hardt et al., 2016; Menon and Williamson, 2018). However, in some domains we may be interested in policies which satisfy a fairness constraint (e.g., the same distribution of treatments are given to each group). Estimating the underlying causal effects would be useful for constrained policy learning.

Incorporating Prior Knowledge: Graphical models (both probabilistic and SCM) permit the specification of prior knowledge when modeling data, and provide a framework for inference that balances these beliefs with evidence from the data. This is a powerful fairness idea—we may believe a priori that a dataset should look a certain way if not for some bias. In the context of a fair machine learning pipeline that considers many datasets, this relates to the AutoML task of learning distributions over datasets that share global parameters (Edwards and Storkey, 2017).

In automated decision making, the focus on intervention over classification (Barabas et al., 2018) suggests the more equitable deployment of machine learning when only biased data are available, but also raises significant technical challenges. We believe causal modeling to be an invaluable tool in addressing these challenges, and hope that this paper contributes to the discussion around how best to understand and make predictions from existing datasets without replicating existing biases.

References

Appendix A Data Generation

We detail our dataset generation progess in Algorithm 1. We denote the outcome YY under interventions do(T=t),do(A=ado(T=t),do(A=a) as yT=t,A=ay_{T=t,A=a}. The subroutines in Algorithms 2 and 3 generate all factual and counterfactual outcomes and treatments for each example, one for each possible setting of AA and/or TT. In Algorithm 1, we have several undefined constant variables. We use the following values for those variables:

for continuous variables, βi∼Cat([0,.1,.2,.3,.4],[.5,.125,.125,.125,.125])\beta_{i}\sim Cat([0,.1,.2,.3,.4],[.5,.125,.125,.125,.125])

for binary variables βi∼Cat([0,.1,.2,.3,.4],[.6,.1,.1,.1,.1])\beta_{i}\sim Cat([0,.1,.2,.3,.4],[.6,.1,.1,.1,.1])

for ZZ, βi=Cat([.4,.6],[.5,.5])\beta_{i}=Cat([.4,.6],[.5,.5])

where Cat(x,p)Cat(x,p) selects values from xx according to the array of probabilities pp.

α0,α1,ζ=0.7,0.4,0.1\alpha_{0},\alpha_{1},\zeta=0.7,0.4,0.1

We also use the function ClipClip, which is defined as:

Appendix B Identifiability of Causal Effects

Here we show that if we can successfully recover the joint distribution P(Z,A,X,T,Y)P(Z,A,X,T,Y), we can recover all three treatment effects we are interested in:

The effect of TT on YY (T→YT\rightarrow Y): \mathdsE(Y∣do(T=1),X,A)−\mathdsE(Y∣do(T=0),X,A)\mathds{E}(Y|do(T=1),X,A)-\mathds{E}(Y|do(T=0),X,A)

The effect of AA on TT (A→TA\rightarrow T): \mathdsE(T∣do(A=1),X)−\mathdsE(T∣do(A=0),X)\mathds{E}(T|do(A=1),X)-\mathds{E}(T|do(A=0),X)

The effect of AA on YY (A→YA\rightarrow Y): \mathdsE(Y∣do(A=1),X)−\mathdsE(Y∣do(A=0),X)\mathds{E}(Y|do(A=1),X)-\mathds{E}(Y|do(A=0),X)

Our proof will closely follow Louizos et al. (2017). For each effects, it will suffice to show that we can recover the first term on the right-hand side of each expression. (The argument for the second term is the same). We will show only the proof for the effect of TT on YY — the others are very similar.

Theorem. Given the causal model in Fig. 1(d), if we recover the joint distribution P(Z,A,X,T,Y)P(Z,A,X,T,Y), then we can recover \mathdsE(Y∣do(T=1),X,A)\mathds{E}(Y|do(T=1),X,A).

By the dodo-calculus, we can reduce further:

If we know the joint distribution P(Z,A,X,T,Y)P(Z,A,X,T,Y), we can identify the value of each term in this expression; hence we can identify the value of the whole expression. ∎

Appendix C Experimental details

We run each model on 500 distinct data seed/model seed pairs, in order to get robust confidence estimates on the error of each model. We parametrize each function in our causal model with a neural network. Our networks between XX and ZZ have a single hidden layer of 20 hidden units. The size of the learned hidden confounder ZZ was 10 units. Each of our TARNets consist of a network outputting a shared representation, and two networks making predictions from that representation. Each of these network have 1 hidden layer with 100 hidden units. The size of the shared representation in the TARNets was 20 units. For simplicity, we set gXσ=1g_{X}^{\sigma}=1 for all experiments (but not gZσg_{Z}^{\sigma})—this amounts to assuming unit variance for the data XX, a sensible assumption because they are normalized during pre-processing. We used ELU non-linear activations (Clevert et al., 2015). We trained our model with ADAM (Kingma and Ba, 2015) with a learning rate of 0.001, calculating the ELBO on a validation set and stopping training after 10 consecutive epochs without improvement. We sample 10 times from the posterior q(Z∣⋅)q(Z|\cdot) at both training and test time for each input example. At training time we compute the average ELBO across the ten samples, while at test time we use the average prediction.