Deep Counterfactual Networks with Propensity-Dropout
Ahmed M. Alaa, Michael Weisz, Mihaela van der Schaar
Introduction
The problem of inferring individualized treatment effects from observational datasets is a fundamental problem in many domains such as precision medicine (Shalit et al., 2017), econometrics (Abadie & Imbens, 2016), social sciences (Athey & Imbens, 2016), and computational advertising (Bottou et al., 2013). A lot of attention has been recently devoted to this problem due to the recent availability of electronic health record (EHR) data in most of the hospitals in the US (Charles et al., 2015), which paved the way for using machine learning to estimate the individual-level causal effects of treatments from observational EHR data as an alternative to the expensive clinical trials.
A typical observational dataset comprises a subject’s features, a treatment assignment indicator (i.e. whether the subject received the treatment), and a “factual outcome” corresponding to the subject’s response. Estimating the effect of a treatment for any given subject requires inferring her “counterfactual outcome”, i.e. her response had she experienced a different treatment assignment. Classical works have focused on estimating “average” treatment effects through variants of propensity score matching (Rubin, 2011; Austin, 2011; Abadie & Imbens, 2016; Rosenbaum & Rubin, 1983; Rubin, 1973). More recent works tackled the problem of estimating “individualized” treatment effects using representation learning (Johansson et al., 2016; Shalit et al., 2017), Bayesian inference (Hill, 2012), and standard supervised learning (Wager & Athey, 2015).
In this paper, we propose a novel approach for individual-level causal inference that casts the problem in a multitask learning framework. In particular, we model a subject’s potential (factual and counterfactual) outcomes using a deep multitask network with a set of layers that are shared across the two outcomes, and a set of idiosyncratic layers for each outcome (see Fig. 1). We handle selection bias in the observational data via a novel propensity-dropout regularization scheme, in which the network is thinned for every subject via a dropout probability that depends on the subject’s propensity score. Our model can provide individualized measures of uncertainty in the estimated treatment effect by applying Monte Carlo propensity-dropout at inference time (Gal & Ghahramani, 2016).
Learning is carried out through an alternate training approach in which we divided the observational data into a “treated batch” and a “control batch”, and then update the weights of the shared and idiosyncratic layers for each batch separately in an alternating fashion. We conclude the paper by conducting a set of experiments on data based on a real-world observational study showing that our algorithm outperforms the state-of-the-art.
Problem Formulation
Model Description
We propose a neural network model for estimating the individualized treatment effect by learning a shared representation for the two potential outcomes. Our model, depicted in Fig. 1, comprises a propensity network (right) and a potential outcomes network (left). The propensity network is a standard feed-forward network with layers and hidden units in the layer, and is trained separately to estimate the propensity score via the samples in . The potential outcomes network is a multitask network (Collobert & Weston, 2008) that comprises shared layers (with hidden units in the shared layer), and idiosyncratic layers (with hidden units in the layer) for potential outcome .
2 Propensity-Dropout
In order to ameliorate the impact of selection bias, we use the outputs of the propensity network to regularize the potential outcomes network. We do so through a dropout scheme that we call propensity-dropout. In propensity-dropout, the dropout procedure is applied in such a way that it assigns “simple models” to subjects with very high or very low propensity scores ( close to 0 or 1), and more “complex models” to subjects with balanced propensity scores ( close to 0.5). That is, we use a different dropout probability for each training example depending on the associated score: the dropout probability is higher for subjects with features that belong in a region of poor treatment assignment overlap in the feature space. We implement the propensity-dropout by using the following formula for the dropout probability:
where and are the weight matrices for the layer of the propensity, shared and idiosyncratic layers, respectively, and are dropout masking vectors, and is any activation function.
3 Training the Model
We train the network in alternating phases, where in each phase, we either use the treated batch or the control batch to update the weights of the shared and idiosyncratic layers. As shown in Algorithm 1, we run this process over a course of epochs; the shared layers are updated in all epochs, wheres only one set of idiosyncratic layers is updated in any given epoch. Dropout is applied as explained in the previous Subsection with . As visualized in Fig. 2, we can think of alternate training as deterministically dropping all units of one of the idiosyncratic layers in every epoch.
We update the weights of all layers in each epoch using the Adam optimizer with default settings and Xavier initialization (Kingma & Ba, 2014).
Experiments
The ground truth counterfactual outcomes are never available in an observational dataset, which hinders the evaluation of causal inference algorithms on real-world data. Following (Hill, 2012; Johansson et al., 2016), we adopt a semi-synthetic experimental setup in which the covariates and treatment assignments are real but outcomes are simulated. We conduct our experiments using the Infant Health and Development Program (IHDP) dataset introduced in (Hill, 2012). (The IHDP is a social program applied to premature infants aiming at enhancing their IQ scores at the age of three.) The dataset comprises 747 subjects (139 treated and 608 control), with 25 covariates associated with each subject. Outcomes are simulated based on the data generation process designated as the “Response Surface B” setting in (Hill, 2012).
We evaluate the performance of a DCN-PD model with (a total of 4 layers), and with 200 hidden units in all layers (ReLU activation), in terms of the mean squared error (MSE) of the estimated treatment effect. We divide the IHDP data into a training set (80) and an out-of-sample testing set (20), and then evaluate the MSE on the testing sample in 100 different experiments, were in each experiment a new realization for the outcomes is drawn from the data generation model in (Hill, 2012). (We implemented the DCN-PD model in a Tensorflow environment.) The propensity network is implemented as a standard 2-layer feed-forward network with 25 hidden layers, and is trained using the Adam optimizer.
The marginal benefits conferred by the propensity-dropout regularization scheme are illustrated in Fig. 3, which depicts box plots for the MSEs achieved by the DCN-PD model, and two DCN models with conventional dropout (dropout probabilities of 0.2 and 0.5 for all layers and all training examples). As we can see in Fig. 3, the DCN-PD model offers a significant improvement over the two DCN models for which the dropout probabilities are uniform over all the training examples. This result implies that the DCN-PD model generalizes better to the true feature distribution when trained with a biased dataset as compared to DCN with regular dropout, which suggests that propensity-dropout is a good regularizer for causal inference.
In order to assess the marginal performance gain achieved by the proposed multitask model when combined with the propensity-dropout scheme, we compare the performance of DCN-PD with other state-of-the-art models in Table 1. In particular, we compare the MSE (averaged over 100 experiments) achieved by the DCN-PD with those achieved by nearest neighbor matching (-NN), Causal Forests with double-sample trees (Wager & Athey, 2015), Bayesian Additive Regression Trees (BART) (Chipman et al., 2010; Hill, 2012), and Balancing neural networks (BNN) (Johansson et al., 2016). (For BNNs, we use 4 layers with 200 hidden units per layer to ensure a fair comparison.) We also provide a direct comparison with a standard single-output feed-forward neural network (with 4-layers and 200 hidden units per layer) that treats the treatment assignment as an input feature (NN-4), and a DCN with a standard dropout with a probability of 0.2. As we can see in Table 1, DCN-PD outperforms all the other models, with the BNN model being the most competitive. (BNN is a strong benchmark as it handles the selection by learning a “balanced representation” for the input features (Johansson et al., 2016).) DCN-PDs significantly outperforms the NN-4 benchmark, which suggests that the multitask modeling framework is a more appropriate conception of causal inference compared to direct modeling by assuming that the treatment assignment is an input feature.