Null-sampling for Interpretable and Fair Representations

Thomas Kehrenberg, Myles Bartlett, Oliver Thomas, Novi Quadrianto

Introduction

Without due consideration for the data collection process, machine learning algorithms can exacerbate biases, or even introduce new ones if proper control is not exerted over their learning . While most of these issues can be solved by controlling and curating data collection in a fairness-conscious fashion, doing so is not always an option, such as when working with historical data. Efforts to address this problem algorithmically have been centred on developing statistical definitions of fairness and learning models that satisfy these definitions. One popular definition of fairness used to guide the training of fair classifiers, for example, is demographic parity, stating that positive outcome rates should be equalised (or invariant) across protected groups.

In the typical setup, we have an input x\bm{x}, a sensitive attribute ss that represents some non-admissible information like gender and a class label yy which is the prediction target. The idea of fair representation learning is then to transform the input x\bm{x} to a representation z\bm{z} which is invariant to ss. Thus, learning from z\bm{z} will not introduce a forbidden dependence on ss. A good fair representation is one that preserves most of the information from x\bm{x} while satisfying the aforementioned constraints.

As unlabelled data is much more freely available than labelled data, it is of interest to learn the representation in an unsupervised manner. This will allow us to draw on a much more diverse pool of data to learn from. While annotations for yy are often hard to come by (and often noisy ), annotations for the sensitive attribute ss are usually less so, as ss can often be obtained from demographic information provided by census data. We thus consider the setting where the representation is learned from data that is only labelled with ss and not yy. This is in contrast to most other representation learning methods. We call the set used to learn the representation the representative set, because its distribution is meant to match the distribution of the deployment setting (and is thus representative).

Once we have learnt the mapping from x\bm{x} to z\bm{z}, we can transform the training set which, in contrast to the representative set, has the yy labels (and ss labels). In order to make our method more widely applicable, we consider an aggravated fairness problem in which the training set contains a strong spurious correlation between ss and yy, which makes it impossible to learn from it a representation which is invariant to ss but not invariant to yy. Non-invariance to yy is important in order to be able to predict yy. The training set thus does not match the deployment setting, thereby rendering the representative set essential for learning the right invariance. From hereon, we will use the terms spurious and sensitive interchangeably, depending on the context, to refer to an attribute of the data we seek invariance to. We can draw a connection between learning in the presence of spurious correlations and what call residual unfairness. Consider the Stop, Question and Frisk (SQF) dataset for example: the data was collected in New York City, but the demographics of the recorded cases do not represent the true demographics of NYC well. The demographic attributes of the recorded individuals might correlate so strongly with the prediction target that the two are nearly indistinguishable. This is the scenario that we are investigating: ss and yy are so closely correlated in the labelled dataset that they cannot be distinguished, but the learning of ss is favoured due to being the “path of least resistance”. The deployment setting (i.e. the test set) does not possess this strong correlation and thus a naïve approach will lead to very unfair predictions. In this case, a disentangled representation is insufficient; the representation needs to be explicitly invariant solely with respect to ss. In our approach, we make use of the (partially labelled) representative set to learn this invariant representation.

While there is a substantial body of literature devoted to the problems of fair representation-learning, exactly how the invariance in question is achieved is often overlooked. When critical decisions, such as who should receive bail or be released from jail, are being deferred to an automated decision making system, it is critical that people be able to trust the logic of the model underlying it, whether it be via semantic or visual explanations. We build on the work of and learn a decomposition (f−1:Zs×Z¬s→Xf^{-1}:Z_{s}\times Z_{\neg s}\rightarrow X) of the data domain (XX) into independent subspaces invariant to ss (Z¬sZ_{\neg s}) and indicative of ss (ZsZ_{s}), which lends an interpretability that is absent from most representation-learning methods. While model interpretability has no strict definition , we follow the intuition of – a simple relationship to something we can understand, a definition which representations in the data domain naturally fulfil.

Whether as a result of the aforementioned sampling bias or simply because the features necessarily co-occur, it is not rare for features to correlate with one another in real-world datasets. Lipstick and gender for example, are two attributes that we expect to be highly correlated and to enforce invariance to gender can implicitly enforce invariance to makeup. This is arguably the desired behaviour. However, unforeseen biases in the data may engender cases which are less justifiable. By baking interpretability into our model (by having representations in the data domain), though we still have no better control over what is learned, we can at least diagnose such pathologies.

To render our representations interpretable, we rely on a simple transformation we call null-sampling to map invariant representations in the data domain. Previous approaches to fair representation learning predominantly rely upon autoencoder models to jointly minimise reconstruction loss and invariance. We discuss first how this can be done with such a model that we refer to as cVAE (conditional VAE), before arguing that the bijectivity of invertible neural networks (INNs) makes them better suited to this task. We refer to the variant of our method based on these as cFlow (conditional Flow). INNs have several properties that make them appealing for unsupervised representation learning. The focus of our approach is on creating invariant representations that preserve the non-sensitive information maximally, with only knowledge of ss and not of the target yy, while at the same time having the ability to easily probe what has been learnt.

Our contribution is thus two-fold: 1) We propose a simple approach to generating representations that are invariant to a feature ss, while having the benefit of interpretability that comes with being in the data domain. We call our model NIFR (Null-sampling for Interpretable and Fair Representations). 2) We explore a setting where the labelled training set suffers from varying levels of sampling bias, demonstrating an approach based on transferring information from a more diverse representative set, with guarantees of the non-spurious information being preserved.

Background

Learning fair representations. Given a sensitive attribute ss (for example, gender or race) and inputs x\bm{x}, a fair representation z\bm{z} of x\bm{x} is then one for which z⊥s\bm{z}\perp s holds, while ideally also being predictive of the class label yy. was the first to propose the learning of fair representations which allow for transfer to new classification tasks. More recent methods are often based on variational autoencoders (VAEs) . The achieved fairness of the representation can be measured with various fairness metrics. These measure, however, usually how fair the predictions of a classifier are and not how fair a representation is.

The appropriate measure of fairness for a given task is domain-specific and there is often not a universally accepted measure. However, Demographic Parity is the most widely used . Demographic Parity demands y^⊥s\hat{y}\perp s where y^\hat{y} refers to the predictions of the classifier. In the context of fair representations, we measure the Demographic Parity of a downstream classifier, f(⋅)f(\cdot), which is trained on the representation zz i.e. f:Z→Y^f:Z\to\hat{Y}.

A core principle of all fairness methods is the accuracy-fairness trade-off. As previously stated, the fair representation should be invariant to ss (→\to fairness) but still be predictive of yy (→\to accuracy). These desiderata cannot, in general, be simultaneously satisfied if ss and yy are correlated.

The majority of existing methods for fair representations also make use of yy labels during training, in order to ensure that z\bm{z} remains predictive of yy. This aspect can, in theory, be removed from the methods, but then there is no guarantee that information about yy is preserved .

Learning fair, transferrable representations. In addition to producing fair representations, want to ensure the representations are transferrable. Here, an adversary is used to remove sensitive information from a representation zz. Auxiliary prediction and reconstruction networks, to predict class label yy and reconstruct the input xx respectively, are trained on top of zz, with ss being ancillary input to the reconstruction.

Also related is who employ a FactorVAE regularised for fairness. The idea is to learn a representation that is both disentangled and invariant to multiple sensitive attributes. This factorisation makes the latent space easily manipulable such that the different subspaces can be freely removed and composed at test time. Zeroing out the dimensions or replacing them with independent noise imparts invariance to the corresponding sensitive attribute. This method closely resembles ours when we use an invertible encoder. However, the emphasis of our approach is on interpretability, information-preservation, and coping with sampling bias - especially extreme cases where ∣ supp(Str×Ytr)∣<∣ supp(Ste×Yte)∣|\,\textrm{supp}(S_{tr}\times Y_{tr})|<|\,\textrm{supp}(S_{te}\times Y_{te})|.

Attempts were made by prior to this work to learn fair representations in the data domain in order to make it interpretable and transferable. In their work, the input is assumed to be additively decomposable in the feature space into a fair and unfair component, which together can be used by the decoder to recover the original input. This allows us to examine representations in a human-interpretable space and confirm that the model is not learning a relationship reliant on a sensitive attribute. Though a first step in this direction, we believe such a linear decomposition is not sufficiently expressive to fully capture the relationship between the sensitive and non-sensitive attributes. Our approach allows for the modelling of more complex relationships.

Learning in the presence of spurious correlations. Strong spurious correlations make the task of learning a robust classifier challenging: the classifier may learn to exploit correlations unrelated to the true causal relationship between the features and label, and thereby fail to generalise to novel settings. This problem was recently tackled by who apply a penalty based on the mutual information between the feature embedding and the spurious variable. While the method is effective under mild biasing, we show experimentally that it is not robust to the range of settings we consider.

Jacobsen et al. explore the vulnerability of traditional neural networks to spurious variables – e.g., textures, in the case of ImageNet – and propose a INN-based solution akin to ours. The INN’s encoding is split such that one partition, zbz_{b} is encouraged to be predictive of the spurious variable while the other serves as the logits for classification of the semantic label. Information related to the nuisance variable is “pulled out” of the logits as a result of maximising log⁡p(s∣zn)\log p(s|z_{n}). This specific approach, however, is incompatible with the settings we consider, due to its requirement that both ss and yy be available at training time.

Viewing the problem from a causal perspective, develop a variant of empirical risk minimisation called invariant risk minimisation (IRM). The goal of IRM is to train a predictor that generalises across a large set of unseen environments; because variables with spurious correlations do not represent a stable causal mechanism, the predictor learns to be invariant to them. IRM assumes that the training data is not iid but is partitioned into distinct environments, e∈Ee\in E. The optimal predictor is then defined as the minimiser of the sum of the empirical risk ReR_{e} over this set. In contrast, we assume possession of only a single source of labelled, albeit spuriously-correlated, data, but that we have a second source of data that is free of spurious correlations, with the benefit being that it only needs to be labelled with respect to ss.

Interpretable Invariances by Null-Sampling

We assume we are given inputs x∈X\bm{x}\in\mathcal{X} and corresponding labels y∈Yy\in\mathcal{Y}. Furthermore, there is some spurious variable s∈Ss\in\mathcal{S} associated with each input x\bm{x} which we do not want to predict. Let XX, SS and YY be random variables that take on the values x\bm{x}, ss and yy, respectively. The fact that both yy and ss are predictive of x\bm{x} implies that I(X;Y),I(X;S)>0\mathcal{I}(X;Y),\mathcal{I}(X;S)>0, where I(⋅;⋅)\mathcal{I}(\cdot;\cdot) is the mutual information. Note, however, that the conditional entropy is non-zero: H(S∣X)≠0H(S|X)\neq 0, i.e., SS is not completely determined by XX.

The difficulty of this setup emerges in the training set: there is a close correspondence between SS and YY, such that for a model that sees the data through the lens of the loss function, the two are indistinguishable. Furthermore, we assume that this is not the case in the test set, meaning the model cannot rely on shortcuts provided by SS if it is to generalise from the training set.

We call this scenario where we only have access to the labels of a biasedly-sampled subpopulation an aggravated fairness problem. These are not uncommon in the real-world. For instance, in long-feedback systems such as mortgage-approval where the demographics of the subpopulation with observed outcomes is not representative of the subpopulation on which the model has been deployed. In this case, ss has the potential to act as a false (or spurious) indicator of the class label and training a model with such a dataset would limit generalisability. Let (Xtr,Str,Ytr)(X^{\mathit{tr}},S^{\mathit{tr}},Y^{\mathit{tr}}) then be the random variables sampled for the training set and (Xte,Ste,Yte)(X^{\mathit{te}},S^{\mathit{te}},Y^{\mathit{te}}) be the random variables for the test set. The training and test sets thus induce the following inequality for their mutual information: I(Str;Ytr)≫I(Ste;Yte)≈0\mathcal{I}(S^{\mathit{tr}};Y^{\mathit{tr}})\gg\mathcal{I}(S^{\mathit{te}};Y^{\mathit{te}})\approx 0.

Our goal is to learn a representation zu\bm{z}_{u} that is independent of ss and transferable between downstream tasks. Complementary to zu\bm{z}_{u}, we refer to some abstract component of the model that absorbs the unwanted information related to ss as B\mathcal{B}, the realisation of which we define with respect to each of the two models to be described. The requirement for zu\bm{z}_{u} can be expressed via mutual information:

However, for the representation to be useful, we need to capture as much relevant information in the data as possible. Thus, the combined objective function:

where θ\theta refers to the trainable parameters of our model fθf_{\theta} and pθ(x)p_{\theta}(\bm{x}) is the likelihood it assigns to the data.

We optimise this loss in an adversarial fashion by playing a min-max game, in which our encoder acts as the generative component. The adversary is an auxiliary classifier gg, which receives zu\bm{z}_{u} as input and attempts to predict the spurious variable ss. We denote the parameters of the adversary as ϕ\phi; for the parameters of the encoder we use θ\theta, as before. The objective from Eq (2) is then

where Lc\mathcal{L}_{c} is the cross-entropy between the predictions for ss and the provided labels. In practice, this adversarial term is realised with a gradient reversal layer (GRL) between zu\bm{z}_{u} and gg as is common in adversarial approaches .

2 The Disentanglement Dilemma

The objective in Eq (3) balances the two desiderata: predicting yy and being invariant to ss. However, in the training set (Xtr,Str,Ytr)(X^{\mathit{tr}},S^{\mathit{tr}},Y^{\mathit{tr}}), yy and ss are so strongly correlated that removing information about ss inevitably removes information about yy. This strong correlation makes existing methods fail under this setting. In order to even define the right learning goal, we require another source of information that allows us to disentangle ss and yy. For this, we assume the existence of another set of samples that follow a similar distribution to the test set, but whilst the sensitive attribute is available, the class labels are not. In reality, this is not an unreasonable assumption, as, while properly annotated data is scarce, unlabelled data can be obtained in abundance (with demographic information from census data, electoral rolls, etc.). Previous work has also considered treated “unlabelled data” as still having ss labels . We are restricted only in the sense that the spurious correlations we want to sever are indicated in the features. We call this the representative set, consisting of XrepX^{\mathit{rep}} and SrepS^{\mathit{rep}}. It fulfils I(Srep;Yrep)≈0\mathcal{I}(S^{\mathit{rep}};Y^{\mathit{rep}})\approx 0 (or rather, it would, if the class labels YrepY^{\mathit{rep}} were available).

We now summarise the training procedure; an outline for the invertible network model (cFlow) can be seen in Fig. 3a. First, the encoder network ff is trained on (Xrep,SrepX^{\mathit{rep}},S^{\mathit{rep}}), during the first phase. The trained network is then used to encode the training set, taking in x\bm{x} and producing the representation, zu\bm{z}_{u}, decorrelated from the spurious variable. The encoded dataset can then be used to train any off-the-shelf classifier safely, with information about the spurious variable having been absorbed by some auxiliary component B\mathcal{B}. In the case of the conditional VAE (cVAE) model, B\mathcal{B} takes the form of the decoder subnetwork, which reconstructs the data conditional on a one-hot encoding of ss, while for the invertible network B\mathcal{B} is realised as a partition of the feature map z\bm{z} (such that z=[zu,zb]\bm{z}=[\bm{z}_{u},\bm{z}_{b}]), given the bijective constraint. Thus, the classifier cannot take the shortcut of learning ss and instead must learn how to predict yy directly. Obtaining the ss-invariant representations, xu\bm{x}_{u}, in the data domain is simply a matter of replacing the B\mathcal{B} component of the decoder’s input for the cVAE, and zb\bm{z}_{b} for cFlow, with a zero vector of equivalent size. We refer to this procedure used to generate xu\bm{x}_{u} as null-sampling (here, with respect to zb)\bm{z}_{b}).

Null-sampling resembles the annihilation operation described in , however we note that the two serve very different roles. Whereas the annihilation operation serves as a regulariser to prevent trivial solutions (similar to ), null-sampling is used to generate the invariant representations post-training.

3 Conditional Decoding

We first describe a VAE-based model similar to that proposed in , before highlighting some of its shortcomings that motivate the choice of an invertible representation learner.

The model takes the form of a class conditional β\beta-VAE , in which the decoder is conditioned on the spurious attribute. We use θenc,θdec∈θ\theta_{enc},\theta_{dec}\in\theta to denote the parameters of the encoder and decoder sub-networks, respectively. Concretely, the encoder component performs the mapping x→zux\rightarrow{\bm{z}_{u}}, while B\mathcal{B} is instantiated as the decoder, B≔pθdec(x∣zu,s)\mathcal{B}\coloneqq p_{\theta_{dec}}(x|z_{u},s), which takes in a concatenation of the learned non-spurious latent vector zu\bm{z}_{u} and a one-hot encoding of the spurious label ss to produce a reconstruction of the input x^\hat{x}. Conditioning on a one-hot encoding of ss, rather than a single value, as done in is the key to visualising invariant representations in the data domain. If I(zu;s)\mathcal{I}(z_{u};s) is properly minimised, the decoder can only derive its information about ss from the label, thereby freeing up zu\bm{z}_{u} from encoding the unwanted information while still allowing for reconstruction of the input. Thus, by feeding a zero-vector to the decoder we achieve x^⊥s\hat{x}\perp s. The full learning objective for the cVAE is given as

where β\beta is a hyperparameter that determines the trade-off between reconstruction accuracy and independence constraints, and p(zu)p(\bm{z}_{u}) is the prior imposed on the variational posterior. For all our experiments, p(zu)p(\bm{z}_{u}) is realised as an Isotropic Gaussian. Fig. 3b summarises the procedure as a diagram.

While we show this setup can indeed work for simple problems, as before us have, we show that it lacks scalability due to disagreement between the components of the loss. Since information about ss is only available to the decoder as a binary encoding, if the relationship between ss and xx is highly non-linear and cannot be summarised by a simple on/off mechanism, as is the case if ss is an attribute such as gender, off-loading information to the decoder by conditioning is no longer possible. As a result, zu\bm{z}_{u} is forced to carry information about ss in order to minimise the reconstruction error.

The obvious solution to this is to allow the encoder to store information about ss in a partition of the latent space as in . However, we question whether an autoencoder is the best choice for this setup, with the view that an invertible model is the better tool for the task. Using an invertible model has several guarantees, namely complete information-preservation and freedom from a reconstruction loss, the importance of which we elaborate on below.

4 Conditional Flow

Invertible Neural Networks. Invertible neural networks are a class of neural network architecture characterised by a bijective mapping between their inputs and output . The transformations are designed such that their inverses and Jacobians are efficiently computable. These flow-based models permit exact likelihood estimation through the warping of a base density with a series of invertible transformations and computing the resulting, highly multi-modal, but still normalised, density, using the change of variable theorem:

where hih_{i} refers to the outputs of the layers of the network and p(z)p(z) is the base density, specifically an Isotropic Gaussian in our case. Training of the invertible neural network is then reduced to maximising log⁡p(x)\log p(x) over the training set, i.e. maximising the probability the network assigns to samples in the training set.

The Benefits of Bijectivity. Using an invertible network to generate our encoding, zu\bm{z}_{u}, carries a number of advantages over other approaches. Ordinarily, the main benefit of flow-based models is that they permit exact density estimation. However, since we are not interested in sampling from the model’s distribution, in our case the likelihood term serves as a regulariser, as it does for . Critically, this forces the mean of each latent dimension to zero enabling null-sampling. The invertible property of the network guarantees the preservation of all information relevant to yy which is independent of ss, regardless of how it is allocated in the output space. Secondly, we conjecture that the encodings are more robust to out-of-distribution data. Whereas an autoencoder could map a previously seen input and a previously unseen input to the same representation, an invertible network sidesteps this due to the network’s bijective property, ensuring all relevant information is stored somewhere. This opens up the possibility of transfer learning between datasets with a similar manifestation of ss, as we demonstrate in the Appendix LABEL:sec:transfer-learning.

Under our framework, the invertible network ff maps the inputs x\bm{x} to a representation zu\bm{z}_{u}: f(x)=zf(\bm{x})=\bm{z}. We interpret the embedding z\bm{z} as being the concatenation of two smaller embeddings: z=[zu,zb]\bm{z}=[\bm{z}_{u},\bm{z}_{b}]. The dimensionality of zb\bm{z}_{b}, and zu\bm{z}_{u}, by complement, is a free parameter (see Appendix LABEL:ssec:partition-size-matters for tuning strategies). As ff is invertible, x\bm{x} can be recovered like so:

where zb\bm{z}_{b} is required for equality of the output dimension and input dimension to satisfy the bijectivity of the network – we cannot output zu\bm{z}_{u} alone, but have to output zb\bm{z}_{b} as well. In order to generate the pre-image of zu\bm{z}_{u}, we perform null-sampling with respect to zb\bm{z}_{b} by zeroing-out the elements of zb\bm{z}_{b} (such that xu=f−1([zu,0→])\bm{x}_{u}=f^{-1}([\bm{z}_{u},\stackrel{{\scriptstyle\rightarrow}}{{0}}])), i.e. setting them to the mean of the prior density, N(z;0,I)\mathcal{N}(z;0,I).

How can we be sure that zu\bm{z}_{u} contains enough information about yy? The importance of the invertible architecture bears out from this consideration. As long as zb\bm{z}_{b} does not contain the information about yy, zu\bm{z}_{u} necessarily must. We can raise or lower the information capacity of zb\bm{z}_{b} by adjusting its size; this should be set to the smallest size sufficient to capture all information about ss, so as not to sacrifice class-relevant information. Appendix LABEL:ssec:size-of-zb explores the effects of the size further.

Experiments

We present experiments to demonstrate that the null-sampled representations are in fact invariant to ss while still allowing a classifier to predict yy from them. We run our cVAE and cFlow models on the coloured MNIST (cMNIST) and CelebA dataset, which we artificially bias, first describing the sampling procedure we follow to do so for non-synthetic datasets. As baselines we have the model of (Ln2L) and the same CNN used to evaluate the cFlow and cVAE models but with the unmodified images as input (CNN). For the cFlow model we adopt a Glow-like architecture , while both subnetworks of the cVAE model comprise gated convolutions , where the encoding size is 256256. For cMNIST, we construct the Ln2L baseline according to its original description, for CelebA, we treat it as an augmentation of the baseline CNN’s objective function. Detailed information regarding model architectures can be found in Appendix LABEL:sec:architectures and LABEL:sec:optimisation-details.

Synthesising Dataset Bias. For our experiments, we require a training set that exhibits a strong spurious correlation, together with a test set that does not. For cMNIST, this is easily satisfied as we have complete control over the data generation process. For CelebA and UCI Adult, on the other hand, we have to generate the split from the existing data. To this end, we first set aside a randomly selected portion of the dataset from which to sample the biased dataset The portion itself is then split further into two parts: one in which (s=−1∧y=−1)∨(s=+1∧y=+1)(s=-1\land y=-1)\lor(s=+1\land y=+1) holds true for all samples, call this part Deq\mathcal{D}_{eq}, and the other part, call it Dopp\mathcal{D}_{opp}, which contains the remaining samples. To investigate the behaviour at different levels of correlation, we mix these two subsets according to a mixing factor η\eta. For η≤12\eta\leq\tfrac{1}{2}, we combine (all of) Deq\mathcal{D}_{eq} with a fraction of 2η2\eta from Dopp\mathcal{D}_{opp}. For η>12\eta>\tfrac{1}{2}, we combine (all of) Dopp\mathcal{D}_{opp} and a fraction of 2(1−η)2(1-\eta) from Deq\mathcal{D}_{eq}. Thus, for η=0\eta=0, the biased dataset is just Deq\mathcal{D}_{eq}, for η=1\eta=1 it is just Dopp\mathcal{D}_{opp} and for η=12\eta=\tfrac{1}{2} the biased dataset is an ordinary subset of the whole data. The test set is simply the data remaining from the initial split.

Evaluation protocol. We evaluate our results in terms of accuracy and fairness. A model that perfectly decouples its predictions from ss will achieve near-uniform accuracy across all biasing-levels. For binary ss/yy we quantify the fairness of a classifier’s predictions using demographic parity (DP): the absolute difference in the probability of a positive prediction for each sensitive group.

We report the results from two image datasets. cMNIST, a synthetic dataset, is a good starting point for evaluating our model due to the direct control we have over the biasing. CelebA, on the other hand, is a more practical and challenging example. We also test our method on a tabular dataset, the Adult dataset.

cMNIST. The coloured MNIST (cMNIST) dataset is a variant of the MNIST dataset in which the digits are coloured. In the training set, the colours have a one-to-one correspondence with the digit class. In the test set (and the representative set), colours are assigned randomly. The colours are drawn from Gaussians with 10 different means. We follow the colourisation procedure outlined by , with the mean colour values selected so as to be maximally dispersed. The full list of such values can be found in Appendix LABEL:sec:color-details. We produce multiple variants of the cMNIST dataset corresponding to different standard deviations σ\sigma for the colour sampling: σ∈{0.00,0.01,...,0.05}\sigma\in\{0.00,0.01,...,0.05\}.

For this specific dataset, we can establish an additional baseline by simply grey-scaling the dataset which only leaves the luminosity as spurious information. We also evaluate the model, with all the associated hyperparameters, from . The only difference between the setups is the dataset creation, including the range of σ\sigma values we consider. Our versions of the dataset, on the whole, exhibit much stronger colour bias, to the point of the mapping the digit’s colour and class being bijective. Fig. 6 shows that the model significantly underperforms even the naïve baseline, aside from at σ=0\sigma=0, where they are on par.

Inspection of the null-samples shows that both the cVAE and cFlow model succeed in removing almost all colour information, which is supported quantitatively by Fig. 6. While the cVAE outperforms cFlow marginally at low σ\sigma values, performance degrades as this increases. This highlights the problems with the conditional decoder we anticipated in Section 3.3. The lower σ\sigma, and therefore the variation in sampled colour, is, the more reliably the ss label, corresponding to the mean of RGB distribution, encodes information about the colour. For higher σ\sigma values, the sampled colours can deviate far from the mean and so the encoder must incorporate information about ss into its representation if it is to minimise the reconstruction loss. cFlow, on the other hand, is consistent across σ\sigma values.

CelebA. To evaluate the effectiveness of our framework on real-world image data we use the CelebA dataset , consisting of 202,599 celebrity images. These images are annotated with various binary physical attributes, including “gender”, “hair color”, “young”, etc, from which we select our sensitive and target attributes. The images are centre cropped and resized to 64×6464\times 64, as is standard practice. For our experiments, we designate “gender” as the sensitive attribute, and “smiling” and “high cheekbones” as target attributes. We chose gender as the sensitive attribute as it a common sensitive attribute in the fairness literature. For the target attributes, we chose attributes that are harder to learn than gender and which do not correlate too strongly with gender in the dataset (“wearing lipstick” for example being an attribute too closely correlated with gender). The model is trained on the representative set (normal subset of CelebA) and is then used to encode the artificially biased training set and the test set. The results for the most strongly biased training set (η=0\eta=0) can be found in Fig. 4. Our method outperforms the baselines in accuracy and fairness.

We also assess performance for different mixing factors (η\eta) which correspond to varying degrees of bias in the training set (see Fig. 5). This is to verify that the model does not harm performance when there is not much bias in the training set. For these experiments, the model is trained once on the representative set and is then used to encode different training sets. The results show that for the intermediate values of η\eta, our model incurs a small penalty in terms of accuracy, but at the same time makes the results fairer (corresponding to an accuracy-fairness trade-off). Qualitative results can be found in Fig. 1 (images from cVAE can be found in Appendix LABEL:sec:qual-results-celeba).

To show that our method can handle multinomial, as well as binary, sensitive attributes, we also conduct experiments with s=hair colors=\textrm{hair color} as a ternary attribute (“Blonde” “Black”, “Brown”), excluding “Red” because of the paucity of samples and the noisiness of their labels. The results for these experiments can be found in Appendix LABEL:ssec:multi-s.

Results for the UCI Adult dataset. The UCI Adult dataset consists of census data and is commonly used to evaluate models focused on algorithmic fairness. Following convention, we designate “gender” as the sensitive attribute ss and whether an individual’s salary is 50,000orgreateras50,000 or greater asy.WeshowtheperformanceofourapproachincomparisontobaselineapproachesinFig.7.Weevaluatetheperformanceofallmodelsformixingfactors(. We show the performance of our approach in comparison to baseline approaches in Fig. 7. We evaluate the performance of all models for mixing factors (\eta)and) and1.ResultsshowninFig.7showthatwematchorexceedthebaseline.Intermsoffairnessmetrics,ourapproachgenerallyoutperformsthebaselinemodelsforbothof. Results shown in Fig. 7 show that we match or exceed the baseline. In terms of fairness metrics, our approach generally outperforms the baseline models for both of\eta$. Detailed results can be found in the Appendix LABEL:ssec:detailed-adult.

We also did experiments to show that the encoder transfers to other tasks. These transfer-learning experiments can be found in Appendix LABEL:sec:transfer-learning.

Conclusion

We have proposed a general and straightforward framework for producing invariant representations, under the assumption that a representative but partially-labelled representative set is available. Training consists of two stages: an encoder is first trained on the representative set to produce a representation that is invariant to a designated spurious feature. This is then used as input for a downstream task-classifier, the training data for which might exhibit extreme bias with respect to that feature. We train both a VAE- and INN-based model according to this procedure, and show that the latter is particularly well-suited to this setting due to its losslessness. The design of the models allows for representations that are in the data domain and therefore exhibit meaningful invariances. We characterise this for synthetic as well as real-world datasets for which we develop a method for simulating sampling bias.

Acknowledgements

This work was in part funded by the European Research Council under the ERC grant agreement no. 851538. We are grateful to NVIDIA for donating GPUs.

References

References