Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing

Yisong Xiao, Aishan Liu, Tianlin Li, Xianglong Liu

Introduction

ML systems have been extensively involved in making social-critical decisions, such as employment and loan (Chan and Wang, 2018; Bono et al., 2021; Berk et al., 2021; Krittanawong et al., 2020; Mukerjee et al., 2002). Despite the promising performance, ML systems can also inadvertently introduce societal biases, which violate the fairness requirements of software design. Recent studies have revealed that some ML systems may be discriminatory based on protected attributes (e.g., gender, race). For example, the hiring system used by Amazon (Dastin, 2018) is showing a strong preference for male candidates, while rating the resumes of female candidates with lower scores. Therefore, it is crucial to conduct exhaustive tests to detect fairness violations and further reduce discrimination, which can relieve the fairness concerns of ML systems in social-critical applications.

To evaluate fairness, extensive studies have been proposed to generate test samples that can violate the fairness requirement and reveal the discrimination of ML systems. Generally, existing studies can be primarily divided into two categories: individual fairness (Dwork et al., 2012; Galhotra et al., 2017) evaluation and group fairness (Feldman et al., 2015; Hardt et al., 2016) evaluation. Individual fairness requires similar decisions for similar individuals, while group fairness requires equal treatment for different groups (grouped by a protected attribute). Individual fairness is capable of identifying discriminatory behaviors that may be ignored by group fairness measures (Galhotra et al., 2017). Specifically, individual fairness allows for a more powerful and granular examination of discriminatory behavior that arises due to changes in only the protected attribute in various contexts, while group fairness measures may fail to detect discrimination when the model treats the same group oppositely in different situations (Zhang et al., 2021). Therefore, our primary objective in this work is to conduct individual fairness testing, which entails generating individual discriminatory instances We use “instance” or “sample” interchangeably. for testing purposes. In other words, we aim to generate two similar instances that only differ in the protected attribute but show different model decisions, which make a system exhibit individual discrimination (Galhotra et al., 2017; Aggarwal et al., 2019).

For individual fairness testing, existing methods primarily generate discriminatory instances in a two-phase generation framework (Galhotra et al., 2017; Udeshi et al., 2018; Aggarwal et al., 2019; Zhang et al., 2020b, 2021; Zheng et al., 2022; Fan et al., 2022). Specifically, they first construct a set of individual discriminatory instances as global seeds via random sampling or gradient-based search; they then conduct an iterative process to search for instances near the global seeds; finally, they utilize these identified discriminatory instances to improve the individual fairness of the tested model via retraining. However, the naturalness of generated instances remains unsatisfactory (i.e., the generated instances often deviate from the original data distribution), which impedes fairness testing in practical application. For example, the generated instances by these methods may violate real-world constraints (e.g., authorize a loan to a 10-year-old individual) (Chen et al., 2022), thus falsely estimate the fairness of models.

To address the problem, this paper proposes a framework named Latent Imitator (LIMI), which can generate natural discriminatory instances that is closer to the original data distribution. Our LIMI approach imitates the decision boundary of the target model and probes latent samples on it to generate more natural discriminatory instances. Previous studies (Mickisch et al., 2020; Karimi et al., 2019) have shown that the decision boundary tends to be close to the majority of natural data points after training, which indicates a close alignment between the decision boundary and the distribution of the training data. Therefore, test cases near the decision boundary are more likely to possess better naturalness and have a higher potential to induce discrimination. However, for the real-world black-box models, we have no access to the detailed model information and fail to derive a semantic decision boundary in the original input domain. Thus, we first approximate a surrogate decision boundary in the latent space of a generative model, which enjoys semantic property (Radford et al., 2015). Specifically, we construct an auxiliary dataset that maps the latent space to the model decision space and then learns a linear hyperplane based on the dataset as a surrogate boundary of the real decision boundary. The surrogate decision boundary depicts a coarse region for test sample generation, and instances near it will possess better discrimination and naturalness. However, there may exist a gap between the surrogate boundary and the real decision boundary, which may induce an inaccurate latent location. We thus design the latent candidates probing strategy, in which we first locate a latent vector at the surrogate boundary with a one-step movement and then probe two potential candidates on either side of it that may be closer to the real boundary. Therefore, we can finely probe discriminatory latent vectors on the real decision boundary for better naturalness.

To evaluate the performance of our LIMI, we conduct extensive experiments under multiple datasets on both traditional ML models and novel DNNs. Compared to 6 state-of-the-art individual fairness testing approaches, our LIMI generates ×\times9.42 more discriminatory instances at ×\times8.71 faster speed, and the distribution of generated instances possesses better naturalness (+19.65%) on average. In addition, benefiting from the naturalness of our generated discriminatory instances, we can improve the model fairness in both individual fairness (45.67% on IFrIF_{r} and 32.81% on IFoIF_{o}) and group fairness (9.86% on SPDSPD and 28.38% on AODAOD) via retraining. Moreover, to better understand our framework, we conduct further investigations and demonstrate that our framework performs well on image datasets and can be flexibly combined with other baselines as a fast global prober. Our main contributions are:

To the best of our knowledge, we are the first to generate natural individual discriminatory instances for fairness testing, which could help reveal the actual model fairness.

We propose LIMI framework that can generate more natural individual discriminatory instances by imitating the decision boundary of the target model and further sampling latent instances on it.

Extensive experiments on several benchmarks demonstrate LIMI outperforms SOTA baselines largely in effectiveness (×\times9.42 instances), efficiency (×\times8.71 speeds), and naturalness (+19.65%) on average.

We publish LIMI as a self-contained toolkit on our website (Anonym, 2023).

Preliminaries

In this section, we first provide a brief review of the relevant background, including the binary classification model, and generative adversarial networks; we then illustrate the problem definition.

Binary Classification Model. Given a dataset D\mathcal{D} with data sample x∈Xx\in X and label y∈Yy\in Y, the binary classification model can be represented as f(x,θ):X→Yf(x,\theta):X\rightarrow Y, where YY only contains two classes and θ\theta denotes the parameters of model ff. The goal of these binary classification models (e.g., Random Forests (RF) (Ho, 1998), Support Vector Machines (SVM) (Cortes and Vapnik, 1995), and Deep neural networks (DNN) (LeCun et al., 2015)) is to learn a decision boundary, which is a surface in the input domain that can separate input instances into two classes. Formally, the decision boundary can be written as F={x:θ(x)=0}F=\{x:\theta(x)=0\}, where θ(x)<0\theta(x)<0 denotes xx is classified as the negative class, and θ(x)>0\theta(x)>0 is the positive one.

Generative Adversarial Networks (GANs). GANs are commonly used to generate high-quality samples for data synthesis. The GAN framework was first proposed for synthetic images by Goodfellow et al. (Goodfellow et al., 2014), which plays a zero-sum game between the generator GG and the discriminator DD. Generator GG is trained to mimic the given real training dataset distribution, which maps a synthetic sample g(z)g(\mathbf{z}) from random latent vector z\mathbf{z}; while discriminator DD is trained to distinguish the generated samples from real samples. The core idea of GAN is the adversarial training process according to the following min-max objective:

where x∼pdatax\sim p_{data} denotes the real distribution, and z∼pz\mathbf{z}\sim p_{\mathbf{z}} denotes the distribution of latent space.

Besides visual images, GANs can also be utilized for tabular data generation. Xu et al. (Xu and Veeramachaneni, 2018) first proposed TGAN to map the continuous noise latent into discrete tabular data. Based on that, Xu et al. (Xu et al., 2019) proposed CTGAN, which invents the mode-specific normalization to overcome the non-Gaussian and multi-modal distribution, and further designs to deal with the imbalanced discrete columns.

2. Problem Definition

Individual discrimination. Let A={A1,A2,...,An}A=\{A_{1},A_{2},...,A_{n}\} and I={I1,I2,...,In}I=\{I_{1},I_{2},...,I_{n}\} represent the attributes set of data samples XX and its input domain. Following (Zhang et al., 2020b), individual discrimination refers to the unequal treatment (e.g., different model decisions) of two similar individuals who only differ in the protected attributes. Here, the protected attributes (e.g., gender, race, and age) have been clearly defined and recognized in various laws and policies (e.g., GDPR and UK equality law). Following the commonly-adopted settings (Galhotra et al., 2017; Udeshi et al., 2018; Aggarwal et al., 2019; Zhang et al., 2020b, 2021; Zheng et al., 2022; Fan et al., 2022; Udeshi et al., 2018), the protected attributes are assumed to be a known prior, and we use PAPA to denote the set of protected attributes and NPANPA to denote the non-protected attributes. For each instance x={x1,x2,...,xn}∈Xx=\{x_{1},x_{2},...,x_{n}\}\in X (xix_{i} is the value of corresponding attribute AiA_{i}), we define that xx is an individual discriminatory instance for model ff when there exists another instance x′x^{\prime} that satisfies

Moreover, we use (x,x′)(x,x^{{}^{\prime}}) to represent an individual discriminatory instance pair for ff. Thus, in this paper, we aim to generate individual discriminatory instance for fairness testing, which could force the target model to produce biased decisions and violates the individual fairness requirements. In accordance with common assumptions in black-box fairness testing (Udeshi et al., 2018; Aggarwal et al., 2019; Fan et al., 2022), we assume that the data samples XX, the corresponding feature space AA, and the protected attributes set PAPA are provided for testers, while the target model ff trained on dataset D={X,Y}\mathcal{D}=\{X,Y\} is black-box (without prior knowledge of inner information like gradients).

Naturalness of generated instances. In addition to individual discrimination, we also prioritize the naturalness of generated discriminatory instances. In the context of tabular data, naturalness is a measure of how well the dataset reflects the underlying distribution of the population (Black and van Nederpelt, 2020). Specifically, naturalness refers to how closely the statistical patterns and relationships in a dataset, such as value distributions and attribute correlations, match the distribution of the population it represents (Brenninkmeijer et al., 2019). To quantify the naturalness of generated data, statistical measures such as the Average Nearest Neighbor Distance (Peterson, 2009), Pearson’s correlation (Sedgwick, 2012), and Kolmogorov-Smirnov statistic (Massey Jr, 1951) can be used. These metrics evaluate the distributional distance/similarity between the original and generated data, with the original data (real-world datasets) serving as the reference population. Generated data with high naturalness has a closer distribution to the original dataset, indicating that the data is a more accurate representation of the population (Wen et al., 2021; Xu et al., 2019). Therefore, we aim to generate individual discriminatory instances with high naturalness, which are more likely to appear in the population and can better reveal unfair behaviors in the real world.

Methodology

In this paper, we propose a framework named Latent Imitator (LIMI) to generate natural individual discriminatory instances for black-box fairness testing. Since the decision boundary of the target model often reflects the nature of the original data distribution, we first coarsely approximate the decision boundary of the target model by deriving a surrogate boundary in the latent space of GAN. Thus, test cases near the surrogate boundary are more likely to serve as discriminatory instances with slight perturbations and possess better naturalness. However, there may exist a gap between the surrogate boundary and the real boundary, which would hinder the accurate location of latent sample generation. Therefore, we then manipulate random latent vectors with a one-step movement to the surrogate boundary and probe two potential candidates, so that we could find potential discriminatory latent vectors that may be more closely located around the decision boundary leading to better naturalness. The overall framework is shown in Fig 1.

As established in previous work (Mickisch et al., 2020; Karimi et al., 2019), the decision boundary of the model is often close to the majority of natural data points after training, which highly reflects the nature of the original data distribution at the same time. Therefore, test cases generated under the guidance of the target model decision boundary could be obviously more natural (i.e., closer to the original distribution). However, in the black-box fairness testing scenario, we cannot directly access the model and fail to derive a semantic decision boundary in the original input domain. Thus, we try to approximate the target model’s decision boundary by deriving a surrogate boundary in the latent space of GAN based on a constructed synthetic latent dataset.

GAN has been widely used to synthesize massive realistic data for software testing (Zhang et al., 2018; Gao and Han, 2019), therefore we use GANs to help us generate the synthetic latent dataset. Specifically, we can derive a sample xx in the input domain from a randomly sampled latent vector z\mathbf{z} by the generator GG of GAN as g(z):Z→Xg(\mathbf{z}):\mathbf{Z}\rightarrow X; the sample g(z)g(\mathbf{z}) can be then classified by the target model with a binary label f(g(z))f(g(\mathbf{z})) in the prediction domain. The pipeline implicitly embodies the model decision process on the latent space, which can be used for the latent decision boundary approximation. Based on the above analysis, we aim to find a surrogate decision boundary Flatent={z:θ(g(z))=0}F_{latent}=\{\mathbf{z}:\theta(g(\mathbf{z}))=0\} in the latent space to characterize the implicit relationship, so that we can approximate the real decision boundary of the target black-box model for subsequent individual discriminatory instances generation. Since studies (Denton et al., 2019; Shen et al., 2020) have illustrated that the latent space of GAN is approximately linearly separable in any binary semantic attribute (e.g., whether to loan), we, therefore, employ a linear hyperplane as the coarse surrogate decision boundary. Coarsely, FlatentF_{latent} can be written as following expression

where ww denotes the normal vector, and bb represents the intercepts.

The latent vector z\mathbf{z} is assigned to an unfavorable sample (e.g., not loan) of model ff when it satisfies wTz+b<0\mathbf{w}^{T}\mathbf{z}+b<0, otherwise, a favorable sample (e.g., loan) if it satisfies wTz+b>0\mathbf{w}^{T}\mathbf{z}+b>0. It is worth noting that unlike a surrogate decision tree produced by LIME (Ribeiro et al., 2016) in the original input domain, the surrogate boundary in the continuous latent space has vector arithmetic property (Radford et al., 2015), which indicates that latent vectors can be manipulated semantically under its guidance. We can easily obtain a sample of the opposite class from the initial random point through latent vector arithmetic (e.g., adding with the normal vector ww). Therefore, the operable (i.e., vector arithmetic property) surrogate boundary establishes a foundation for the subsequent probing process.

To better obtain the surrogate decision boundary (i.e., a linear hyperplane), we need to construct an auxiliary latent dataset, which maps the latent space to the model decision space. Concretely, we sample a set of latent vectors Zinit\mathbf{Z}_{init} randomly and then utilize the generator of GAN to synthesize samples in the original input domain XX; we then label each sample with the corresponding prediction result based on the target tested model. Formally, the auxiliary latent dataset Daux\mathcal{D}_{aux} can be expressed as

Note that the fitness of our surrogate boundary is highly related to the effectiveness of subsequent test case generation. Thus, we try to obtain a more precise boundary by refining the auxiliary dataset Daux\mathcal{D}_{aux} with the ranking and resampling process. Since high-confidence samples tend to be more concentrated, while low-confidence samples are often distributed at the twists and turns of the decision boundary, denser high-confidence samples have a higher probability of reflecting the real condition of the lumpy neural network boundary (Guan and Loew, 2020). Therefore, we first rank the auxiliary latent dataset based on each classification score assigned by model ff, then retain latent vectors whose scores are higher than threshold ϵ\epsilon. To guarantee sufficient high-quality training samples, ϵ\epsilon is empirically set as 0.7. Considering the imbalance of the synthetic data, we adopt a random over-sampling strategy (Santoso et al., 2017) that adds more samples from the minority class in Daux\mathcal{D}_{aux} by simply replicating, so as to construct a class-balanced dataset. Based on the constructed auxiliary dataset Daux\mathcal{D}_{aux}, we finally train our linear representation model (i.e., SVM with linear kernel) to derive a surrogate decision boundary FlatentF_{latent}.

To sum up, we first approximate the boundary of the target model by deriving a surrogate decision boundary in the latent space of GAN based on a constructed synthetic latent dataset. The surrogate decision boundary depicts the target region for the generation of individual discriminatory instances (i.e., possess both natural and discriminatory properties) in the latent space coarsely, which will provide strong guidance for the later latent probing process.

2. Latent Candidates Probing

Existing studies (Zhang et al., 2020b; Moosavi-Dezfooli et al., 2016; Brendel et al., 2017) have revealed the fact that data points near the decision boundary are more likely to mislead the model’s predictions where only a slight perturbation is required. Therefore, after obtaining a latent surrogate decision boundary, we probe the latent space for the random initial samples and move them near the surrogate boundary, so that we could find potential natural and discriminatory latent vectors.

With the surrogate boundary hh obtained in latent space, we can easily calculate the distance dd from any random latent vector z\mathbf{z} to hh, which reflects the decision score of model ff under the corresponding test case xx:

When a latent vector z\mathbf{z} gets closer to the surrogate boundary, the uncertainty of model prediction increases extremely, while the corresponding synthetic sample still stays in the central distribution of the original dataset. Therefore, when we generate a test case xx from latent z\mathbf{z} near our surrogate boundary and modify its protected attribute value to get another instance x′x^{{}^{\prime}}, they will most likely get different prediction results from the target model. This assumption exploits the inherent adversarial robustness property of machine learning models (Szegedy et al., 2013), that is a model can be misled by a slight perturbation added to the input. In this way, we can obtain an individual discriminatory instance with maximum possibility and acquire naturalness simultaneously.

Based on the above analysis, given a random latent vector z\mathbf{z} as an initial point, we can naively make a search process that perturbs it toward the surrogate boundary iteratively, which is similar to the existing fairness testing methods (Zhang et al., 2020b, 2021; Zheng et al., 2022). The perturbing process can be simply defined as

where dirdir is the direction (±1\pm 1) to the surrogate boundary, sps_{p} denotes the step size of the iterative search process, and wu\mathbf{w}_{u} represents the unit of w\mathbf{w} (i.e., wu=wwTw\mathbf{w}_{u}=\frac{\mathbf{w}}{\sqrt{\mathbf{w}^{T}\mathbf{w}}}), which directs the shortest path.

However, we note the tedious nature of the iterative search process that it is time-consuming which would highly influence the speed of test case generation. Thus, we aim to accelerate the generation process, where we replace the iterative searching process with the one-step probing strategy. The refined latent manipulation can be formally written as

where bu=bwTwb_{u}=\frac{b}{\sqrt{\mathbf{w}^{T}\mathbf{w}}} is also the normalized unit. Due to the implicit constraints of the surrogate boundary in latent space, the manipulated latent vectors still maintain their naturalness. We can verify the property of z0\mathbf{z}_{0} (d(z0)=0d(\mathbf{z}_{0})=0) by taking it into the distance function. Though the single-step probing strategy facilitates efficiency significantly, it may lose accuracy due to the coarseness of our surrogate boundary.

Notice that, in the first coarse step of our approach, we utilize a simple linear hyperplane to approximate tested ML models, while the real decision boundary of these models will be nonlinear and more complex. For example, the boundary of a simple neural network is a subset of tropical hypersurface (Alfarra et al., 2022) due to its non-linear activation. Therefore, the manipulated latent vector z0\mathbf{z}_{0} may not accurately locate at the real decision boundary, which fails to reflect strong naturalness. To solve this potential problem, we propose an additional latent candidate probing strategy where we introduce two candidates on the side of z0\mathbf{z}_{0} to bridge the approximation gap between our surrogate boundary and the real decision boundary, therefore, leading to better naturalness of the generated instances. Specifically, these two candidates can be calculated with the following expressions

where λ\lambda represents a hyperparameter that controls the walking length along direction wu\mathbf{w}_{u}. For concrete tabular data, small λ\lambda would result in generating duplicate test cases, making the candidates useless; on the contrary, big λ\lambda can incur those candidates to move far away from decision boundaries, which deviates from the original data distribution. We will conduct ablation studies in Section 5.

Therefore, given a set of random latent vectors Zinit\mathbf{Z}_{init}, for each z∈Zinit\mathbf{z}\in\mathbf{Z}_{init} we can calculate potential individual discriminatory latent vectors near the model decision boundary to construct Didz={(z0,z+,z−):z∈Zinit}\mathcal{D}_{idz}=\{(\mathbf{z}_{0},\mathbf{z}_{+},\mathbf{z}_{-}):\mathbf{z}\in\mathbf{Z}_{init}\}.

To sum up, we design two probing strategies to obtain latent vectors that are most likely to generate natural discriminatory test cases, one is to probe latent vectors located at the surrogate boundary, and another is to probe two candidates to bridge the boundary approximation gap. With only a few lightweight vector calculations, we can finely probe many latent vectors that are distributed near the surrogate boundary, which will benefit the naturalness of generated discriminatory instances.

3. Overall Testing Process

Algorithm 1 shows the overall testing process of our proposed LIMI approach. In general, given a black-box target model fθf_{\theta} and a generator GG both trained on dataset XX, our testing framework can be roughly divided into two phases, i.e., latent boundary approximation and latent candidates probing. In the first phase, we construct an auxiliary dataset Daux\mathcal{D}_{aux} that consists of latent vectors and corresponding decision labels from model fθf_{\theta}, and then adopt a linear kernel SVM to derive a surrogate decision boundary in the latent space based on Daux\mathcal{D}_{aux}. In the second phase, we manipulate random latent vectors to the surrogate boundary in one step and probe two candidates near them for better locating the real decision boundary, and then collect the probed latent vectors to construct potential discriminatory vectors Didz\mathcal{D}_{idz}.

After obtaining the latent candidates, we utilize the generator GG to synthesize Didz\mathcal{D}_{idz} as the set of potential individual discriminatory instances as

Finally, we check whether a test case in Dtest\mathcal{D}_{test} is a discriminatory instance by modifying the value of the selected protected attribute. Once a candidate sample in the tuple (g(z0),g(z+),g(z−))(g(z_{0}),g(z_{+}),g(z_{-})) is determined as a discriminatory instance, we will add it into the discriminatory set Didi\mathcal{D}_{idi} and break out to stop the testing of the left samples in this tuple immediately.

It is worth noting that our approach has no iterative search process, which is much more efficient than the search-based generation framework of other approaches (Zhang et al., 2020b, 2021; Zheng et al., 2022). Literally, the search-based framework needs to compute guidance (i.e., direction to decision boundary) before each iteration, while our method avoids it due to the latent vector arithmetic property (Radford et al., 2015). Though researchers have been optimizing the guidance computation methods (e.g., modifying the loss function of gradient calculations) to improve search efficiency in their long-term work, the search process still consumes a long time. In our framework, the surrogate decision boundary depicts a probing region coarsely, then the tedious iteration search is replaced with a latent probing process, which significantly accelerates the testing process.

Evaluation

In this section, we evaluate the performance of LIMI in terms of effectiveness, efficiency, the naturalness of generated instances, and the utility for fairness improvement. We first outline the experimental setup and then conduct the evaluation by answering the following research questions.

RQ1: How effective and efficient is LIMI in finding individual discriminatory instances?

RQ2: How natural are the individual discriminatory instances generated by LIMI?

RQ3: How useful are the generated individual discriminatory instances for improving the model’s individual and group fairness?

We evaluate LIMI on four public tabular datasets, which are widely adopted in the fairness testing literature (Zhang et al., 2020b; Udeshi et al., 2018; Aggarwal et al., 2019). The following is a brief description of these datasets.

Adult Income Dataset (adu, 2017) is used for income prediction with 48,842 samples including 32,561 train samples (7,841 favorable vs 24,720 unfavorable) and 16,281 test samples (3,846 favorable vs 12,435 unfavorable), and the label denotes whether an individual’s annual income is over \50K$. The protected attributes in 13 attributes (5 numerical and 8 categorical) of Adult are gender, race, and age respectively.

German Credit Dataset (cre, 1994) is used for credit risk level prediction (i.e., good or bad) with 600 samples (300 favorable vs 300 unfavorable) and 20 features (7 numerical and 13 categorical). The protected attributes are gender and age.

Bank Marketing Dataset (ban, 2014) labels whether a client will subscribe to a term deposit or not. There are 45,211 samples (5,289 favorable vs 39,922 unfavorable) and 16 features (6 numerical and 10 categorical), and the protected attribute is age.

Medical Expenditure Panel Survey Dataset (mep, 2015) is used for health care needs prediction with 15,675 samples (2,628 favorable vs 13,047 unfavorable) and 40 features (4 numerical and 36 categorical), and the protected attribute is gender.

We use Adult, Credit, Bank, and Meps for simplicity. To ensure the rationality of subsequent analysis, we also examine the correlation of attributes using Spearman’s rank order method (Spearman, 1961). We find that the protected attributes are not significantly correlated to any other attributes (average correlation among attributes, Adult: 0.09, Credit:0.10, Bank:0.06, Meps:0.05), which is consistent with the observations in (Zhang et al., 2021).

1.2. Binary Classification Models

We employ three commonly-used binary classification models to evaluate LIMI. The traditional machine learning models are RF and SVM implemented in scikit-learn (Kramer and Kramer, 2016) library. Especially, the RF consists of 100 trees, and the SVM is implemented with an RBF kernel, following (Fan et al., 2022). The DNN model is a six-layer fully-connected neural network, with architecture (Zhang et al., 2020b) in detail. And we optimize the DNN models for 1000 epochs by the Adam optimizer with a learning rate of 0.001. These models all perform well in classification, with an average accuracy of 91.60%. Our training and testing settings align with commonly-adopted practices in (Fan et al., 2022; Zhang et al., 2020b, 2021; Zheng et al., 2022), and we provide more information on our implementation and results on our website (Anonym, 2023).

1.3. Baselines

We compare LIMI with 6 state-of-the-art approaches, which can be divided into two types based on the tested model. The first type includes Aequitas (Udeshi et al., 2018), SymbGen (Aggarwal et al., 2019), and ExpGA (Fan et al., 2022), which provide the testing ability for all machine learning models. The second type includes ADF (Zhang et al., 2020b), EIDIG (Zhang et al., 2021), and NeuronFair (Zheng et al., 2022), which are designed especially for DNN. We use the implementation of these approaches directly from corresponding GitHub repositories and adhere to the best settings reported in their papers.

1.4. Evaluation Metrics

The following three aspects of LIMI are evaluated, including the effectiveness and efficiency of generation, the naturalness of generated individual discriminatory instances, and the fairness improved after retraining.

Effectiveness and efficiency Evaluation. We use #Didi\#\mathcal{D}_{idi} and EGSEGS metrics to measure the effectiveness and efficiency of the individual discriminatory instances generation procedure. #Didi\#\mathcal{D}_{idi} denotes the absolute quantity of generated individual discriminatory instances. EGSEGS denotes the speed of effective test case generation (i.e., how many individual discriminatory instances can the approach generate per second):

where TimeTime represents the consumed time during the generation procedure. Thus, a larger value of #Didi\#\mathcal{D}_{idi} means the method is more effective, and a higher value of EGSEGS means the method is more efficient.

Naturalness Evaluation. Naturalness measures how similar the distribution of generated instances is compared with the original dataset. To thoroughly evaluate the naturalness, we adopt a comprehensive metric similarly in (Wen et al., 2021), which considers both the distribution of each column and the correlation between pairs of columns in evaluation.

As for the distribution, we use Kolmogorov-Smirnov Test (for numerical column) and the Total Variation Distance (for categorical column) to calculate its score. As for the correlation, we use Pearson Coefficient (for two numerical columns) and the Contingency Similarity (for categorical column with any kind of column) to calculate its score. And we compute the average of these two scores to measure the naturalness of Didi\mathcal{D}_{idi}, denoted as ATNATN (Average Tabular Naturalness):

where mean(⋅)\rm{mean}(\cdot) calculates the average value of set, XpX_{p} denotes the pth column of XX, and KS is an inverted version of Kolmogorov-Smirnov Test. The range of ATNATN is 0 to 1, and we say that a method presents a better naturalness if its generated Didi\mathcal{D}_{idi} achieves a higher value of ATNATN. We use the implementation provided by SDMetrics (DataCebo, 2022) library for evaluation. Besides ATNATN, we also evaluate naturalness utilizing both classifier-based detection (LogisticRegression, SVC) and distance-based method (Average Nearest Neighbor Euclidean distances). Their details can be found on our website (Anonym, 2023).

Fairness evaluation. Both individual fairness and group fairness are taken into evaluation. To measure individual fairness, we first follow previous work (Udeshi et al., 2018) to randomly sample a large set of instances and check the ratio of discriminatory instances in the set (denoted as IFrIF_{r}); moreover, we calculate the proportion of individual discriminatory instances that exists in the original dataset (denoted as IFoIF_{o}):

To measure group fairness, we adopt Statistical Parity Difference (SPD) and Average Odds Difference (AOD) metrics following (Hort et al., 2021). The SPD requires that a decision should be independent of the protected attributes (Barocas and Selbst, 2016), which measures the difference in positive classification between different demographic groups:

where y^\hat{y} denotes the model prediction and AA is the group of the protected attribute.

The AOD represents the average of the differences in True Positive Rate (TPR) and False Positive Rate (FPR) between privileged and unprivileged groups (Hardt et al., 2016):

Under the above definitions, lower fairness metric values indicate a fairer model, while larger values denote a higher level of discrimination in the model.

1.5. Implementation details

We adopt the CTGAN (Xu et al., 2019) as the GAN model in experiments. We train the model with 300 epochs and the batch size is set as 500 for each tabular dataset. The average consumed training time is 13 minutes. For the surrogate boundary, we implement it with a linear SVM trained on the refined randomly sampled latent dataset Daux\mathcal{D}_{aux} (containing 100K latent vectors). Without specification, we set the hyperparameter λ\lambda as 0.3. We conduct our experiments on a server with Intel(R) Xeon(R) Gold 6230R CPU @ 2.10GHz, 256GB system memory, and an NVIDIA GeForce RTX-2080-Ti GPU.

2. RQ1: Effectiveness and Efficiency

For the baselines, we follow their original two-stage settings (Aggarwal et al., 2019; Zhang et al., 2020b, 2021; Zheng et al., 2022; Fan et al., 2022; Udeshi et al., 2018), which first search 1000 instances to construct individual discriminatory seeds in the global phase and then search 1000 neighbor instances for each seed (the maximum number of total test cases is 1K ×\times 1K = 1M). As for our approach, we use the entire randomly sampled dataset Zinit\mathbf{Z}_{init} (contains 1M latent vectors) as start points and set the maximum number of test cases as 1M for fair comparisons. We opt to maintain the original two-stage framework of the baselines since that directly combining the two search numbers and restricting the total number of test cases to 1M would disrupt the framework and adversely affect their performance. Following (Fan et al., 2022), we constrain the testing time as one hour. To mitigate contingency, we repeat the generation process 5 times and report the average results.

The results on DNN and RF are shown in Table 1 and 2, while the results on SVM testing can be found on our website (Anonym, 2023). Note that the gradient-based methods (i.e., ADF, EIFIG, and NeuronFair) are unavailable for the traditional machine learning models, and we only apply them to DNN models. From the results, we could make several observations as follows:

As for the generation effectiveness, our approach generates more individual discriminatory instances than other baselines and outperforms them largely (+135,534 on average). Specifically, for DNN models, as shown in Table 1, the average #Didi\#\mathcal{D}_{idi} on four datasets of LIMI is 149,711, which is 18 times and 3 times better than NeuronFair and ExpGA, respectively; for RF models, as shown in Table 2, our method finds 2.5 times and 2.2 times more discriminatory instances than SG and ExpGA; for SVM models, we also achieve good performance, especially on the Credit and Meps datasets. The above results demonstrate the outstanding performance of our method on the test instances generation effectiveness, which can be attributed to our surrogate decision boundary approximation and candidate probing strategies.

As for the generation efficiency, our LIMI has a faster generation speed and outperforms other baselines by 8.71 times on average. For example, as shown in Table 1, LIMI can generate almost 50 discriminatory instances per second on average while the value of NueronFair and ExpGA is 2 and 24, respectively. This indicates that our LIMI testing method is capable to give rapid feedback to software engineers, which will be preferable in time-constrained industrial software testing. We speculate the main reason is the direct vector calculation employed in the latent candidates probing phase, which can reduce the computational complexity and circumvent the tedious iterative search process.

Moreover, we observe that our method shows better stability across different combinations of datasets and target models. Specifically, for ExpGA, when testing the RF model on the Adult dataset with race as the protected attribute, #Didi\#\mathcal{D}_{idi} is only 286, which is much lower than its performance in other test conditions (e.g., #Didi\#\mathcal{D}_{idi} is 10,283 on the Adult with gender as the protected attribute); for AEQUITAS and SG, when testing SVM models on the Adult dataset, #Didi\#\mathcal{D}_{idi} even reduces to 0, denoting a test failure happened. In fact, both SG and ExpGA highly depend on the explanation results produced by the local explainer, which constrains their generation when the interpretation is weak. By contrast, our LIMI only requires the prediction results of the tested model and learns the surrogate decision boundary from the model decision to guide testing, which shows better tests for different models/datasets combinations.

To answer this question, we randomly select the same number of instances as the original dataset from our constructed Didi\mathcal{D}_{idi} and then measure ATNATN for each sampled set. To ensure sampling adequacy, we repeat 10 times and report the average ATNATN values. Considering the #Didi\#\mathcal{D}_{idi} values of some methods are less than the quantity of the original dataset, we remove the time limit and obtain sufficient discriminatory instances for testing.

From the results in Table 3, we identify that the ATNATN values of our method are significantly higher than that of other baselines on all datasets. Specifically, the ATNATN values of our method in all settings are higher than 80%, indicating that the generated instances by LIMI are more natural. In other words, the distribution of generated instances is much closer to the real data distribution. This observation could be used to further explain the outstanding performance of LIMI as it generates test cases near the real decision boundary of the target model that reflects the nature of the original data distribution. In more detail, the white-box methods (i.e., ADF, EIDIG, and NeuronFair) utilize gradients to perturb instances to the decision boundary and show second naturalness (69.11%, 69.03%, and 69.42% on average on all datasets). The impaired naturalness in these white-box methods can be attributed to modifying directly on the raw input (i.e., attribute value) without any restriction, unlike LIMI which probes instances in the semantic latent space implicitly limited by the GAN. This reason has also been observed in image data (Riccio and Tonella, 2020). As for black-box methods (i.e., AEQUITAS, SG, and ExpGA), they show the worst naturalness (58.28%, 48.59%, and 55.34% on average on all datasets). We speculate that the randomness in AEQUITAS and SG, as well as the crossover and mutation operator employed in ExpGA, induce the worst naturalness. As for the classifier-based and distance-based measures, our LIMI still outperforms other baselines, and we also observe consistent trends in naturalness (LIMI ¿ white-box methods ¿ black-box methods). Further details are reported on our website (Anonym, 2023).

To better illustrate the naturalness, we visualize the feature similarity between the generated and original data for NeuronFair, ExpGA, and our LIMI. In particular, we randomly select 1000 discriminatory instances on the DNN models and adopt PCA (Daffertshofer et al., 2004) to translate their features into one dimension; we then report and visualize the feature distribution of these instances. As shown in Figure 2, we can identify that the distribution of instances generated by LIMI (shown in orange) is the closest to the original data distribution, which indicates the better naturalness of our LIMI generation and further verifies our motivation.

4. RQ3: Fairness Improvement

To further demonstrate the importance of the naturalness of the generated discriminatory instances, we follow (Zhang et al., 2020b) and retrain the models using the generated instances to improve model fairness. Given the number of generated instances is much larger than the original dataset, we randomly select the instances with 30% size of the original dataset for retraining, which is almost the same quantity to (Zhang et al., 2020b). We use the selected instances and the original dataset to retrain DNN models and use the fairness metrics (i.e., IFrIF_{r}, IFoIF_{o}, SPDSPD, and AODAOD) to evaluate the fairness improvement. We also evaluate the performance of retraining RF models where we achieve similar results to DNNs. Due to the space limitation, the results can be found on our website (Anonym, 2023). For fair comparisons, we repeat the procedure 5 times and report the average values to avoid the effect of randomness.

The results regarding both the individual and group fairness are shown in Table 4, where column ‘Before’ denotes the model trained on the original dataset, and columns named by the approaches denote the model retrained with discriminatory instances they generated. From the results, we have the following observations:

As for the accuracy, the retrained models have a slight degradation (less than 1% on average), indicating that the models retrained with discriminatory instances can still perform well in real-world datasets.

As for the individual fairness, our LIMI outperforms other baselines significantly in terms of both IFrIF_{r} and IFoIF_{o}. On average, LIMI achieves individual fairness improvement of 45.67% in terms of IFrIF_{r}, and 32.81% in terms of IFoIF_{o}. In more detail, the models retrained by our method achieve 3.11 times and 2.22 times than NeuronFair, and 8.20 times and 3.84 times than ExpGA on individual fairness under two metrics. The above results indicate that individual discriminatory instances generated by our method show better effectiveness in the model fairness improvement.

As for the group fairness, we also notice that retrained models enjoy improvements on almost all datasets. However, the improvements are comparatively slight (+0.02 in terms of SPDSPD, and +0.03 in terms of AODAOD on average). For instance, on the Adult dataset, LIMI achieves 15.94% and 49.56% improvements of SPDSPD and AODAOD respectively; by contrast, retraining by NeuronFair and ExpGA even make the models more unfair between different groups on the Adult and Bank datasets. We speculate that the main reason is that the distribution of instances generated by LIMI is closer to the original dataset, while instances generated by NeuronFair and ExpGA deviate from the real distribution. We also observe that for small dataset Credit, the model overfits on the dataset and behaves even with no discrimination. Thus, directly retraining models using new discriminatory instances would not increase fairness on the Credit. We will further study it in the future.

Investigations and Analyses

Besides the main experiments conducted in the previous section, this section conducts further investigations and analyses to better understand our proposed framework in terms of fitness, stability, scalability, and generalizability.

Fitness. We first evaluate the fitness of our surrogate decision boundary to the target model through the Area Under Receiver Operating Characteristic Curve (AUC) metric (Hanley and McNeil, 1982), which measures the fitness in the class-imbalanced latent samples. Specifically, we evaluate our surrogate boundaries (i.e., linear SVM) both on the training set (100K latent samples with high classification confidence) and the entire set (1M random latent samples). The fitness results are shown in Table 5. Specifically, on Adult, Bank, and Meps datasets, our surrogate boundaries achieve over 93% AUC on the high-confidence training set, and over 84% AUC on the imbalanced entire set, which shows a proper approximation of the real decision boundary of the tested model. However, the AUC value is relatively lower on the Credit dataset, especially on the DNN model (67.82% for the entire latent samples). We attribute this observation to the small size of the Credit dataset (i.e., only 600 samples), which would cause an over-fitted DNN that fails to generalize well to a large number of new samples.

Stability. During the potential probing process, the walking stepsize λ\lambda along the unit vector is critical for finding the candidates. Thus, we hereby investigate the influence of hyperparameter λ\lambda on extra ablation studies. Specifically, we use our LIMI to test a DNN model within an hour and set λ\lambda as 0, 0.1, 0.2, 0.3, 0.4, and 0.5, respectively. As shown in Table 6, the performance of LIMI is relatively stable with only slight fluctuations for different λ\lambda values. Specifically, we observe that λ=0.3\lambda=0.3 generate 2 times more instances than λ=0\lambda=0, indicating the effectiveness of our candidate probing strategy; in the worst case setting (λ=0.5\lambda=0.5), #Didi\#\mathcal{D}_{idi} drops 18.61% on average, however, it still achieves comparable performance to other baselines (c.f. Table 1). Thus, we set λ=0.3\lambda=0.3 in our main experiments. Overall, our LIMI behaves stably when the hyperparameter λ\lambda falls in the range of [0.1,0.5][0.1,0.5]. Ulteriorly, we analyze the influence of latent boundary approximation, where we replace the approximated surrogate boundary with a randomly generated boundary. The average #Didi\#\mathcal{D}_{idi} value (56,637 for approximated boundary vs 8,418 for randomly generated boundary) on the Adult dataset, demonstrating the effectiveness of our boundary approximation module.

Scalability. In our main experiments, we directly compared our LIMI with the other two-phase baselines. Here, we further investigate the potential of our LIMI as a fast global prober. In other words, we use LIMI as a global seed generator to generate a seed set, and then combine it with other local methods (i.e., gradient-based method, genetic algorithm) to generate more instances around seeds. Specifically, we set the search number in the global phase as 40K (1K ×\times 40 iters), the global seed set constraint as 1K, the iteration in the local phase as 1K, and the target model is DNN. The results are shown in Figure 3, where “ADF+LIMI” indicates the test process consisting of a global phase of LIMI and a local phase of ADF. We could observe that, by combining our LIMI with other local methods, we could generate more discriminatory instances, i.e., LIMI enhances the generation quantity around 1.65 times on average. We conjecture the reason is that LIMI could find sufficient discriminatory instances in the global phase, therefore facilitating the local methods that are limited by insufficient seeds, which usually occurs when testing small datasets (e.g., Credit) or relying on an imperfect explainer.

Generalizability. Besides tabular data, we here intend to evaluate the generalizability of our approach to image data. Unlike tabular data, the attributes of the image (e.g., Gender, Smile, and BigLips) are difficult to modify directly from the input domain. Therefore, we slightly adjust the modification of protected attributes in our LIMI framework as follows: (1) we first approximate another linear hyperplane hph_{p} which separates the binary protected attribute (e.g., separating Gender to male and female); (2) Based on hph_{p}, we then modify the protected attribute of potential discriminatory candidates as follows

where wp\mathbf{w}_{p} and bpb_{p} represent the unit normal vector and the unit intercept of the hyperplane hph_{p} respectively. Moreover, (zi′(\mathbf{z}_{i}^{{}^{\prime}},zi)\mathbf{z}_{i}) forms a test case pair, which is different in the protected attribute.

In particular, we here test a ResNet-50 (He et al., 2016)) model trained on the CelebA dataset (Liu et al., 2015) for smile detection, utilize a pre-trained Progressive GAN (Karras et al., 2017) from GAN Zoo (HDGAN, 2021), and select the protected attribute as “gender”. During one hour of testing, LIMI obtains around 5K discriminatory samples that differ in gender from 50K test cases. As shown in Figure 4, we can observe that LIMI modifies the gender attribute obviously, and reveals the gender discrimination in the smile classifier which prefers giving smile predictions to the face of a female.

Threats to Validity

A well-trained GAN. Our LIMI relies on a well-trained GAN, such as CTGAN (Xu et al., 2019) and Progressive GAN (Karras et al., 2017), to probe sufficient test cases in the latent space. For tabular data, we can easily obtain a well-trained GAN in just a few minutes; while for image data, due to the high dimensionality and unstructured nature, it is relatively difficult and requires more time. However, it is convenient for testers to access off-the-shelf GANs for testing in the image domain since the pre-trained weights are widely open-sourced (HDGAN, 2021).

Protected attribute. We follow the most common settings (Aggarwal et al., 2019; Zhang et al., 2020b, 2021; Zheng et al., 2022; Fan et al., 2022; Udeshi et al., 2018) and only consider the single protected attribute for each fairness testing in our main experiment. However, testing for multiple protected attributes at a time will not hamper the performance of LIMI, but will certainly consume more time since all the possible combinations of the protected attributes perturbation need to be attempted. We here also study the combination of multiple protected attributes (gender&race, gender&age, race&age on Adult; gender&age on Credit), where our LIMI still outperforms others (average results over datasets are shown in Table 7). More detailed results can be found on our website (Anonym, 2023).

Generalization to textual data. In this paper, we follow the commonly-used settings (Udeshi et al., 2018; Aggarwal et al., 2019; Zhang et al., 2020b, 2021; Zheng et al., 2022) and primarily verify the effectiveness of our method on tabular data. Our method can be extended to other domains, such as the studies in Section 5 demonstrate the potential and generalizability to the image data. However, it is non-trivial to simply extend our current framework to textual data, since training GANs for text generation (Yu et al., 2017; Zhou et al., 2020) presents unique challenges owing to the sequential nature of textual data. The training challenge often results in mode collapse (Metz et al., 2016), thus the generator would produce limited output and fail to capture the full range of possible outputs resulting in the weak ability to generate natural samples. However, we are still interested in exploring the possibilities of LIMI in generalizing to Natural Language Processing (NLP) tasks in the future.

Related Work

Deep learning models easily exhibit some undesirable behaviors on concerns such as robustness, privacy, and other trustworthiness issues (Liu et al., 2023, 2019, 2021; Guo et al., 2023; Liu et al., 2020a, b; Zhang et al., 2020a; Wang et al., 2021; Tang et al., 2021). To reveal the discordance between existing and required fairness conditions of a given software system (Chen et al., 2022), a long line of work has been dedicated to performing model fairness testing such as test input generation (Xie and Wu, 2020; Black et al., 2020; Perera et al., 2022; Udeshi et al., 2018; Aggarwal et al., 2019; Zhang et al., 2020b, 2021; Zheng et al., 2022; Fan et al., 2022; Sharma et al., 2021; Díaz et al., 2018) and test oracle identification (Hardt et al., 2016; Barr et al., 2014; Rajan et al., 2022; Sharma and Wehrheim, 2019; Chakraborty et al., 2020; Hort et al., 2021; Chakraborty et al., 2021; Barocas and Selbst, 2016). In this paper, we primarily focus on generating individual discriminatory instances for machine learning model fairness testing, which can be roughly divided into white-box and black-box fairness testing based on the access to the target model.

For black-box fairness testing, testers have limited or even without any knowledge of the internal working of ML model (e.g., model architecture and gradients). Galhotra et al. (Galhotra et al., 2017) first formally defined software fairness and discrimination, then proposed a fairness test method named THEMIS, which randomly generates test cases to measure the software discrimination. However, the random test input generation in THEMIS is inefficient. To accelerate the generation, Udeshi et al. (Udeshi et al., 2018) proposed AEQUITAS, a two-phase search-based generating approach. In global search phase, AEQUITAS also randomly samples the input space to discover the discriminatory inputs as test seeds. While in local phase, AEQUITAS designs three different strategies to explore the neighborhood of global seeds systematically, which directs the probability of the attributes to be perturbed. Subsequently, Agarwal et al. (Aggarwal et al., 2019) presented SG, which combines symbolic execution and local explainability to detect individual discrimination. SG first utilizes the local explainer like LIME to construct a decision tree path to approximate the model decision process, and then leverages the symbolic execution to cover different tree paths to generate test cases. The ExpGA (Fan et al., 2022) also uses the interpretable model to search for high-quality test seeds in the global phase, and employs the genetic algorithm to generate a large amount of discriminatory offspring rapidly.

For white-box fairness testing, testers have complete knowledge of the target model and can fully access it. For example, Zhang et al. (Zhang et al., 2020b) proposed ADF, which firstly deals with the fairness testing problem of DNNs. ADF also contains two search phases, and requires the gradient information to search discriminatory instances near the decision boundary of DNN. Following the framework of ADF, Zhang et al. (Zhang et al., 2021) proposed EIDIG, achieving better efficiency by integrating a momentum term in global phase and exploiting the prior information of gradient in local phase. Later, Zheng et al. (Zheng et al., 2022) proposed NeuronFair, further improving the performance by only calculating the gradients of biased neurons that are identified by NeuronFair rather than the whole model. Due to the dependence on gradient information of model architecture, these methods cannot deal with the fairness testing of traditional ML models, which significantly limits their practical application.

Though these methods have shown progress in fairness testing, there is no guarantee that the generated instances are legitimate or natural. In other words, existing techniques consider the generated instances effective as long as they can flip the predicted outcome after changing protected attribute, which may fail to obey the real-world constraints resulting in unnatural test samples (e.g., extreme values like 10-year-old children are authorized with the loan (Chen et al., 2022)). In contrast, this paper proposes the LIMI framework to generate natural individual discriminatory instances by probing nearby the decision boundary for black-box ML model fairness testing.

We note that both DEEPJANUS (Riccio and Tonella, 2020) and LIMI share a similar motivation of exploring nearby the decision boundary for natural samples, however, there are fundamental differences in our approaches and purposes. Our LIMI is designed to derive a surrogate boundary in the latent space and probe potential discriminatory latent candidates, with the aim of generating discriminatory instances for individual fairness testing. In contrast, DEEPJANUS (Riccio and Tonella, 2020) employs an evolutionary algorithm within a model representation of the input domain to generate boundary cases for image classifiers, with the goal of characterizing the frontier of behaviors.

2. GAN-Based Software Testing

Due to the strong generative capability, a variety of works have been proposed to explore the application of GAN in software testing. Some researchers directly utilize existing GANs to facilitate the testing process. Zhang et al. (Zhang et al., 2018) utilized GAN to synthesize driving scenes with various weather conditions in the metamorphic testing of autonomous driving systems. Similarly, Gao and Han (Gao and Han, 2019) utilized DCGAN (Radford et al., 2015) and CycleGAN (Zhu et al., 2017) for style transferring in coverage testing for deep learning systems. Rather than the image domain, Guo et al. (Guo et al., 2022) attempted three kinds of GANs to learn the execution path information of software in order to generate test data that can achieve full test coverage. Other researchers propose variants of GAN to solve their problems specifically. Bao et al. (Bao et al., 2019) proposed ACTGAN to capture the hidden structures of good configurations and generate potentially better configurations, which accelerates the configuration tuning process to a large extent. Porres et al. (Porres et al., 2021) proposed an online GAN algorithm for automatic performance test generation, which can generate a high number of tests to reveal performance defects.

By contrast, this paper primarily focuses on black-box fairness testing, where we resort to the strong data-fitting ability of GANs to better generate natural individual discriminatory instances.

Conclusion

This paper proposes LIMI framework to generate natural individual discriminatory instances for fairness testing. LIMI first coarsely approximates the decision boundary of the tested model by deriving a surrogate linear boundary in the semantic latent space of GAN; LIMI then manipulates the random latent vectors to the surrogate boundary with a one-step movement and further conduct vector calculation to probe two potential discriminatory candidates on either side of it. Therefore, we could generate individual discriminatory instances closer to the real decision boundary and thus acquire better naturalness. Extensive experiments demonstrate that our LIMI can generate a larger number of natural discriminatory instances with a higher speed than 6 SOTA methods in 7 benchmarks. Moreover, the model fairness can be improved by retraining with the natural discriminatory instances generated by LIMI.

Acknowledgement. This work was supported by the National Key R&D Program of China (2022ZD0116310), the National Natural Science Foundation of China (62022009 and 62206009), and the State Key Laboratory of Software Development Environment.

References