CARLA: A Python Library to Benchmark Algorithmic Recourse and Counterfactual Explanation Algorithms
Martin Pawelczyk, Sascha Bielawski, Johannes van den Heuvel, Tobias Richter, Gjergji Kasneci
Introduction
Machine learning (ML) methods have found their way into numerous everyday applications and have become an indispensable asset in various sensitive domains, like disease diagnostics , criminal justice , or credit risk scoring . While ML models bear the great potential to provide effective support in human decision making processes, their predictions may have considerable impact on personal lives, where the final decisions might be disadvantageous for an end user. For example, the rejection of a loan or the denial of parole might have negative effects on the future development of the corresponding person’s life.
When ML systems involve humans in the loop, it is crucial to build a strong foundation for long-term acceptance of these methods. To this end, it is critical (1) to explain the predictions of a model and (2) to offer constructive means for the improvement of those predictions to the advantage of the end–user. Counterfactual explanations – popularized by the seminal work of – provide means for prescriptive model explanations by suggesting actionable feature changes (e.g., increase income) that allow individuals to achieve favourable outcomes in the future (e.g., insurance approval).
When counterfactual explainability is employed in systems that involve humans in the loop, the community refers to it as recourse. Algorithmic recourse subsumes precise recipes on how to obtain desirable outcomes after being subjected to an automated decision, emphasizing feasibility constraints that have to be taken into account. Those explanations are found by making the smallest possible change to an input vector to influence the prediction of a pretrained classifier in a positive way; for example, from ‘loan denial’ to ‘loan approval’, subject to the constraint that an individual’s sex may not change. As documented in recent reviews, there exists a quickly growing literature with available methods (see Figure 1 and ), reflecting the insight that the understanding of complex machine learning models is an elementary ingredient for a wide and safe technology adoption.
In practice, the counterfactual explanation (CE) that an individual receives crucially depends on the method that computes the recourse suggestions. Hence, there is a substantial need for a standardized benchmarking platform, which ensures that methods can be compared in a transparent and meaningful way. Researchers need to be able to easily evaluate their proposed methods against the overwhelming diversity of already available methods and practitioners need to make sure that they are using the right recourse mechanism for the problem at hand. Therefore, a standardized framework for comparison and quality assurance is an essential and indispensable prerequisite.
In this work, we present CARLA (Counterfactual And Recourse LibrAry), a python library with the following merits: First, CARLA provides competitive baselines for researchers to benchmark new counterfactual explanation and recourse methods for the standardized and transparent comparison of CE methods on different integrated data sets. Second, CARLA is a common framework with more than 10 counterfactual explanation methods in combination with the possibility to easily integrate new methods into a commonly accessible and easily distributable Python library. Moreover, the built-in integrated evaluation measures allow users to plug-in their custom black-box predictive models into the available counterfactual explanation methods and conduct extensive evaluations in comparison with other recourse mechanisms across different data sets. The same is true for researchers, who can use CARLA to extensively benchmark available counterfactual methods on popular data sets across various ML models. Third, CARLA supports popular optimization frameworks such as Tensorflow and PyTorch , and provides a generic abstraction layer to support custom implementations. Users can can define problem–specific data set characteristics like immutable features and explicitly specify hyperparameters for the chosen counterfactual explanation method.
The remainder of this work is structured as follows: Section 2 presents related work, Section 3 formally introduces the recourse problem, Section 4 presents the benchmarking process. In Section 5, we describe our main findings, before concluding in Section 6. Appendices A - E describe CARLA’s software architecture and usage instructions, as well as additional experimental results, used ML classifiers, data sets and hyperparameters settings.
Related Work
Explainable machine learning is concerned with the problem of providing explanations for complex ML models. Towards this goal, various streams of research follow different explainability paradigms which can be categorized into the following groups .
Local input attribution techniques seek to explain the behaviour of ML models instance by instance. Those methods aim to understand how all inputs available to the model are being used to arrive at a certain prediction. Some popular approaches for model explanations aim at explainability by design . For white-box models – the internal model parameters are known – gradient-based approaches, e.g. (for deep neural networks), and rule-based or probabilistic approaches for tree ensembles, e.g. have been proposed. In cases where the parameters of the complex models cannot be accessed, model-agnostic approaches can prove useful. This group of approaches seeks to explain a model’s behavior locally by applying surrogate models , which are interpretable by design and are used to explain individual predictions of black-box ML models.
2 Counterfactual Explanations
The main purpose of counterfactual explanations is to suggest constructive interventions to the input of a complex model so that the output changes to the advantage of an end user. By emphasizing both the feature importance and the recommendation aspect, counterfactual explanation methods can be further divided into three different groups: independence-based, dependence-based, and causality-based approaches.
In the class of independence-based methods, where the input features of the predictive model are assumed to be independent, some approaches use combinatorial solvers or evolutionary algorithms to generate recourse in the presence of feasibility constraints . Notable exceptions from this line of work are proposed by , who use decision trees, random search, support vector machines (SVM) and information networks that are aligned with the recourse objective. Another line of research deploys gradient-based optimization to find low-cost counterfactual explanations in the presence of feasibility and diversity constraints . The main problem with these approaches is that they abstract from input correlations. That implies that the intervention costs (i.e., the costs of changing the input to achieve the proposed counterfactual state) are too optimistically estimated. In other words, the estimated costs do not reflect the true costs that an individual would need to incur in practical scenarios, where feature dependencies are usually present: e.g., income is dependent on tenure, and if income changes, tenure also changes (see Figure 2 for a schematic comparison).
In the class of causality-based approaches, all methods make use of Pearl’s causal modelling framework . As such, they usually require knowledge of the system of causal structural equations or the causal graph . The authors of show that these models can generate minimum-cost recourse, if the access to the true causal data generating process was available. However, in practical scenarios, the guarantee for such minimum-cost recommendations is vacuous, since, in complex settings, the causal model is likely to be miss-specified . Since these methods usually require the true causal graph – which is the limiting factor in practice – we have not considered them at this point, but we plan to do that in the future.
Dependence-based methods bridge the gap between the strong independence assumption and the strong causal assumption. This class of models builds recourse suggestions on generative models . The main idea is to change the geometry of the intervention space to a lower dimensional latent space, which encodes different factors of variation while capturing input dependencies. To this end, these methods primarily use variational autoencoders (VAE) . In particular, Mahajan et al. demonstrate how to encode various feasibility constraints into VAE-based models. Most recently, proposed CLUE, a generative recourse model that takes a classifier’s uncertainty into account. Work that deviates from this line of research was done by . The authors of provide FACE, which uses a shortest path algorithm on graphs to find counterfactual explanations. In contrast, Kanamori et al. use integer programming techniques to account for input dependencies.
Preliminaries
In this Section, we review the algorithmic recourse problem and draw a distinction between two observational (i.e., non–causal) methods.
Assuming inputs are pairwise statistically independent, the recourse problem is defined as follows:
where is the set of admissible changes made to the factual input . For example, could specify that no changes to sensitive attributes such as age or sex may be made. For example, using the independent input assumption, existing approaches use mixed-integer linear programming to find counterfactual explanations. In the next paragraph, we present a problem formulation that relaxes the strong independence assumption by introducing generative models.
2 Recourse for Correlated Inputs
where is the set of admissible changes in the -dimensional latent space. For example, would ensure that the counterfactual latent space lies within range of . The problem in () is an abstraction from how the problem is usually solved in practice: most existing approaches first train a type of autoencoder model (e.g., a VAE), and then use the model’s trained decoder as a deterministic function to find counterfactual explanations . Our benchmarked explanation models roughly fit in one of these two categories.
Benchmarking Process
In this Section, we provide a brief explanation model overview and introduce a variety of explanation measures used to evaluate the quality of the generated counterfactual explanations. In Table 1 we present a concise explanation model overview.
Ustun et al. provide a method to generate minimal cost actions for linear classification models such as logistic regression models. AR requires the linear model’s coefficients, and uses these coefficients for its search for counterfactual explanations. To provide reasonable actions it is possible to restrict to user–specified constraints (e.g., has_phd can only change from False to True) or to set a subset of inputs as immutable (e.g., age). The problem to find these changes is a discrete optimization problem. Given a set of actions, AR finds the action which minimizes a defined cost function, using integer programming solvers like CPLEX or CBC.
AR--LIME (I)I\mathbf{(I)}
Most classification tasks do not have linearly separable classes and complex non–linear models usually provide more accurate predictions. Non–linear models are not per se interpretable and usually do not provide coefficients similar to linear models. We use a reduction to apply AR to non–linear models by computing a local linear approximation for the point of interest , using LIME . For an arbitrary black–box model , LIME estimates post–hoc local explanations in form of a set of linear coefficients per instance. Using the coefficients we apply AR.
CEM (I)I\mathbf{(I)}
CLUE (D)D\mathbf{(D)}
Antorán et al. propose CLUE, a generative recourse model that takes a classifier’s uncertainty into account. This model suggests feasible counterfactual explanations that are likely to occur under the data distribution. The authors use a variational autoencoder (VAE) to estimate the generative model. Using the VAE’s decoder, CLUE uses an objective that guides the search of CEs towards instances that have low uncertainty measured in terms of the classifier’s entropy.
DICE (I)I\mathbf{(I)}
Mothilal et al. suggest DICE, which is an explanation model that seeks to generate minimum costs counterfactual explanations according to () subject to a diversity constraint which aims to promote a diverse set of counterfactual explanations. Diversity is achieved by using the whole range of suggested changes, while still keeping proximity to a given input. Regarding the optimization problem, DICE uses gradient descent to find a solution that trades-off proximity and diversity. Domain knowledge – in form of feature ranges or immutability constraints – can be added.
FACE (D)D\mathbf{(D)}
The authors of provide FACE, which uses a shortest path algorithm (for graphs) to find counterfactual explanations from high–density regions. Those explanations are actual data points from either the training or test set. Immutability constraints are enforced by removing incorrect neighbors from the graph. We implemented two variants of this model: the first variant uses an epsilon–graph (FACE--EPS), whereas the second variant uses a knn–graph (FACE--KNN).
Growing Spheres (GS) (I)I\mathbf{(I)}
Growing Spheres – suggested in – is a random search algorithm, which generates samples around the factual input point until a point with a corresponding counterfactual class label was found. The random samples are generated around using growing hyperspheres. For binary input dimensions, the method makes use of Bernoulli sampling. Immutable features are readily specified by excluding them from the search procedure.
REVISE (D)D\mathbf{(D)}
Joshi et al. propose a generative recourse model. This model suggests feasible counterfactual explanations that are likely to occur under the data distribution. The authors use a variational autoencoder (VAE) to estimate the generative model. Using the VAE’s decoder, REVISE uses the latent space to search for CEs. No handling of immutable features exists.
Wachter et al. (Wachter) (I)I\mathbf{(I)}
2 Evaluation Measures for Counterfactual Explanation Methods
As algorithmic recourse is a multi–modal problem we introduce a variety of measures to evaluate the methods’ performances. We use six baseline evaluation measures. Besides distance measures it is important to consider measures that emphasize the quality of recourse.
Constraint violation
This measure counts the number of times the CE method violates user-defined constraints. Depending on the data set, we fixed a list of features which should not be changed by the used method (e.g., sex, age or race).
yNN
We use a measure that evaluates how much data support CEs have from positively classified instances. Ideally, CEs should be close to positively classified individuals which is a desideratum formulated by Laugel et al. . We define the set of individuals who received an undesirable prediction under as . The counterfactual instances (instances for which the label was successfully changed) corresponding to the set are denoted by . We use a measure that captures how differently neighborhood points around a counterfactual instance are classified:
Redundancy
Success Rate
Some generated counterfactual explanations do not alter the predicted label of the instance as anticipated. To keep track how often the generated CE does hold its promise, the success rate shows the fraction of respective models’ correctly determined counterfactuals.
Average Time
By measuring the average time a CE method needs to generate its result, we evaluate the effectiveness and feasibility for real–time prediction settings.
Experimental Evaluation
Using CARLA we conduct extensive empirical evaluations to benchmark the presented counterfactual explanations methods using three real-world data sets. Our main findings are displayed in Figure 3, and Table 2. We split the benchmarking evaluation by CE method category. In the following Sections, we provide an overview over the used data sets (see Table 3) and the classification models. Detailed information on hyperparameter search for the CE methods is provided in Appendix E.
The Adult data set originates from the 1994 Census database, consisting of 14 attributes and 48,842 instances. The classification consists of deciding whether an individual has an income greater than 50,000 USD/year. Since several CE methods cannot handle non-binary categorical data, we binarized these features by partitioning them into the most frequent value, and its counterpart (e.g., US and Non-US, Husband and Non-Husband). The features age, sex and race are set as immutable. The Give Me Some Credit (GMC) data set from a 2011 Kaggle Competition is a credit scoring data set, consisting of 150,000 observations and 11 features. The classification task consists of deciding whether an instance will experience financial distress within the next two years (SeriousDlqin2yrs is 1) or not. We dropped missing data, and set age as immutable.
Black-box models
We briefly describe how the black–box classifiers were trained. CARLA supports different ML libraries (e.g., Pytorch, Tensorflow) to estimate these classifiers as the implementations of the various explanation methods work particular ML libraries only. The first model is a multi-layer perceptron, consisting of three hidden layers with 18, 9 and 3 neurons, respectively. To allow a more extensive comparison (AR only works on linear models) between CE methods, we chose logistic regression models as the second classification model for which we evaluate the CE methods. Detailed information on the classifiers’ training for each data set is provided in Appendix C.
Benchmarking
Conclusion and Broader Impact of CARLA
The current implementations of recourse methods, mentioned in Section 4.1 are based on the original implementation of the respective research groups. Researchers mostly implement their experiments and models for specific ML frameworks and data sets. For example, some explanation methods are restricted to Tensorflow and are not applicable to Pytorch models. In the future, we will extend CARLA to decouple each recourse method from the frameworks and data contraints.
When trying to combine different CE methods into a common benchmarking framework we encountered the following issues: First, a great number of repositories only contain remarks about installation and script calls to recreate the results from the corresponding research papers. Second, missing information about interfaces for data sets or black–box models further complicated the process of integrating different CE methods into the benchmarking workflow. In order to add more CE methods and data sets to CARLA, we are currently in contact with several authors in this exciting and rapidly growing field. With a growing open-source community, CARLA can evolve to be the main library for generating counterfactual explanations and benchmarks for recourse methods. Therefore we are continuously expanding the catalog of explanation methods and data sets, and welcome researchers to add their own recourse methods to the library. To facilitate this process, we provide a step-by-step user-guide to integrate new CE methods into CARLA, which we present in Appendix A.
The rapidly growing number of available CE methods calls for standardized and efficient ways to assure the quality of a new technique in comparison with other approaches on different data sets. Quality assurance is a key aspect of actionable recourse, since complex models and CE mechanisms can have a considerable impact on personal lives. In this work, we presented CARLA, a versatile benchmarking platform for the standardized and transparent comparison of CE methods on different integrated data sets. In the explainability field, CARLA bears the potential to help researchers and practitioners alike to efficiently derive more realistic and use–case–driven recourse strategies and assure their quality through extensive comparative evaluations. We hope that this work contributes to further advances in explainability research.
References
Checklist
The checklist follows the references. Please read the checklist guidelines carefully for information on how to answer these questions. For each question, change the default [TODO] to [Yes] , [No] , or [N/A] . You are strongly encouraged to include a justification to your answer, either by referencing the appropriate section of your paper or providing a brief inline description. For example:
Did you include the license to the code and datasets? [Yes] See Section LABEL:gen_inst.
Did you include the license to the code and datasets? [No] The code and the data are proprietary.
Did you include the license to the code and datasets? [N/A]
Please do not modify the questions and only use the provided macros for your answers. Note that the Checklist section does not count towards the page limit. In your paper, please delete this instructions block and only keep the Checklist section heading above along with the questions/answers below.
Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes] . As we state in the abstract, our goal is to provide a Python framework for benchmarking counterfactual explanation methods. Users can easily evaluate our results by accessing our Github repository, where we host our Python framework and our benchmarking results.
Did you describe the limitations of your work? [Yes] . In Section 6, we discuss the current limitations of our approach. The counterfactual explanation methods are based on the original implementation of the respective research groups. Researchers mostly implement their experiments and models for specific ML frameworks and data sets. For example, some explanation methods are restricted to Tensorflow and are not applicable to Pytorch models.
Did you discuss any potential negative societal impacts of your work? [N/A] . We discuss the broader impact of our benchmarking library in Section 6; we mainly see positive impacts on the literature of algorithmic recourse.
Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes] . We have read the ethics review guidelines and attest that our paper conforms to the guidelines.
If you are including theoretical results…
Did you state the full set of assumptions of all theoretical results? [N/A] . We did not provide theoretical results.
Did you include complete proofs of all theoretical results? [N/A] . We did not provide theoretical results.
Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] . Details of implementations, data sets and instructions can be found here: Appendices A, C, E, and our Github repository.
Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] . Please see Appendices E and C.
Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [Yes] . Error bars have been reported for our cost comparisons in terms of the 25th and 75ht percentiles of the cost distribution, see for example Figure 3.
Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] . All models are evaluated on an i7-8550U CPU with 16 Gb RAM, running on Windows 10.
If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…
If your work uses existing assets, did you cite the creators? [Yes] . The data sets, which are publicly available are appropriately cited in Section 5. We cite and link to any additional code used, for example .
Did you mention the license of the assets? [Yes] . All assets are publicly available and attributed.
Did you include any new assets either in the supplemental material or as a URL? [Yes] . Our implementation and code is accessible through our Github repository.
Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A] . We use publicly available data sets without any personal identifying information.
Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A] . We use publicly available data sets without any personal identifying information.
If you used crowdsourcing or conducted research with human subjects…
Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A] . We did not use crowdsourcing or conduct research with human subjects.
Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A] . We did not use crowdsourcing or conduct research with human subjects.
Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A] . We did not use crowdsourcing or conduct research with human subjects.
Appendix A CARLA’s Software Interface
In the following, we introduce our open-source benchmarking software CARLA. we describe the architecture in more detail and provide examples of different use-cases and their implementation.
The purpose of this Python library is to provide a simple and standardized framework to allow users to apply different state-of-the-art recourse methods to arbitrary data sets and black-box-models. It is possible to compare different approaches and save the evaluation results, as described in Section 4.2. For research groups, CARLA provides an implementation interface to integrate new recourse methods in an easy-to-use way, which allows to compare their method to already existing methods.
A simplified visualization of the CARLA software architecture is depicted in Figure 4. For every component (Data, MLModel, and RecourseMethod) the library provides the possibility to use existing methods from our catalog, or extend the users custom methods and implementations. The components represent an interface to the key parts in the process of generating counterfactual explanations. Data provides a common way to access the data across the software and maintains information about the features. MLModel wraps each black-box model and stores details on the encoding, scaling and feature order specific to the model. The primary purpose of RecourseMethod is to provide a common interface to easily generate counterfactual examples.
Besides the possibility to use pretrained black-box-models and preprocessed data, CARLA provides an easy way to load and define own data sets and model structures independent of their framework (e.g., Pytorch, Tensorflow, sklearn). The following sections will give an overview and provide example implementations of different use cases.
A.2 CARLA for Research Groups
One of the most exciting features of CARLA is, that research groups can make use of the RecourseMethod-wrapper to implement their own method to generate counterfactual examples. This opens up a way of standardized and consistent comparisons between different recourse methods. Strong and weak points of new algorithms can be stated, benchmarked and analysed in forthcoming publications with the help of CARLA.
In Figure 5, we show how an implementation of a custom recourse method can be structured. After defining the recourse method in the shown way, it can be used with the library to generate counterfactuals for a given data set and benchmark its results against other methods. Research groups have the choice to do this using our provided catalog of data sets, recourse methods and black-box models (Figure 6) or use their own models and data sets (see Figures 7 and 8).
A.3 CARLA as a Recourse Library
A common usage of the package is to generate counterfactual examples. This can be done by loading black-box-models and data sets from our provided catalogs, or by user-defined models and datasets via integration with the defined interfaces. Figure 6 shows an implementation example of a simple use-case, applying a recourse method to a pre-defined data set and model from our catalog. After importing both catalogs, the only necessary step is to describe the data set name (e.g., adult, give me some credit, or compas) and the model type (e.g., ann, or linear) the user wants to load. Every recourse method contains the same properties to generate counterfactual examples.
To give users the possiblity to explore their own black-box-model on a custom data set, we implemented in CARLA easy-to-use interfaces, that are able to wrap every possible model or data set. These interfaces specify particular properties users have to implement, to be able to work with the library. Figure 7 shows an example implementation of the data wrapper, and Figure 8 depicts the same for an arbitrary black-box-model. After defining data set and black-box model classes, users simply need to call the canonical methods and generate counterfactual examples, similar to the process in Figure 6.
A.4 Benchmarking Recourse Methods
Besides the generation of counterfactual examples, the focus of CARLA lies on benchmarking recourse methods. Users are able to compute evaluation measures to make qualitative statements about usability and applicability.
All measurements, which are described in Section 4.2, are implemented in the Benchmarking class of CARLA and can be used for every wrapped recourse method. Figure 9 shows an example implementation of a benchmarking process based on the variables of Figure 6.
Appendix B Additional Experimental Results
In this Section, we depict the missing experiments from the COMPAS data set in Figure 11 and Table 4. These results underline the trends that we have already highlighted in Section 5.
Appendix C ML Classifiers
In this section, we describe how the black–box models were fitted. CARLA supports different ML libraries to estimate these models (e.g., Pytorch, Tensorflow) as the implementations of the various explanation methods work with a particular ML library. We note that the various explanation methods rely on different binary feature encodings. DICE, for example, requires that binary inputs are supplied as one–hot vectors, while FACE needs binary features encoded in a single column. If this was the case, we fitted two ML models, using the same hyperparameters, and generated CEs with respect to the same set of samples.
To ensure similar behavior between the different ML libraries and encoding variations, each black-box model type has the same structure (e.g., number of hidden layer, number of neurons), and training parameters (e.g., learning rate, epochs, etc.).
The first model is a multi-layer perceptron, consisting of three hidden layers with 18, 9 and 3 neurons, respectively. We use ReLu activation functions and binary cross entropy to calculate class probabilities. Optimization of the loss function is done by RMSProp using a learning rate of 0.002 for every data set. By performing 25 epochs on COMPAS and 10 epochs on Adult and GMC we reached acceptable performance. Further increasing epochs gave rise to very marginal performance increases. For Adult we use a batch–size of 1024, for COMPAS 25 and for GMC 2048.
To allow a more extensive comparison between CE methods, we choose linear models as the second black–box model category for which we evaluate the CE methods. Again, we optimized these models with RMSProp using a binary cross entropy loss. For Adult, we used 100 epochs and a batch–size of 2048, for COMPAS we choose 25 epochs and batch–size of 128, and for GMC we chose 10 epochs with a batch–size of 2048. The learning rate on every data set is set 0.002. Table 5 provides an overview of the model’s classification accuracies.
Appendix D COMPAS Data Set Description
The COMPAS data set contains data for more than 10,000 criminal defendants in Florida. It is used by the jurisdiction to score defendant’s likelihood of reoffending. We kept a small part of the raw data as features like name, id, casenumbers or date-time were dropped. The classification task consists of classifying an instance into high risk of recidivism (score_text is high). By converting the feature race into white and non-white, we keep the categorical input binary. Similar to Adult, the immutable features for COMPAS are age, sex and race.
Appendix E Hyperparameter Search for the Counterfactual Explanation and Recourse Methods
We generated counterfactual explanations for instances from , the set of factuals with negative class predictions.
It frequently occurred that the action with the lowest cost did not flip the prediction of the black-box classifier. To overcome this problem, we let AR compute a flipset of 150 actions per instance, and subsequently search this set for low–cost CEs. For AR--LIME, we used LIME and required sampling around the instance to make sure that the coefficients at were truly local.
CEM
CLUE
We use the default hyperparameters from , which are set as a function of the data set dimension . Performing hyperparameter search did not yield results that were improving distances while keeping the same success rate.
DICE
Since DICE is able to compute a set of counterfactuals for a given instance, we only chose to generate one CE per input instance. We use a grid search for the proximity and diversity weights.
FACE
To determine the strongest hyperparameters for the graph size we conducted a grid search. We found that values of gave rise to the best balance of success rate and costs. For the epsilon graph, a radius of 0.25 yields the strongest results to balance between high yNN and low cost.
GS
We chose 0.02 as the step size with which the sphere is grown. Lower values yield similar results at the costs of higher computational time, while higher values gave worse results.
REVISE
The grid search to find an acceptable learning rate and similarity weight yielded and for about 1500 iterations.