Diffusion Visual Counterfactual Explanations

Maximilian Augustin, Valentyn Boreiko, Francesco Croce, Matthias Hein

Introduction

It can be argued that one of the main problems hindering the widespread use of machine learning and image classification in particular, is the missing possibility to explain the decisions of black-box models such as neural networks. This is not only a problem for decisions affecting humans where the current draft for AI regulation in Europe requires “transparency”, but it is a pressing problem in all applications of machine learning. The reason, to some extent, is that humans would like to understand but also control if the learning algorithm has captured the “concepts” of the underlying classes or if it just predicts well using spurious features, artefacts in the data set, or other sources of error. In this paper, we focus on model-agnostic explanations which can, in principle, be applied to any image classifier and do not rely on the specific structure of the classifier such as decision trees or linear classifiers. In this area, in particular for image classification, sensitivity based explanations , explanations based on feature attributions , saliency maps , Shapley additive explanations , and local fits of interpretable models have been proposed. Moreover, proposed counterfactual explanations (CEs), which are instance-specific explanations. They can be applied to any classifier and it has been argued that they are close to the human justification of decisions using counterfactual reasoning: “I would recognize it as zebra (instead of horse) if it had black and white stripes.” For a given classifier almost all methods for the generation of CEs try to solve the following problem: “Given a target class cc what is the minimal change δ\delta of input xx, such that x+δx+\delta is classified as class cc with high probability and is a realistic instance of my data generating distribution?” From the perspective of “debugging” existing machine learning models, CEs are interesting as they construct a concrete input x+δx+\delta with a different classification that allows the developer to test if the model has learned the correct features.

The main reason why visual counterfactual explanations (VCEs), that is CEs for image classification, are not widely used is that the tasks of generating CEs and adversarial examples are very related. Even imperceivable changes of the image can already change the prediction of the classifier, however, the resulting noise patterns do not show the user if the classifier has picked up the right class-specific features. One possible solution is using adversarially robust models , which have been shown to produce semantically meaningful VCEs by directly maximizing the probability of the target class in image space. These approaches have the downside that they only can generate VCEs for robust models which is a significant restriction as these models are not competitive in terms of prediction accuracy. The second approach is to restrict the generation of VCEs using a generative model or constraining the set of potential image manipulations . However, these approaches are either restricted to datasets with a small number of classes, cannot provide explanations for arbitrary classifiers, or generate VCEs that look realistic but have so little in common with the original image that not much insight can be gained.

Recently, trained a StyleGAN2 model to discover and manipulate class attributes. While their approach yields impressive results on smaller-scale datasets with few similar classes (for example, different bird species), the authors did not demonstrate that the method scales to complex tasks such as ImageNet with hundreds of classes where different classes require different sets of attributes. Another disadvantage is that the StyleGAN model needs to be retrained for every classifier, making it prohibitively expensive to explain multiple large models. Moreover, have proposed a loss to do VCEs in the latent space of a GAN/VAE. They show promising results for MNIST/FMNIST, but no code is available for ImageNet. Furthermore, to explain a classifier, the conditional information during the generation of explanations should ideally come only from the classifier itself, but rely on a conditional GAN which might introduce a bias.

In this paper, we overcome the aforementioned challenges and generate our Diffusion Visual Counterfactual Explanations (DVCEs) for arbitrary ImageNet classifiers (see Fig 1). We use the progress in the generation of realistic images using diffusion processes , which recently were able to outperform GANs . Similar to , we use a classifier and a distance-type regularization to guide the generation of the images. Two modifications to the diffusion process are key elements for our DVCEs: i) a combination of distance regularization and starting point of the diffusion process together with an adaptive reparameterization lets us generate VCEs visually close to the original image in a controlled way so that hyperparameters can be fixed across images and even across models, ii) our cone regularization of the gradient of the classifier via an adversarially robust model ensures that the diffusion process does not converge to trivial non-semantic changes but instead produces realistic images of the target class which achieve high confidence by the classifier. Our approach can be employed for any dataset where a generative diffusion model and an adversarially robust model are available. In a qualitative comparison, user study, and a quantitative evaluation (Sec. 4.1), we show that our DVCEs achieve higher realism (according to FID) and have more meaningful features (according to the user study) than both the recent methods of and .

Diffusion models

Diffusion models are generative models that consist of two steps: a forward diffusion process (and thus a Markov process) that transforms the data distribution to a prior distribution (which is usually assumed to be a standard normal distribution) and the reverse diffusion process that transforms the prior distribution back to the data distribution. The existence of the (continuous-time) reverse diffusion process was closer investigated in . In the discrete-time setting, a Markov chain {x1,...,xT}\{x_{1},...,x_{T}\} for any data point x0x_{0} in the forward direction is defined via a Markov process that adds noise to the data point at each timestep:

where {β1,...,βT}\{\beta_{1},...,\beta_{T}\} is some variance schedule, chosen such that q(xT)∼N(0,I)q(x_{T})\sim\mathcal{N}(0,I). Note that given x0x_{0}, it is possible to sample from q(xt∣x0)q(x_{t}|x_{0}) in closed form instead of applying noise tt times using:

where α‾t:=∏k=1t(1−βk)\overline{\alpha}_{t}:=\prod_{k=1}^{t}(1-\beta_{k}). In , the authors have shown that the reverse transitions q(xt−1∣xt)q(x_{t-1}|x_{t}) approach diagonal Gaussian distributions as T→∞T\rightarrow\infty and thus one can use a DNN with parameters θ\theta to approximate q(xt−1∣xt)q(x_{t-1}|x_{t}) as pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) by predicting the mean and the diagonal covariance matrix μθ(xt,t)\mu_{\theta}(x_{t},t) and Σθ(xt,t)\Sigma_{\theta}(x_{t},t):

To sample from pθ(x0)p_{\theta}(x_{0}), one samples xTx_{T} from N(0,I)\mathcal{N}(0,I) and then follows the reverse process by repeatedly sampling from the transition probabilities pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}).

Instead of predicting μθ(xt,t)\mu_{\theta}(x_{t},t) directly, it has been shown in that the best performing parameterization uses a neural network ϵθ(xt,t)\epsilon_{\theta}(x_{t},t) to approximate the source noise ϵ\epsilon in (3) and the loss used for training resembles that of a denoising model:

The mean of the reverse step μθ\mu_{\theta} in (4) can be derived using Bayes theorem as:

The issue is that efficient sampling from the original diffusion model is possible only because the reverse process is made of normal distributions. As we need to sample from pθ,ϕ(xt−1∣xt,y)p_{\theta,\phi}(x_{t-1}|x_{t},y) hundreds of times to obtain a single sample from the data distribution, it is not possible to use MCMC-samplers with high complexity to sample from each of the individual transitions. In , they proposed to solve it by approximating pθ,ϕ(xt−1∣xt,y)p_{\theta,\phi}(x_{t-1}|x_{t},y) with slightly shifted versions of pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) to make closed-form sampling possible. Such transition kernels are given by:

which we further adapt for the goal of generating VCEs and use in our experiments.

Diffusion Visual Counterfactual Explanations

A VCE xx for a chosen target class yy, a given classifier pϕ(y∣⋅)p_{\phi}(y|\cdot), and an input x^\hat{x} should satisfy the following criteria: i) validity: the VCE xx should be classified by pϕ(y∣⋅)p_{\phi}(y|\cdot) as the desired target class yy with high predicted probability, ii) realism: the VCE should be as close as possible to a natural image, iii) minimality/closeness: the difference between the VCE xx and the original image x^\hat{x} should be the minimal semantic modification necessary to change the class, in particular, the generated image xx should be close to x^\hat{x} while being valid and realistic, e.g. by changing the object in the image and leaving the background unchanged. Note that targeted adversarial examples are valid but do not show meaningful semantic changes in the target class for a non-robust model and are not realistic.

The l1.5l_{1.5}-SVCEs of change the image in order to maximize the predicted probability of the classifier into the target class inside an l1.5l_{1.5}-ball around the image which is a targeted adversarial example. Thus this only works for robust classifiers and they use an ImageNet classifier that was trained to be multiple-norm robust (MNR), which we denote in this paper as MNR-RN50 (see Sec. 4.2). The realism of the l1.5l_{1.5}-SVCEs comes purely from the generative properties of robust classifiers , which can lead to artefacts. In contrast, our Diffusion Visual Counterfactual Explanations (DVCEs) work for any classifier and our DVCEs are more realistic due the better generative properties of diffusion models. An approach similar to our DVCE framework is Blended Diffusion (BD) which manipulates the image inside a masked region. One can adapt BD for the generation of VCEs by using as mask the whole image. DVCE and BD share the same diffusion model, but BD cannot be applied to arbitrary classifiers and requires image-specific hyperparameter tuning, see Fig. 2.

which adapts to the predicted mean of the diffusion model that we use to change μt\mu_{t} in (10) to

This adaptive parameterization allows for fine-grained control of the influence of the classifier and distance regularization so that now the hyperparameters CcC_{c} and CdC_{d} have the same influence across images and even classifiers. It facilitates the generation of DVCEs as otherwise hyperparameter finetuning would be necessary for each image as in BD, see Fig. 2 for a comparison. However, even with our adaptive parameterization, it is still not easy to produce semantically meaningful changes close to x^\hat{x} as can be seen in App. B.2. Thus, as in , we vary the starting point of the diffusion process and observe in App. B.2 that starting from step T2\frac{T}{2} of the forward diffusion process, together with the adaptive parameterization and using as the distance the l1l_{1}-distance, provides us with sparse but semantically meaningful changes. In our experiments, we set T=200T=200.

2 Cone Projection for Classifier Guidance

3 Final Scheme for Diffusion Visual Counterfactuals

Our solution for a non-adversarially robust classifier pϕ(y∣⋅)p_{\phi}(y|\cdot) is to use Algorithm 1 of by replacing the update step with:

Experiments

In this section, we evaluate the quality of the DVCE. We compare DVCE to existing works in Sec. 4.1. In Sec. 4.2, we compare DVCEs for various state-of-the-art ImageNet models and show how DVCEs can be used to interpret differences between classifiers. For our DVCEs, we use for all experiments the fixed hyperparameters given in Sec. 3.3. The diffusion model used for DVCE is taken from and has been trained class-unconditionally on 256x256 ImageNet images using a modified UNet architecture. The user study, discovery of spurious features using VCEs, further experiments, and ablations are in the appendix.

We compare DVCEs with VCEs produced by BD (BDVCEs) and l1.5l_{1.5}-SVCEs. As the latter only works for adversarially robust classifiers, we use the multiple-norm robust ResNet50 from , MNR-RN50, as the classifier to create VCEs for all three methods. As this model is robust on its own, we do not use the cone projection for its DVCEs.

First, in Fig. 4 we present a qualitative evaluation, where we transform one image into two different classes that are close to the true one in the WordNet hierarchy. The radius of the l1.5l_{1.5}-ball of is chosen as the smallest r∈{50,75,100,150}r\in\{50,75,100,150\} such that the confidence in the target class is larger than 0.900.90 per image. For BD, we select the image with the smallest classifier and regularization weight that reaches confidence larger than 0.9 from the set of parameters discussed in Sec. 3.3. If 0.9 is not reached by any setting for one of the two baselines, we show the image that achieves the highest confidence. As Fig. 4 shows, DVCE is the only method that satisfies all desired properties of VCEs. For example, for “mashed potato”, DVCE preserves the bowl and only changes the content into either guacamole or carbonara. Our qualitative comparison shows that the same hyperparameter setting of DVCE can handle different classes and transfer between similar classes, such as different snakes or wolf types, as well as different object sizes, e.g. cheetah and snake.

In contrast, both l1.5l_{1.5}-SVCEs and BDVCEs require different hyperparameters for different images to achieve high confidence in the target class. Even with the six parameter configurations, BD is not able to always produce images with high confidence in the target class. More problematic for VCEs is that the resulting images can have high confidence but are neither realistic (cheetah →\rightarrow tiger) nor resemble the original image at all (dingo →\rightarrow timber wolf/white wolf). Even if the method works with the given parameters, for example, mashed potato →\rightarrow guacamole or carbonara, the overall image quality cannot match that of DVCE as often the images contain overly bright colors. In the case of the volcano VCE, DVCE shows class features like lava whereas BDVCE can not clearly be labeled as a volcano. For the l1.5l_{1.5}-SVCEs, one often needs large radii to achieve the desired confidence of 0.90, which often results in images that do not look realistic.

As noted in , a quantitative analysis of VCEs using FID scores is difficult as methods not changing the original image have low FID score. Thus, we have developed a cross-over evaluation scheme, where one partitions the classes into two sets and only analyzes cross-over VCEs (more details are in App. E). We show the results in Tab. 1. In terms of closeness, DVCEs are worse than l1.5l_{1.5}-SVCEs, which can be expected as they optimize inside an l1.5l_{1.5}-ball. However, DVCE are the most realistic ones (FID score) and have similar validity as l1.5l_{1.5}-SVCEs. BDVCEs are the worst in all categories.

Moreover, we have conducted a user study (2020 users), in which participants decided if the changes of the VCE are meaningful or subtle and if the generated image is realistic, see App. C for details. The percentage of total images having the three different properties is (order: DVCE/l1.5l_{1.5}-SVCE/BD): meaningful - 62.0%, 48.4%, 38.7%; realism - 34.7%, 24.6%, 52.2%; subtle - 45.0%, 50.6%, 31.0%. This confirms that DVCEs generate more meaningful features in the target classes. While the result regarding realism seems to contradict the quantitative evaluation, this is due to fact that realism means that the user considered the image realistic irrespectively if it shows the target class or not.

2 Model comparison

Non-robust models. In Fig. 5 we show that DVCEs can be generated for various state-of-the-art ImageNet models. We use a Swin-TF, a ConvNeXt and a Noisy-Student EfficientNet . Both the Swin-TF and the ConvNeXt are pretrained on ImageNet21k whereas the EfficientNet uses noisy-student self-training on a large unlabeled pool. As all models do not yield perceptually aligned gradients, we use the cone projection with 30∘30^{\circ} angle described in Sec. 3.2 with the robust model from the previous section. All other parameters are identical to the previous experiment. We highlight that DVCE is the first method capable of explaining (in the sense of VCEs) arbitrary classifiers on a task as challenging as ImageNet. Overall, DVCEs satisfy all desired properties of VCEs. It also allows us to inspect the most important features of each model and class. For the stupa and church classes, for example, it seems like all models use different roof and tower structures as the most prominent feature as they spent most of their budget on changing those.

Robust models. Here, we evaluate the DVCEs of different robust models trained to be multiple-norm adversarially robust. They are generated by multiple norm-finetuning an initially lpl_{p}-robust model to become robust with respect to l1l_{1}-, l2l_{2}- and l∞l_{\infty}-threat model, specifically an l2l_{2}-robust ResNet50 resulting in MNR-RN50 used in the revious sections, an l∞l_{\infty}-robust XCiT-transformer model called MNR-XCiT in the following and an l∞l_{\infty}-robust DeiT-transformer called MNR-DeiT. Their multiple-norm robust accuracies can be found in . In , different lpl_{p}-ball adversarially robust models were compared for the generation of VCEs and they showed that for their l1.5l_{1.5}-SVCEs the multiple norm robust model was better (both in terms of FID and qualitatively) than individual lpl_{p}-norm robust classifiers. This is the reason why we also use a multiple-norm robust model for the cone-projection of a non-robust classifier. In Fig. 6, we show DVCEs of the three different classes for two examples images and two target classes each. All of the multiple-norm robust models have similarly good DVCEs showing classifier-specific variations in the semantic changes. This proves again the generative properties of adversarially robust models, in particular of robust transformers. More examples and further experiments are in App. B.4.

Limitations, Future Work and Societal impact

In comparison to GANs, diffusion-based approaches can be expensive to evaluate as they require multiple iterations of the reverse process to create one sample. This can be an issue when creating a large amount of VCEs and makes deployment in a time-sensitive setting challenging. We also rely on a robust model during the creation of VCEs for the cone projection. While we show that standard classifiers do not yield the desired gradients, training robust models can be challenging. An interesting direction for future research is thus to replace the cone projection with another “denoising” procedure for the gradient of a non-robust model. From a theoretical standpoint, using conditional sampling with reverse transitions of the form (7) is justified, in practice, however, we have to approximate the reverse transitions with shifted normal distributions of the form (9). This approximation relies on the conditioning function having low curvature, and it can be hard to verify this once we start adding a classifier, distance and other possible terms and it is unclear how this influences the outcome of the diffusion process. The quantitative evaluation of VCEs is difficult, as the standard FID metric compares the distribution of features of a classifier over a test and a generated dataset. However, for VCEs it is not only important to generate realistic images, but also to achieve high confidence and to create meaningful changes. Moreover, metrics such as IM1, IM2 for VCEs rely on a well-trained (V)AE for every class, which is difficult to achieve for a dataset with 1000 classes and high-resolution images. Future research should therefore try to develop metrics for the quantitative evaluation of VCEs and we think that the evaluation in Tab. 1 is the first step in that direction. DVCEs and VCEs in general help to discover biases of the classifiers and thus have a positive societal impact, however one can abuse them for unintended purposes as any conditional generative model.

Conclusion

We have proposed DVCEs, a novel way to create VCEs for any state-of-the-art ImageNet classifier using diffusion models and our cone projection. DVCEs can handle vastly different image configurations, object sizes, and classes and satisfy all desired properties of VCEs.

Acknowledgement

The authors acknowledge support by the DFG Excellence Cluster Machine Learning - New Perspectives for Science, EXC 2064/1, Project number 390727645 and the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A as well as the DFG grant 389792660 as part of TRR 248.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes] See Sec. 5.

Did you discuss any potential negative societal impacts of your work? [Yes] See Sec. 5.

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] See supplemental material.

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [N/A] We didn’t train models.

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [N/A] We have conducted qualitative experiments in the main paper for the same seed and show the diversity of VCEs across seeds in App. B.3.

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] See App. D

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes]

Did you mention the license of the assets? [Yes] Yes, directly in the citation, when applicable.

Did you include any new assets either in the supplemental material or as a URL? [Yes] We include our code in the supplemental material.

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A] ImangeNet and ImageNet21k are public datasets.

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A]

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [Yes] We include the results from the user study together with the screenshots of instructions in App. C.

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A] Not applicable.

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [Yes] See App. C.

Appendix A Appendix

Overview of Appendix

In the following, we provide a brief overview of the additional experiments reported in the Appendix.

In App. B, we show how varying l1l_{1} regularization weight (B.1), the starting point of the diffusion process TT (B.2), diversity (B.3), and different levels of robustness (B.4) influence the DVCEs. In B.5, we add the ablation study of the angle used in the cone projection.

In App. C, we show the results of our user study which shows that both quantitatively and qualitatively DVCEs generate more meaningful features compared to l1.5l_{1.5}-SVCEs and BDVCEs .

In App. D, we describe the hardware and resources used.

In App. E, we explain the quantitative evaluation of the realism, validity, and closeness of our DVCEs.

In App. F, we show, how our DVCEs can help to uncover spurious features.

In App. G, we show some of the failure cases of our method.

Appendix B Ablation study of DVCEs

In this section, we start by evaluating the impact of the distance regularization term on the diffusion process. We want VCEs to resemble the original image in the overall appearance and only change class-specific features to transfer the image into the target class. To achieve sparse changes, we use l1l_{1}-distance regularization. In Fig. 7, we vary the regularization strengths from 0.050.05 to 0.250.25. As can be seen, all regularization strengths generate meaningful target class-specific features, however, regularization CdC_{d} of 0.150.15 gives a good trade-off between being close to the original image as well as being realistic. For the rest of the evaluation, we, therefore, chose a weight of 0.150.15.

B.3 Diversity

Note that the diffusion process is by design stochastic. This means that, unlike optimization-based VCEs, it is possible to generate a diverse set of VCEs from the same starting image. In Fig. 9, we visualize different DVCEs obtained for the non-robust Swin-TF using the cone projection.

B.4 Different robust models

In this section we first show more examples in Fig. 10 for the qualitative comparison of different robust models described in Sec. 4.2.

Further, in Fig. 11, we compare cone projection of 33 robust and 33 non-robust models.

B.5 Cone projection

In this section, we are going to compare different parameters for the cone projection from Sec. 3.2, which allows us to explain non-robust classifiers. Remember that we project the gradient of the robust classifier onto a cone centered around the gradient of the target classifier. Thus, larger angles for the cone allow the method to deviate more and more from the target model gradient. In Fig. 12, we show the resulting DVCEs for various angles between 1°1\degree and 50°50\degree. As shown in Fig. 3, the pure gradient of the target model is unable to visually guide the diffusion process, thus angles smaller than or equal to 15°15\degree often do not produce class-specific features in the resulting images. For angles of at least 30°30\degree, we can see class-specific features in all images. Although it is possible to use even larger angles, we use 30°30\degree throughout the entire paper because this keeps the direction close to the target model’s gradient, which is important as we want to explain the target model and not the robust one.

Appendix C User study

In this section, we discuss a user study that we performed to compare the l1.5l_{1.5}-SVCEs , BDVCEs , and our DVCEs.

Whereas in Fig. 4 we provide a qualitative comparison showing the generated VCEs, the user study provides us also with a quantitative comparison. The images for DVCEs were generated the same as for Fig. 4, see Sec. 4.1 for details. However, for BDVCEs, maximizing confidence leads to images far away from the original one (images such as leopard, tiger, timber wolf, and white wolf in Fig. 4), thus we have selected one of the settings of the hyperparameters discussed in Sec. 4.1, where the changes were looking the closest on average while introducing the meaningful features of the target classes. For l1.5l_{1.5}-SVCEs we have selected the highest radius, r=150r=150, as this still leads to smaller changes (see Tab. 1) in all the metrics, compared to DVCEs and BDVCEs, while maximizes the confidence. To generate the images, we randomly selected images from the ImageNet test set and then generated VCEs for two target classes (different from the original class of the image) which belong same WordNet category as the original class.

In this study, the users, who participated voluntarily and without payment, were shown the original images and the VCEs of the three different methods in random order. The users were researchers in machine learning and related areas but none of them is working on VCEs themself or had seen the generated VCEs before. Following , we asked 2020 users to rate 4444 VCE, if the following three properties are satisfied for a given VCE (no or multiple answers are allowed): i) “Which images have meaningful features in the target class?” (meaningful), ii) “Which images look realistic?” realism, iii) “Which images show subtle, yet understandable changes?” (subtle). The screenshot with instructions and the shown images can be seen in Fig. 13.

In the following we report the percentages of the images for DVCEs, l1.5l_{1.5}-SVCEs, and BDVCEs which were considered to have one of the three different properties: meaningful - 62.0%, 48.4%, 38.7%; realism - 34.7%, 24.6%, 52.2%; subtle - 45.0%, 50.6%, 31.0%. In Fig. 16 we report additionally all original images together with their changes and provide the percentages for each image individually.

Our DVCEs achieve with 62.0%62.0\% the highest percentage for meaningful changes, whereas BDVCEs achieve only 38.7%38.7\%. However, BDVCEs are considered to be realistic in 52.2%52.2\% of the cases whereas our DVCEs in 34.7%34.7\%. The reason for this seemingly contradictory result is that the BDVCEs are often not able to realize a meaningful class change, in the sense that the generated images do not show the corresponding class-specific feature e.g. in Fig. 16 for the change siamang →\rightarrow gorilla, the BDVCEs look very good (realism 80%80\%) but are considered to be only 20%20\% meaningful (in fact they did not change the original image) or for jellyfish →\rightarrow brain coral the shown image looks realistic but shows no features of the target class. The reason is that it despite BD has the advantage that we select the image with the highest confidence from six parameter settings whereas for our method we only have one fixed parameter setting, it is sometimes not able to reach high confidence in the target class and does either too little changes (siamang →\rightarrow gorilla) or quite significant changes but meaningless ones (jellyfish →\rightarrow brain coral). Regarding l1.5l_{1.5}-SVCEs - they have the most subtle changes 50.6%50.6\% vs. 45.0%45.0\% for DVCEs but the images show often artefacts (please use zoom into images) which is the reason why they are considered the least realistic ones. In total, the user study shows that DVCE performs best among the three methods if one considers all three categories.

Appendix D Resources and hardware used

All the experiments were done on Tesla V100 GPUs, and for the generation of a batch of 66 DVCEs without cone projection and blended VCEs 33 minutes were required. For the batch of 66 DVCEs with cone projections, 55 minutes were required, as one additional model was loaded and the gradient with respect to it was calculated.

Appendix E Quantitative evaluation

In this section, we discuss the quantitative evaluation presented in Tab. 1 to complement the user study and the qualitative evaluation of VCEs. The images for DVCEs were generated the same as in App. C. For BDVCEs, because generating images for the FID evaluation is costly, we have chosen one of the settings of the hyperparameters (that achieves high confidence and such that the resulting images look similar to the original ones on average) discussed in Sec. 4.1. For l1.5l_{1.5}-SVCEs we have selected the highest radius, r=150r=150 for the same reasons as discussed in App. C.

Realism is assessed using Fréchet Inception Distance (FID) , which was also used in . However, in , the authors noticed that there is a risk of having low FID scores for the VCEs that don’t change the starting images significantly. To overcome this problem, we propose the following simple “crossover” evaluation:

divide the classes in the WordNet clusters into two disjoint sets A and B with resp. 55045504 and 45764576 images so that one gets a roughly balanced split of ImageNet classes.

compute VCEs for the test set images of classes A and B with targets (425425 different classes) in B resp. A (352352 different classes) (crossover). More precisely, as targets we use classes from the same WordNet cluster (see the first step) but which are in the other set. This ensures subtle changes of semantically similar classes and rules out meaningless class changes like “granny smith →\rightarrow container ship”.

determine two subsets of the ImageNet trainset for the classes in A resp. B, such that in each subset the distribution of the labels corresponds to the one of A resp. B. Then we compute FID scores once between the training set corresponding to A and the VCEs generated with a target in A and for the training set corresponding to B and the VCEs with targets in B and report the average of the two FID scores.

This way, we make sure that the original images for VCEs don’t come from the same distribution as the images from the subset of the training set of ImageNet. As a sanity check, we evaluated the FID scores of the training set of classes in A to the original images of the classes in B and vice versa which yields an average of 41.5. As all methods achieve a smaller FID score (lower is better), they are all able to produce features of the target classes. However, DVCE stands out with an FID score of 17.6 compared to 27.9 (BDVCEs) and 25.6 (l1.5l_{1.5}-SVCEs).

Validity is evaluated as the mean confidence of the model achieved on the VCEs for the selected target classes. The mean confidence of DVCEs is almost the same as that of l1.5l_{1.5}-SVCEs which maximize the confidence over the l1.5l_{1.5}-ball and thus are expected to have high mean confidence. So there is little difference in validity, whereas BDVCEs have significantly lower confidence and thus have worse performance regarding validity. This is again due to the problem that the parameters leading to high confidence of the classifier but still leading to an image related to the original class are extremely difficult to find (if they exist at all) as they are image-specific.

Closeness is assessed with the set of metrics - l1,l1.5,l2,l_{1},l_{1.5},l_{2}, as well as the perceptual LPIPS metric for the AlexNet model. As l1.5l_{1.5}-SVCEs directly manipulate the image without the modification by a generative model, it is not surprising that in terms of lpl_{p}-distances they outperform DVCEs and BDVCEs.

Summary From the qualitative and the quantitative results we deduce that DVCEs and l1.5l_{1.5}-SVCEs are equally valid but DVCEs outperform in terms of image quality (realism) all other approaches significantly while remaining close to the original images.

Appendix F Spurious features

In this section, we want to show how DVCEs can be used as a “debugging” tool for ImageNet classifiers (resp. the training set). First, we show that DVCE finds the same spurious features as discovered in and we show also more spurious features found by our method.

Reproducing spurious features found in . In Fig. 17, we first reproduce the spurious features found by with l1.5l_{1.5}-SVCE for ϵ1.5=150\epsilon_{1.5}=150 for the MNR-RN50 model. We generate the DVCEs for MNR-RN50, the non-robust Swin-TF , and ConvNeXt . We find similar spurious features for the target classes “white shark” and “tench”. As discussed in , these spurious features are a consequence of the selection of the images in the training set. For “white shark” it contains a lot of images containing cages to protect the diver and thus the models pick up this “co-occurrence”. Even more extreme for “tench” where a large fraction of images show the angler holding the tench in their hands leading to the fact that counterfactuals for tench contain human faces and hands. The spurious text feature for the class “granny smith” shown by the MNR-RN50 model is not visible for Swin-TF, and ConvNeXt. One key difference is that they have been trained on ImageNet-22k and then fine-tuned to ImageNet which might have reduced this spurious feature but it requires a more detailed analysis to be sure about this.

New spurious features for non-robust models. We use our DVCEs here to find novel spurious features picked up by all classifiers (but to a mixed extent) MNR-RN50, ConvNeXt and Swin-TF: features of flowers (and not the bee itself) increase the confidence in the target class “bee” and features of the tiger face increase the confidence in the target class “tiger cat” as can be seen in Fig. 18. Here, we show both the target class and the ground truth label for all the VCEs. We also show in the rightmost column the samples from the trainset of ImageNet from the respective target classes. In both cases, they show why the classifier has picked up these spurious features. Images of bees often show flowers, most of the time much larger than the bee itself. Whereas the DVCE for “tiger cat” shows that the training set of this class is completely broken as it contains images of “tigers” (note that “tiger” is a separate class in ImageNet).

Appendix G Failure cases

From what we have seen, generally, there are two failure cases of our method: i) sometimes DVCEs can be blurry, which seems to be an artefact of the diffusion model as it happens for BDVCEs and is visible also in the original diffusion paper (see top left image in Fig. 15 of ), ii) DVCEs can fail when the change is difficult to realize (i.e. the original image is not similar to the target class).