Prediction Poisoning: Towards Defenses Against DNN Model Stealing Attacks

Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz

Introduction

Effectiveness of state-of-the-art DNN models at a variety of predictive tasks has encouraged their usage in a variety of real-world applications e.g., home assistants, autonomous vehicles, commercial cloud APIs. Models in such applications are valuable intellectual property of their creators, as developing them for commercial use is a product of intense labour and monetary effort. Hence, it is vital to preemptively identify and control threats from an adversarial lens focused at such models. In this work we address model stealing, which involves an adversary attempting to counterfeit the functionality of a target victim ML model by exploiting black-box access (query inputs in, posterior predictions out).

Stealing attacks dates back to Lowd & Meek (2005), who addressed reverse-engineering linear spam classification models. Recent literature predominantly focus on DNNs (specifically CNN image classifiers), and are shown to be highly effective (Tramèr et al., 2016) on complex models (Orekondy et al., 2019), even without knowledge of the victim’s architecture (Papernot et al., 2017b) nor the training data distribution. The attacks have also been shown to be highly effective at replicating pay-per-query image prediction APIs, for as little as $30 (Orekondy et al., 2019).

Defending against stealing attacks however has received little attention and is lacking. Existing defense strategies aim to either detect stealing query patterns (Juuti et al., 2019), or degrade quality of predicted posterior via perturbation. Since detection makes strong assumptions on the attacker’s query distribution (e.g., small L2L_{2} distances between successive queries), our focus is on the more popular perturbation-based defenses. A common theme among such defenses is accuracy-preserving posterior perturbation: the posterior distribution is manipulated while retaining the top-1 label. For instance, rounding decimals (Tramèr et al., 2016), revealing only high-confidence predictions (Orekondy et al., 2019), and introducing ambiguity at the tail end of the posterior distribution (Lee et al., 2018). Such strategies benefit from preserving the accuracy metric of the defender. However, in line with previous works (Tramèr et al., 2016; Orekondy et al., 2019; Lee et al., 2018), we find models can be effectively stolen using just the top-1 predicted label returned by the black-box. Specifically, in many cases we observe <<1% difference between attacks that use the full range of posteriors (blue line in Fig. 2) to train stolen models and the top-1 label (orange line) alone. In this paper, we work towards effective defenses (red line in Fig. 2) against DNN stealing attacks with minimal impact to defender’s accuracy.

The main insight to our approach is that unlike a benign user, a model stealing attacker additionally uses the predictions to train a replica model. By introducing controlled perturbations to predictions, our approach targets poisoning the training objective (see Fig. 2). Our approach allows for a utility-preserving defense, as well as trading-off a marginal utility cost to significantly degrade attacker’s performance. As a practical benefit, the defense involves a single hyperparameter (perturbation utility budget) and can be used with minimal overhead to any classification model without retraining or modifications.

We rigorously evaluate our approach by defending six victim models, against four recent and effective DNN stealing attack strategies (Papernot et al., 2017b; Juuti et al., 2019; Orekondy et al., 2019). Our defense consistently mitigates all stealing attacks and further shows improvements over multiple baselines. In particular, we find our defenses degrades the attacker’s query sample efficiency by 1-2 orders of magnitude. Our approach significantly reduces the attacker’s performance (e.g., 30-53% reduction on MNIST and 13-28% on CUB200) at a marginal cost (1-2%) to defender’s test accuracy. Furthermore, our approach can achieve the same level of mitigation as baseline defenses, but by introducing significantly lesser perturbation.

Contributions. (i) We propose the first utility-constrained defense against DNN model stealing attacks; (ii) We present the first active defense which poisons the attacker’s training objective by introducing bounded perturbations; and (iii) Through extensive experiments, we find our approach consistently mitigate various attacks and additionally outperform baselines.

Related Literature

Model stealing attacks (also referred to as ‘extraction’ or ‘reverse-engineering’) in literature aim to infer hyperparameters (Oh et al., 2018; Wang & Gong, 2018), recover exact parameters (Lowd & Meek, 2005; Tramèr et al., 2016; Milli et al., 2018), or extract the functionality (Correia-Silva et al., 2018; Orekondy et al., 2019) of a target black-box ML model. In some cases, the extracted model information is optionally used to perform evasion attacks (Lowd & Meek, 2005; Nelson et al., 2010; Papernot et al., 2017b). The focus of our work is model functionality stealing, where the attacker’s yardstick is test-set accuracy of the stolen model. Initial works on stealing simple linear models (Lowd & Meek, 2005) have been recently succeeded by attacks shown to be effective on complex CNNs (Papernot et al., 2017b; Correia-Silva et al., 2018; Orekondy et al., 2019) (see Appendix B for an exhaustive list). In this work, we works towards defenses targeting the latter line of DNN model stealing attacks.

Since ML models are often deployed in untrusted environments, a long line of work exists on guaranteeing certain (often orthogonal) properties to safeguard against malicious users. The properties include security (e.g., robustness towards adversarial evasion attacks (Biggio et al., 2013; Goodfellow et al., 2014; Madry et al., 2018)) and integrity (e.g., running in untrusted environments (Tramer & Boneh, 2019)). To prevent leakage of private attributes (e.g., identities) specific to training data in the resulting ML model, differential privacy (DP) methods (Dwork et al., 2014) introduce randomization during training (Abadi et al., 2016; Papernot et al., 2017a). In contrast, our defense objective is to provide confidentiality and protect the functionality (intellectual property) of the ML model against illicit duplication.

Model stealing defenses are limited. Existing works (which is primarily in multiclass classification settings) aim to either detect stealing attacks (Juuti et al., 2019; Kesarwani et al., 2018; Nelson et al., 2009; Zheng et al., 2019) or perturb the posterior prediction. We focus on the latter since detection involves making strong assumptions on adversarial query patterns. Perturbation-based defenses are predominantly non-randomized and accuracy-preserving (i.e., top-1 label is unchanged). Approaches include revealing probabilities only of confident classes (Orekondy et al., 2019), rounding probabilities (Tramèr et al., 2016), or introducing ambiguity in posteriors (Lee et al., 2018). None of the existing defenses claim to mitigate model stealing, but rather they only marginally delay the attack by increasing the number of queries. Our work focuses on presenting an effective defense, significantly decreasing the attacker’s query sample efficiency within a principled utility-constrained framework.

Preliminaries

Model Functionality Stealing. Model stealing attacks are cast as an interaction between two parties: a victim/defender VV (‘teacher’ model) and an attacker AA (‘student’ model). The only means of communication between the parties are via black-box queries: attacker queries inputs x∈X\bm{x}\in\mathcal{X} and defender returns a posterior probability distribution y∈ΔK=P(y∣x)=FV(x)\bm{y}\in\Delta^{K}=P(\bm{y}|\bm{x})=F_{V}(\bm{x}), where ΔK={y⪰0,1Ty=1}\Delta^{K}=\{\bm{y}\succeq 0,\bm{1}^{T}\bm{y}=1\} is the probability simplex over KK classes (we use KK instead of K−1K-1 for notational convenience). The attack occurs in two (sometimes overlapping) phases: (i) querying: the attacker uses the black-box as an oracle labeler on a set of inputs to construct a ‘transfer set’ of input-prediction pairs Dtransfer={(xi,yi)}i=1B\mathcal{D}^{\text{transfer}}=\{(\bm{x}_{i},\bm{y}_{i})\}_{i=1}^{B}; and (ii) training: the attacker trains a model FAF_{A} to minimize the empirical risk on Dtransfer\mathcal{D}^{\text{transfer}}. The end-goal of the attacker is to maximize accuracy on a held-out test-set (considered the same as that of the victim for evaluation purposes).

Knowledge-limited Attacker. In model stealing, attackers justifiably lack complete knowledge of the victim model FVF_{V}. Of specific interest are the model architecture and the input data distribution to train the victim model PV(X)P_{V}(X) that are not known to the attacker. Since prior work (Hinton et al., 2015; Papernot et al., 2016; Orekondy et al., 2019) indicates functionality largely transfers across architecture choices, we now focus on the query data used by the attacker. Existing attacks can be broadly categorized based on inputs {x∼PA(X)}\{\bm{x}\sim P_{A}(X)\} used to query the black-box: (a) independent distribution: (Tramèr et al., 2016; Correia-Silva et al., 2018; Orekondy et al., 2019) samples inputs from some distribution (e.g., ImageNet for images, uniform noise) independent to input data used to train the victim model; and (b) synthetic set: (Papernot et al., 2017b; Juuti et al., 2019) augment a limited set of seed data by adaptively querying perturbations (e.g., using FGSM) of existing inputs. We address both attack categories in our paper.

Defender’s Assumptions. We closely mimic an assumption-free scenario similar to existing perturbation-based defenses. The scenario entails the knowledge-limited defender: (a) unaware whether a query is malicious or benign; (b) lacking prior knowledge of the strategy used by an attacker; and (c) perturbing each prediction independently (hence circumventing Sybil attacks). For added rigor, we also study attacker’s countermeasures to our defense in Section 5.

Approach: Maximizing Angular Deviation between Gradients

Maximizing Angular Deviation (MAD). The core idea of our approach is to perturb the posterior probabilities y\bm{y} which results in an adversarial gradient signal that maximally deviates (see Fig. 2) from the original gradient (Eq. 1). More formally, we add targeted noise to the posteriors which results in a gradient direction:

to maximize the angular deviation between the original and the poisoned gradient signals:

The above presents a challenge of black-box optimization problem for the defense since the defender justifiably lacks access to the attacker model FF (Eq. 5). Apart from addressing this challenge in the next few paragraphs, we also discuss (a) solving a non-standard and non-convex constrained maximization objective; and (b) preserving accuracy of predictions via constraint (8).

Estimating G{\bm{G}}. Since we lack access to adversary’s model FF, we estimate the jacobian G=∇wlog⁡Fsur(x;w){\bm{G}}=\nabla_{\bm{w}}\log F_{\text{sur}}(\bm{x};\bm{w}) (Eq. 5) per input query x{\bm{x}} using a surrogate model FsurF_{\text{sur}}. We empirically determined (details in Appendix E.1) choice of architecture of FsurF_{\text{sur}} robust to choices of adversary’s architecture FF. However, the initialization of FsurF_{\text{sur}} plays a crucial role, with best results on a fixed randomly initialized model. We conjecture this occurs due to surrogate models with a high loss provide better gradient signals to guide the defender.

Variant: MAD-argmax. Within our defense formulation, we encode an additional constraint (Eq. 8) to preserve the accuracy of perturbed predictions. MAD-argmax variant helps us perform accuracy-preserving perturbations similar to prior work. But in contrast, the perturbations are constrained (Eq. 7) and are specifically introduced to maximize the MAD objective. We enforce the accuracy-preserving constraint in our solver by iterating over extremes of intersection of sets Eq.(6) and (8): ΔkK={y⪰0,1Ty=1,yk≥yj,k≠j}⊆ΔK\Delta^{K}_{k}=\{\bm{y}\succeq 0,\bm{1}^{T}\bm{y}=1,y_{k}\geq y_{j},k\neq j\}\subseteq\Delta^{K}.

Experimental Results

Victim Models and Datasets. We set up six victim models (see column ‘FVF_{V}’ in Table 1), each model trained on a popular image classification dataset. All models are trained using SGD (LR = 0.1) with momentum (0.5) for 30 (LeNet) or 100 epochs (VGG16), with a LR decay of 0.1 performed every 50 epochs. We train and evaluate each victim model on their respective train and test sets.

Attack Strategies. We hope to broadly address all DNN model stealing strategies during our defense evaluation. To achieve this, we consider attacks that vary in query data distributions (independent and synthetic; see Section 3) and strategies (random and adaptive). Specifically, in our experiments we use the following attack models: (i) Jacobian-based Data Augmentation ‘JBDA’ (Papernot et al., 2017b); (ii,iii) ‘JB-self’ and ‘JB-top3’ (Juuti et al., 2019); and (iv) Knockoff Nets ‘knockoff’ (Orekondy et al., 2019); We follow the default configurations of the attacks where possible. A recap and implementation details of the attack models are available in Appendix D.

Effectiveness of Attacks. We evaluate accuracy of resulting stolen models from the attack strategies as-is on the victim’s test set, thereby allowing for a fair head-to-head comparison with the victim model (additional details in Appendix A and D). The stolen model test accuracies, along with undefended victim model FVF_{V} accuracies are reported in Table 1. We observe for all six victim models, using just 50K black-box queries, attacks are able to significantly extract victim’s functionality e.g., >>87% on MNIST. We find the knockoff attack to be the strongest, exhibiting reasonable performance even on complex victim models e.g., 74.6% (0.93×\timesAcc(FVF_{V})) on Caltech256.

How Good are Existing Defenses? Most existing defenses in literature (Tramèr et al., 2016; Orekondy et al., 2019; Lee et al., 2018) perform some form of information truncation on the posterior probabilities e.g., rounding, returning top-kk labels; all strategies preserve the rank of the most confident label. We now evaluate model stealing attacks on the extreme end of information truncation, wherein the defender returns just the top-1 ‘argmax’ label. This strategy illustrates a rough lower bound on the strength of the attacker when using existing defenses. Specific to knockoff, we observe the attacker is minimally impacted on simpler datasets (e.g., 0.2% accuracy drop on CIFAR10; see Fig. A5 in Appendix). While this has a larger impact on more complex datasets involving numerous classes (e.g., a maximum of 23.4% drop observed on CUB200), the strategy also introduces a significant perturbation (L1L_{1}=1±\pm0.5) to the posteriors. The results suggest existing defenses, which largely the top-1 label, are largely ineffective at mitigating model stealing attacks.

2 Results

In the follow sections, we demonstrate the effectiveness of our defense rigorously evaluated across a wide range of complex datasets, attack models, defense baselines, query, and utility budgets. For readability, we first evaluate the defense against attack models, proceed to comparing the defense against strong baselines and then provide an analysis of the defense.

Figure 3 presents evaluation of our defenses MAD (Eq. 4-7) and MAD-argmax (Eq. 4-8) against the four attack models. To successfully mitigate attacks as a defender, we want the defense curves (colored solid lines with operating points denoted by thin crosses) to move away from undefended accuracies (denoted by circular discs, where ϵ\epsilon=0.0) to ideal defense performances (cyan cross, where Acc(Def.) is unchanged and Acc(Att.) is chance-level).

We observe from Figure 3 that by employing an identical defense across all datasets and attacks, the effectiveness of the attacker can be greatly reduced. Across all models, we find MAD provides reasonable operating points (above the diagonal), where defender achieves significantly higher test accuracies compared to the attacker. For instance, on MNIST, for <<1% drop in defender’s accuracy, our defense simultaneously reduces accuracy of the jbtop3 attacker by 52% (87.3%→\rightarrow35.7%) and knockoff by 29% (99.1%→\rightarrow69.8%). We find similar promising results even on high-dimensional complex datasets e.g., on CUB200, a 23% (65.1%→\rightarrow41.9%) performance drop of knockoff for 2% drop in defender’s test performance. Our results indicate effective defenses are achievable, where the defender can trade-off a marginal utility cost to drastically impede the attacker.

2.2 MAD Defense vs. Baseline Defenses

We now study how our approach compares to baseline defenses, by evaluating the defenses against the knockoff attack (which resulted in the strongest attack in our experiments). From Figure 4, we observe:

(i) Utility objective = L1L_{1} distance (Fig. 4a): Although random-noise and reverse-sigmoid reduce attacker’s accuracy, the strategies in most cases involves larger perturbations. In contrast, MAD and MAD-argmax provides similar non-replicability (i.e., Acc(Att.)) with significantly lesser perturbation, especially at lower magnitudes. For instance, on MNIST (first column), MAD (L1L_{1} = 0.95) reduces the accuracy of the attacker to under 80% with 0.63×\times the perturbation as that of reverse-sigmoid and random-noise (L1≈L_{1}\approx 1.5).

(ii) Utility objective = argmax-preserving (Fig. 4b): By setting a hard constraint on retaining the label of the predictions, we find the accuracy-preserving defenses MAD-argmax and reverse-sigmoid successfully reduce the performance of the attacker by at least 20% across all datasets. In most cases, we find MAD-argmax in addition achieves this objective by introducing lesser distortion to the predictions compared to reverse-sigmoid. For instance, in Fig. 4a, we find MAD-argmax consistently reduce the attacker accuracy to the same amount at lesser L1L_{1} distances. In reverse-sigmoid, we attribute the large L1L_{1} perturbations to a shift in posteriors towards a uniform distribution e.g., mean entropy of perturbed predictions is 3.02 ±\pm 0.16 (max-entropy = 3.32) at L1L_{1}=1.0 for MNIST; in contrast, MAD-argmax displays a mean entropy of 1.79 ±\pm 0.11. However, common to accuracy-preserving strategies is a pitfall that the top-1 label is retained. In Figure 7 (see overlapping red and yellow cross-marks), we present the results of training the attacker using only the top-1 label. In line with previous discussions, we find that the attacker is able to significantly recover the original performance of the stolen model for accuracy-preserving defenses MAD-argmax and reverse-sigmoid.

(iii) Non-replicability vs. utility trade-off (Fig. 4b): We now compare our defense MAD (blue lines) with baselines (rand-noise and dp-sgd) which trade-off utility to mitigate model stealing. Our results indicate MAD offers a better defense (lower attacker accuracies for similar defender accuracies). For instance, to reduce the attacker’s accuracy to <<70%, while the defender’s accuracy significantly degrades using dp-sgd (39%) and rand-noise (56.4%), MAD involves a marginal decrease of 1%.

2.3 Analysis

Ablative Analysis. We present an ablation analysis of our approach in Figure 9. In this experiment, we compare our approach MAD and MAD-argmax to: (a) G=I\bm{G}=\bm{I}: We substitute the jacobian G\bm{G} (Eq. 5) with a K×KK\times K identity matrix; and (b) y∗\bm{y}^{*}=rand: Inner maximization term (Eq. 4) returns a random extreme of the simplex. Note that both (a) and (b) do not use the gradient information to perturb the posteriors.

From Figure 9, we observe: (i) poor performance of y∗\bm{y}^{*}=rand, indicating random untargeted perturbations of the posterior probability is a poor strategy; (ii) G=I\bm{G}=\bm{I}, where the angular deviation is maximized between the posterior probability vectors is a slightly better strategy; (ii) MAD outperforms the above approaches. Consequently, we find using the gradient information (although a proxy to the attacker’s gradient signal) within our formulation (Equation 4) is crucial to providing better model stealing defenses.

Subverting the Defense. We now explore various strategies an attacker can use to circumvent the defense. To this end, we evaluate the following strategies: (a) argmax: attacker uses only the most-confident label during training; (b) arch-*: attacker trains other choices of architectures; (c) nquery: attacker queries each image multiple times; (d) nquery+aug: same as (c), but with random cropping and horizontal flipping; and (e) opt-*: attacker uses an adaptive LR optimizer e.g., ADAM (Kingma & Ba, 2014).

We present results over the subversion strategies in Figure 9. We find our defense robust to above strategies. Our results indicate that the best strategy for the attacker to circumvent our defense is to discard the probabilities and rely only on the most confident label to train the stolen model. In accuracy-preserving defenses (see Fig. 7), this previously resulted in an adversary entirely circumventing the defense (recovering up to 1.0×\times original performance). In contrast, we find MAD is nonetheless effective in spite of the strategy, maintaining a 9% absolute accuracy reduction in attacker’s stolen performance.

Conclusion

In this work, we were motivated by limited success of existing defenses against DNN model stealing attacks. While prior work is largely based on passive defenses focusing on information truncation, we proposed the first active defense strategy that attacks the adversary’s training objective. We found our approach effective in defending a variety of victim models and against various attack strategies. In particular, we find our attack can reduce the accuracy of the adversary by up to 65%, without significantly affecting defender’s accuracy.

Acknowledgement. This research was partially supported by the German Research Foundation (DFG CRC 1223). We thank Paul Swoboda and David Stutz for helpful discussions.

References

Appendix A Overview and Notation

Appendix B Related Work: Extension

A summary of existing model stealing attacks and defenses is presented in Table A2.

Appendix C Detailed Algorithm

We present a detailed algorithm (see Algorithm 1) for our approach described in Section 4.

The algorithm roughly follows four steps:

Predict (L2): Obtains posterior probability predictions y{\bm{y}} for input x{\bm{x}} using a victim model FV(x;wV)F_{V}({\bm{x}};{\bm{w}}_{V}).

Maximize MAD Objective (L4): We find the optimal direction y∗{\bm{y}}^{*} which maximizes the MAD objective (Eq. 3). To compute the arg max⁡\operatorname*{arg\,max}, we iterative over the KK extremes of the probability simplex ΔK\Delta^{K} to find y∗{\bm{y}}^{*} which maximizes the objective. The extreme yk{\bm{y}}_{k} denotes a probability vector with yk=1y_{k}=1.

Appendix D Attack Models: Recap and Implementation Details

Jacobian Based Data Augmentation (jbda) (Papernot et al., 2017b). The motivation of the approach is to obtain a surrogate of the victim black-box classifier, with an end-goal of performing evasion attacks (Biggio et al., 2013; Goodfellow et al., 2014). We restrict discussions primarily to the first part of constructing the surrogate. To obtain the surrogate (the stolen model), the authors depend on an unlabeled ‘seed’ set, typically from the same distribution as that used to train the victim model. As a result, the attacker assumes (mild) knowledge of the input data distribution and the class-label of the victim.

The key idea behind the approach is to query perturbations of inputs, to obtain a reasonable approximation of the decision boundary of the victim model. The attack strategy involves performing the following steps in a repeated manner: (i) images from the substitute set (initially the seed) D\mathcal{D} is labeled by querying the victim model FVF_{V} as an oracle labeler; (ii) the surrogate model FAF_{A} is trained on the substitute dataset; (iii) the substitute set is augmented using perturbations of existing images: Dρ+1=Dρ∪{x+λρ+1⋅sgn(JF[FA(x)])  :  x∈Dρ}\mathcal{D}_{\rho+1}=\mathcal{D}_{\rho}\cup\{{\bm{x}}+\lambda_{\rho+1}\cdot\text{sgn}(J_{F}[F_{A}({\bm{x}})])\;:\;{\bm{x}}\in\mathcal{D}_{\rho}\}, where JJ is the jacobian function.

We use a seed set of: 100 (MNIST and FashionMNIST), 500 (CIFAR10, CUB200, Caltech256) and 1000 (CIFAR100). We use the default set of hyperparameters of Papernot et al. (2017b) in other respects.

Jacobian Based {self, top-k} (jbself, jbtop3) (Juuti et al., 2019) . The authors generalize the above approach, by extending the manner in which the synthetic samples are produced. In jbself, the jacobian is calculated w.r.t to kk nearest classes and in jb-self, w.r.t the maximum a posterior class predicted by FAF_{A}.

Knockoff Nets (knockoff) (Orekondy et al., 2019) . Knockoff is a recent attack model, which demonstrated model stealing can be performed without access to seed samples. Rather, the queries to the black-box involve natural images (which can be unrelated to the training data of the victim model) sampled from a large independent data source e.g., ImageNet1K. Consequently, no knowledge of the input data distribution nor the class-label space of the victim model is required to perform model stealing. The paper proposes two strategies on how to sample images to query: random and adaptive. We use the random strategy in the paper, since adaptive resulted in marginal increases in an open-world setup (which we have).

As the independent data sources in our knockoff attacks, we use: EMNIST-Letters (when stealing MNIST victim model), EMNIST (FashionMNIST), CIFAR100 (CIFAR10), CIFAR10 (CIFAR100), ImageNet1k (CUB200, Caltech256). Overlap between query images and the training data of the victim models are purely co-incidental.

We use the code from the project’s public github repository.

Evaluating Attacks. The resulting replica model FAF_{A} from all the above attack strategies are evaluated on a held-out test set. We remark that the replica model is evaluated as-is, without additional finetuning or modifications. Similar to prior work, we evaluate the accuracies of FAF_{A} on the victim’s held-out test set. Evaluating both stolen and the victim model on the same test set allows for fair head-to-head comparison.

Appendix E Supplementary Analysis

In this section, we present additional analysis to supplement Section 5.2.3.

Central to our defense is estimating the jacobian matrix G=∇wlog⁡F(x;w){\bm{G}}=\nabla_{\bm{w}}\log F(\bm{x};\bm{w}) (Eq. 5), where F(⋅;w)F(\cdot;\bm{w}) is the attacker’s model. However, a defender with black-box attacker knowledge (where FF is unknown) requires determining G{\bm{G}} by instead using a surrogate model FsurF_{\text{sur}}. We determine choice of FsurF_{\text{sur}} empirically by studying two factors: (a) architecture of FsurF_{\text{sur}}: choice of defender’s surrogate architecture robust to varying attacker architectures (see Fig. A3); and (b) initialization of FsurF_{\text{sur}}: initialization of the surrogate model parameters plays a crucial role in providing a better defense. We consider four choices of initialization: {‘rand’, ‘early’, ‘mid’, ‘late’} which exhibits approximately {chance-level 25%, 50%, 75%} test accuracies respectively. We observe (see Fig. A3) that a randomly initialized model, which is far from convergence, provides better gradient signals in crafting perturbations.

E.2 Run-time Analysis

Appendix F Additional Plots

We present evaluation of all attacks considered in the paper on an undefended model in Figure A4. Furthermore, specific to the knockoff attack, we analyze how training using only the top-1 label (instead of complete posterior information) affects the attacker in Figure A5.

F.2 Budget vs. Accuracy

We plot the budget (i.e., number of distinct black-box attack queries to the defender) vs. the test accuracy of the defender/attacker in Figure A6. The figure supplements Figure 2 and the discussion found in Section 5.2.1 of the main paper.

F.3 Attacker argmax

In Figure A7, we perform the non-replicability vs. utility evaluation (complementing Fig. 7 in the main paper) under a special situation: the attacker discards the probabilities and only uses the top-1 ‘argmax’ label to train the stolen model. Relevant discussion can be found in Section 5.2.2.

F.4 Black-box Angular Deviations

F.5 MAD Ablation Experiments

We present the ablation experiments covering all defender models in Figure A9. Relevant discussion is available in Section 5.2.3 of the main paper under “Ablative Analysis”.