Knockoff Nets: Stealing Functionality of Black-Box Models

Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz

Introduction

Machine Learning (ML) models and especially deep neural networks are deployed to improve productivity or experience e.g., photo assistants in smartphones, image recognition APIs in cloud-based internet services, and for navigation and control in autonomous vehicles. Developing and engineering such models for commercial use is a product of intense time, money, and human effort – ranging from collecting a massive annotated dataset to tuning the right model for the task. The details of the dataset, exact model architecture, and hyperparameters are naturally kept confidential to protect the models’ value. However, in order to be monetized or simply serve a purpose, they are deployed in various applications (e.g., home assistants) to function as blackboxes: input in, predictions out.

Large-scale deployments of deep learning models in the wild has motivated the community to ask: can someone abuse the model solely based on blackbox access? There has been a series of “inference attacks” which try to infer properties (e.g., training data , architecture ) about the model within the blackbox. In this work, we focus on model functionality stealing: can one create a “knockoff” of the blackbox model solely based on observed input-output pairs? In contrast to previous works , we work with minimal assumptions on the blackbox and intend to purely steal the functionality.

We formulate model functionality stealing as follows (shown in Figure 1). The adversary interacts with a blackbox “victim” CNN by providing it input images and obtaining respective predictions. The resulting image-prediction pairs are used to train a “knockoff” model. The adversary’s intention is for the knockoff to compete with the victim model at the victim’s task. Note that knowledge transfer approaches are a special case within our formulation, where the task, train/test data, and white-box teacher (victim) model are known to the adversary.

Within this formulation, we spell out questions answered in our paper with an end-goal of model functionality stealing:

Can we train a knockoff on a random set of query images and corresponding blackbox predictions?

What makes for a good set of images to query?

How can we improve sample efficiency of queries?

What makes for a good knockoff architecture?

Related Work

Privacy, Security and Computer Vision. Privacy has been largely addressed within the computer vision community by proposing models which recognize and control privacy-sensitive information in visual content. The community has also recently studied security concerns entailing real-world usage of models e.g., adversarial perturbations in black- and white-box attack scenarios. In this work, we focus on functionality stealing of CNNs in a blackbox attack scenario.

Model Stealing. Stealing various attributes of a blackbox ML model has been recently gaining popularity: parameters , hyperparameters , architecture , information on training data and decision boundaries . These works lay the groundwork to precisely reproduce the blackbox model. In contrast, we investigate stealing functionality of the blackbox independent of its internals. Although two works are related to our task, they make relatively stronger assumptions (e.g., model family is known, victim’s data is partly available). In contrast, we present a weaker adversary.

Knowledge Distillation. Distillation and related approaches transfer the knowledge from a complex “teacher” to a simpler “student” model. Within our problem formulation, this is a special case when the adversary has strong knowledge of the victim’s blackbox model e.g., architecture, train/test data is known. Although we discuss this, a majority of the paper makes weak assumptions of the blackbox.

Active Learning. Active Learning (AL) aims to reduce labeling effort while gathering data to train a model. Ours is a special case of pool-based AL , where the learner (adversary) chooses from a pool of unlabeled data. However, unlike AL, the learner’s image pool in our case is chosen without any knowledge of the data used by the original model. Moreover, while AL considers the image to be annotated by a human-expert, ours is annotated with pseudo-labels by the blackbox.

Problem Statement

We now formalize the task of functionality stealing (see also Figure 2).

Functionality Stealing. In this paper, we introduce the task as: given blackbox query access to a “victim” model FV:X→YF_{V}:\mathcal{X}\rightarrow\mathcal{Y}, to replicate its functionality using “knockoff” model FAF_{A} of the adversary. As shown in Figure 2, we set it up as a two-player game between a victim VV and an adversary AA. Now, we discuss the assumptions in which the players operate and their corresponding moves in this game.

Victim’s Move. The victim’s end-goal is to deploy a trained CNN model FVF_{V} in the wild for a particular task (e.g., fine-grained bird classification). To train this particular model, the victim: (i) collects task-specific images x∼PV(X)\bm{x}\sim P_{V}(X) and obtains expert annotations resulting in a dataset DV={(xi,yi)}\mathcal{D}_{V}=\{(\bm{x}_{i},y_{i})\}; (ii) selects the model FVF_{V} that achieves best performance (accuracy) on a held-out test set of images DVtest\mathcal{D}_{V}^{\text{test}}. The resulting model is deployed as a blackbox which predicts output probabilities y=FV(x)\bm{y}=F_{V}(\bm{x}) given an image x\bm{x}. Furthermore, we assume each prediction incurs a cost (e.g., monetary, latency).

Adversary’s Unknowns. The adversary is presented with a blackbox CNN image classifier, which given any image x∈X\bm{x}\in\mathcal{X} returns a KK-dim posterior probability vector y∈K, ∑kyk=1\bm{y}\in^{K},\ \sum_{k}y_{k}=1. We relax this later by considering truncated versions of y\bm{y}. We assume remaining aspects to be unknown: (i) the internals of FVF_{V} e.g., hyperparameters or architecture; (ii) the data used to train and evaluate the model; and (iii) semantics over the KK classes.

Adversary’s Attack. To train a knockoff, the adversary: (i) interactively queries images {xi∼πPA(X)}\{\bm{x}_{i}\stackrel{{\scriptstyle\pi}}{{\sim}}P_{A}(X)\} using strategy π\pi to obtain a “transfer set” of images and pseudo-labels {(xi,FV(xi))}i=1B\{(\bm{x}_{i},F_{V}(\bm{x}_{i}))\}_{i=1}^{B}; and (ii) selects an architecture FAF_{A} for the knockoff and trains it to mimic the behaviour of FVF_{V} on the transfer set.

Objective. We focus on the adversary, whose primary objective is training a knockoff that performs well on the task for which FVF_{V} was designed i.e., on an unknown DVtest\mathcal{D}_{V}^{\text{test}}. In addition, we address two secondary objectives: (i) sample-efficiency: maximizing performance within a budget of BB blackbox queries; and (ii) understanding what makes for good images to query the blackbox.

Victim’s Defense. Although we primarily address the adversary’s strategy in the paper, we briefly discuss victim’s counter strategies (in Section 6) of reducing informativeness of predictions by truncation e.g., rounding-off.

Remarks: Comparison to Knowledge Distillation (KD). Training the knockoff model is reminiscent of KD approaches , whose goal is to transfer the knowledge from a larger teacher network TT (white-box) to a compact student network SS (knockoff) via the transfer set. We illustrate key differences between KD and our setting in Figure 3: (a) Independent distribution PAP_{A}: FAF_{A} is trained on images x∼PA(X)\bm{x}\sim P_{A}(X) independent to distribution PVP_{V} used for training FVF_{V}; (b) Data for supervision: Student network SS minimize variants of KD loss:

where yTτ=softmax(aT/τ)\bm{y}^{\tau}_{T}=\text{softmax}(\bm{a}_{T}/\tau) is the softened posterior distribution of logits a\bm{a} controlled by temperature τ\tau. In contrast, the knockoff (student) in our case lacks logits aT\bm{a}_{T} and true labels ytrue\bm{y}_{\text{true}} to supervise training.

Generating Knockoffs

In this section, we elaborate on the adversary’s approach in two steps: transfer set construction (Section 4.1) and training knockoff FAF_{A} (Section 4.2).

The goal is to obtain a transfer set i.e., image-prediction pairs, on which the knockoff will be trained to imitate the victim’s blackbox model FVF_{V}.

Selecting PA(X)P_{A}(X). The adversary first selects an image distribution to sample images. We consider this to be a large discrete set of images. For instance, one of the distributions PAP_{A} we consider is the 1.2M images of ILSVRC dataset .

Sampling Strategy π\pi. Once the image distribution PA(X)P_{A}(X) is chosen, the adversary samples images x∼πPA(X)\bm{x}\stackrel{{\scriptstyle\pi}}{{\sim}}P_{A}(X) using a strategy π\pi. We consider two strategies.

In this strategy, we randomly sample images (without replacement) x∼iidPA(X)\bm{x}\stackrel{{\scriptstyle\text{iid}}}{{\sim}}P_{A}(X) to query FVF_{V}. This is an extreme case where adversary performs pure exploration. However, there is a risk that the adversary samples images irrelevant to learning the task (e.g., over-querying dog images to a birds classifier).

1.2 Adaptive Strategy

We now incorporate a feedback signal resulting from each image queried to the blackbox. A policy π\pi is learnt:

Supplementing PA\bm{P_{A}}. To encourage relevant queries, we enrich images in the adversary’s distribution by associating each image xi\bm{x}_{i} with a label zi∈Zz_{i}\in Z. No semantic relation of these labels with the blackbox’s output classes is assumed or exploited. As an example, when PAP_{A} corresponds to 1.2M images of the ILSVRC dataset, we use labels defined over 1000 classes. These labels can be alternatively obtained by unsupervised measures e.g., clustering or estimating graph-density . We find using labels aids understanding blackbox functionality. Furthermore, since we expect labels {zi∈Z}\{z_{i}\in Z\} to be correlated or inter-dependent, we represent them within a coarse-to-fine hierarchy, as nodes of a tree as shown in Figure 4b.

Actions. At each time-step tt, we sample actions from a discrete action space zt∈Zz_{t}\in Z i.e., adversary’s independent label space. Drawing an action is a forward-pass (denoted by a blue line in Figure 4b) through the tree: at each level, we sample a node with probability πt(z)\pi_{t}(z) . The probabilities are determined by a softmax distribution over the node potentials: πt(z)=eHt(z)∑z′Ht(z′)\pi_{t}(z)=\frac{e^{H_{t}(z)}}{\sum_{z^{\prime}}H_{t}(z^{\prime})}. Upon reaching a leaf-node, a sample of images is returned corresponding to label ztz_{t}.

Learning the Policy. We use the received reward rtr_{t} for an action ztz_{t} to update the policy π\pi using the gradient bandit algorithm . This update is equivalent to a backward-pass through the tree (denoted by a green line in Figure 4b), where the node potentials are updated as:

where α=1/N(z)\alpha=1/N(z) is the learning rate, N(z)N(z) is the number of times action zz has been drawn, and rˉt\bar{r}_{t} is the mean-reward over past Δ\Delta time-steps.

Rewards. To evaluate the quality of sampled images xt\bm{x}_{t}, we study three rewards. We use a margin-based certainty measure to encourage images where the victim is confident (hence indicating the domain FVF_{V} was trained on):

To prevent the degenerate case of image exploitation over a single label, we introduce a diversity reward:

To encourage images where the knockoff prediction y^t=FA(xt)\hat{\bm{y}}_{t}=F_{A}(\bm{x}_{t}) does not imitate FVF_{V}, we reward high loss:

We sum up individual rewards when multiple measures are used. To maintain an equal weighting, each reward is individually rescaled to and subtracted with a baseline computed over past Δ\Delta time-steps.

As a product of the previous step of interactively querying the blackbox model, we have a transfer set {(xt,FV(xt)}t=1B, xt∼πPA(X)\{(\bm{x}_{t},F_{V}(\bm{x}_{t})\}_{t=1}^{B},\ \bm{x}_{t}\stackrel{{\scriptstyle\pi}}{{\sim}}P_{A}(X). Now we address how this is used to train a knockoff FAF_{A}.

Selecting Architecture FAF_{A}. Few works have recently explored reverse-engineering the blackbox i.e., identifying the architecture, hyperparameters, etc. We however argue this is orthogonal to our requirement of simply stealing the functionality. Instead, we represent FAF_{A} with a reasonably complex architecture e.g., VGG or ResNet . Existing findings in KD and model compression indicate robustness to choice of reasonably complex student models. We investigate the choice under weaker knowledge of the teacher (FVF_{V}) e.g., training data and architecture is unknown.

Training to Imitate. To bootstrap learning, we begin with a pretrained Imagenet network FAF_{A}. We train the knockoff FAF_{A} to imitate FVF_{V} on the transfer set by minimizing the cross-entropy (CE) loss: LCE(y,y^)=−∑kp(yk)⋅log⁡p(yk^)\mathcal{L}_{\text{CE}}(\bm{y},\hat{\bm{y}})=-\sum_{k}p(y_{k})\cdot\log p(\hat{y_{k}}). This is a standard CE loss, albeit weighed with the confidence p(yk)p(y_{k}) of the victim’s label. This formulation is equivalent to minimizing the KL-divergence between the victim’s and knockoff’s predictions over the transfer set.

Experimental Setup

We now discuss the experimental setup of multiple victim blackboxes (Section 5.1), followed by details on the adversary’s approach (Section 5.2).

We choose four diverse image classification CNNs, addressing multiple challenges in image classification e.g., fine-grained recognition. Each CNN performs a task specific to a dataset. A summary of the blackboxes is presented in Table 1 (extended descriptions in appendix).

Training the Black-boxes. All models are trained using a ResNet-34 architecture (with ImageNet pretrained weights) on the training split of the respective datasets. We find this architecture choice achieve strong performance on all datasets at a reasonable computational cost. Models are trained using SGD with momentum (of 0.5) optimizer for 200 epochs with a base learning rate of 0.1 decayed by a factor of 0.1 every 60 epochs. We follow the train-test splits suggested by the respective authors for Caltech-256 , CUBS-200-2011 , and Indoor-Scenes . Since GT annotations for Diabetic-Retinopathy test images are not provided, we reserve 200 training images for each of the five classes for testing. The number of test images per class for all datasets are roughly balanced. The test images of these datasets DVtest\mathcal{D}_{V}^{\text{test}} are used to evaluate both the victim and knockoff models.

After these four victim models are trained, we use them as a blackbox for the remainder of the paper: images in, posterior probabilities out.

In this section, we elaborate on the setup of two aspects relevant to transfer set construction (Section 4.1).

Our approach for transfer set construction involves the adversary querying images from a large discrete image distribution PAP_{A}. In this section, we present four choices considered in our experiments. Any information apart from the images from the respective datasets are unused in the random strategy. For the adaptive strategy, we use image-level labels (chosen independent of blackbox models) to guide sampling.

PA=PV\bm{P_{A}=P_{V}}. For reference, we sample from the exact set of images used to train the blackboxes. This is a special case of knowledge-distillation with unlabeled data at temperature τ=1\tau=1.

PA=\bm{P_{A}=} ILSVRC . We use the collection of 1.2M images over 1000 categories presented in the ILSVRC-2012 challenge.

PA=\bm{P_{A}=} OpenImages . OpenImages v4 is a large-scale dataset of 9.2M images gathered from Flickr. We use a subset of 550K unique images, gathered by sampling 2k images from each of 600 categories.

PA=D2\bm{P_{A}=D^{2}}. We construct a dataset wherein the adversary has access to all images in the universe. In our case, we create the dataset by pooling training data from: (i) all four datasets listed in Section 5.1; and (ii) both datasets presented in this section. This results in a “dataset of datasets” D2D^{2} of 2.2M images and 2129 classes.

Overlap between PAP_{A} and PVP_{V}. We compute overlap between labels of the blackbox (KK, e.g., 256 Caltech classes) and the adversary’s dataset (ZZ, e.g., 1k ILSVRC classes) as: 100×∣K∩Z∣/∣K∣100\times|K\cap Z|/|K|. Based on the overlap between the two image distributions, we categorize PAP_{A} as:

PA=PV\bm{P_{A}=P_{V}}: Images queried are identical to the ones used for training FVF_{V}. There is a 100% overlap.

Closed-world (PA=D2\bm{P_{A}=D^{2}}): Blackbox train data PVP_{V} is a subset of the image universe PAP_{A}. There is a 100% overlap.

Open-world (PA∈\bm{P_{A}\in} {ILSVRC, OpenImages}): Any overlap between PVP_{V} and PAP_{A} is purely coincidental. Overlaps are: Caltech256 (42% ILSVRC, 44% OpenImages), CUBS200 (1%, 0.5%), Indoor67 (15%, 6%), and Diabetic5 (0%, 0%).

2.2 Adaptive Strategy

In the adaptive strategy (Section 4.1.2), we make use of auxiliary information (labels) in the adversary’s data PAP_{A} to guide the construction of the transfer set. We represent these labels as the leaf nodes in the coarse-to-fine concept hierarchy tree. The root node in all cases is a single concept “entity”. We obtain the rest of the hierarchy as follows: (i) D2D^{2}: we add as parents the dataset the images belong to; (ii) ILSVRC: for each of the 1K labels, we obtain 30 coarse labels by clustering the mean visual features of each label obtained using 2048-dim pool features of an ILSVRC pretrained Resnet model; (iii) OpenImages: We use the exact hierarchy provided by the authors.

Results

Training Phases. The knockoff models are trained in two phases: (a) Online: during transfer set construction (Section 4.1); followed by (b) Offline: the model is retrained using transfer set obtained thus far (Section 4.2). All results on knockoff are reported after step (b).

Evaluation Metric. We evaluate two aspects of the knockoff: (a) Top-1 accuracy: computed on victim’s held-out test data DVtest\mathcal{D}_{V}^{\text{test}} (b) sample-efficiency: best performance achieved after a budget of BB queries. Accuracy is reported in two forms: absolute (xx%) or relative to blackbox FVF_{V} (x×x\times).

In each of the following experiments, we evaluate our approach with identical hyperparameters across all blackboxes, highlighting the generalizability of model functionality stealing.

In this section, we analyze influence of transfer set {(xi,FV(xi)}\{(\bm{x}_{i},F_{V}(\bm{x}_{i})\} on the knockoff. For simplicity, for the remainder of this section we fix the architecture of the victim and knockoff to a Resnet-34 .

Reference: PA=PV\bm{P_{A}=P_{V}} (KD). From Table 2 (second row), we observe: (i) all knockoff models recover 0.920.92-1.05×1.05\times performance of FVF_{V}; (ii) a better performance than FVF_{V} itself (e.g., 3.8% improvement on Caltech256) due to regularizing effect of training on soft-labels .

Can we learn by querying randomly from an independent distribution? Unlike KD, the knockoff is now trained and evaluated on different image distributions (PAP_{A} and PVP_{V} respectively). We first focus on the random strategy, which does not use any auxiliary information.

We make the following observations from Table 2 (random): (i) closed-world: the knockoff is able to reasonably imitate all the blackbox models, recovering 0.840.84-0.97×0.97\times blackbox performance; (ii) open-world: in this challenging scenario, the knockoff model has never encountered images of numerous classes at test-time e.g., >>90% of the bird classes in CUBS200. Yet remarkably, the knockoff is able to obtain 0.810.81-0.96×0.96\times performance of the blackbox. Moreover, results marginally vary (at most 0.04×0.04\times) between ILSVRC and OpenImages, indicating any large diverse set of images makes for a good transfer set.

Upon qualitative analysis, we find the image and pseudo-label pairs in the transfer set are semantically incoherent (Fig. 6a) for output classes non-existent in training images PAP_{A}. However, when relevant images are presented at test-time (Fig. 6b), the adversary displays strong performance. Furthermore, we find the top predictions by knockoff relevant to the image e.g., predicting one comic character (superman) for another.

How sample-efficient can we get? Now we evaluate the adaptive strategy (discussed in Section 4.1.2). Note that we make use of auxiliary information of the images in these tasks (labels of images in PAP_{A}). We use the reward set which obtained the best performance in each scenario: {certainty} in closed-world and {certainty, diversity, loss} in open-world.

From Figure 5, we observe: (i) closed-world: adaptive is extremely sample-efficient in all but one case. Its performance is comparable to KD in spite of samples drawn from a 3636-188×188\times larger image distribution. We find significant sample-efficiency improvements e.g., while CUBS200-random reaches 68.3% at BB=60k, adaptive achieves this 6×\times quicker at BB=10k. We find comparably low performance in Diabetic5 as the blackbox exhibits confident predictions for all images resulting in poor feedback signal to guide policy; (ii) open-world: although we find marginal improvements over random in this challenging scenario, they are pronounced in few cases e.g., 1.5×1.5\times quicker to reach an accuracy 57% on CUBS200 with OpenImages. (iii) as an added-benefit apart from sample-efficiency, from Table 2, we find adaptive display improved performance (up to 4.5%) consistently across all choices of FVF_{V}.

What can we learn by inspecting the policy? From previous experiments, we observed two benefits of the adaptive strategy: sample-efficiency (although more prominent in the closed-world) and improved performance. The policy πt\pi_{t} learnt by adaptive (Section 4.1.2) additionally allows us to understand what makes for good images to query. πt(z)\pi_{t}(z) is a discrete probability distribution indicating preference over action zz. Each action zz in our case corresponds to labels in the adversary’s image distribution.

We visualize πt(z)\pi_{t}(z) in Figure 7, where each bar represents an action and its color, the parent in the hierarchy. We observe: (i) closed-world (Fig. 7 top): actions sampled with higher probabilities consistently correspond to output classes of FVF_{V}. Upon analyzing parents of these actions (the dataset source), the policy also learns to sample images for the output classes from an alternative richer image source e.g., “ladder” images in Caltech256 sampled from OpenImages instead; (ii) open-world (Fig. 7 bottom): unlike closed-world, the optimal mapping between adversary’s actions to blackbox’s output classes is non-trivial and unclear. However, we find top actions typically correspond to output classes of FVF_{V} e.g., indigo bunting. The policy, in addition, learns to sample coarser actions related to the FVF_{V}’s task e.g., predominantly drawing from birds and animals images to knockoff CUBS200.

What makes for a good reward? Using the adaptive sampling strategy, we now address influence of three rewards (discussed in Section 4.1.2). We observe: (i) closed-world (Fig. 8 left): All reward signals in adaptive helps with the sample efficiency over random. Reward cert (which encourages exploitation) provides the best feedback signal. Including other rewards (cert+div+L\mathcal{L}) slightly deteriorates performance, as they encourage exploration over related or unseen actions – which is not ideal in a closed-world. Reward uncert, a popular measure used in AL literature underperforms in our setting since it encourages uncertain (in our case, irrelevant) images. (ii) open-world (Fig. 8 right): All rewards display only none-to-marginal improvements for all choices of FVF_{V}, with the highest improvement in CUBS200 using cert+div+L\mathcal{L}. However, we notice an influence on learnt policies where adopting exploration (div + L\mathcal{L}) with exploitation (cert) goals result in a softer probability distribution π\pi over the action space and in turn, encouraging related images.

Can we train knockoffs with truncated blackbox outputs? So far, we found adversary’s attack objective of knocking off blackbox models can be effectively carried out with minimal assumptions. Now we explore the influence of victim’s defense strategy of reducing informativeness of blackbox predictions to counter adversary’s model stealing attack. We consider two truncation strategies: (a) top-kk: top-kk (out of KK) unnormalized posterior probabilities are retained, while rest are zeroed-out; (b) rounding rr: posteriors are rounded to rr decimals e.g., round(0.127, rr=2) = 0.13. In addition, we consider the extreme case “argmax”, where only index k=arg max⁡kykk=\operatorname*{arg\,max}_{k}y_{k} is returned.

From Figure 9 (with KK = 256), we observe: (i) truncating yi\bm{y}_{i} – either using top-kk or rounding – slightly impacts the knockoff performance, with argmax achieving 0.76-0.84×\times accuracy of original performance for any budget BB; (ii) top-kk: even small increments of kk significantly recovers the original performance – 0.91×0.91\times at k=2k=2 and 0.96×0.96\times at k=5k=5; (iii) rounding: recovery is more pronounced, with 0.99×0.99\times original accuracy achieved at just r=2r=2. We find model functionality stealing minimally impacted by reducing informativeness of blackbox predictions.

2 Architecture choice

In the previous section, we found model functionality stealing to be consistently effective while keeping the architectures of the blackbox and knockoff fixed. Now we study the influence of the architectural choice FAF_{A} vs. FVF_{V}.

How does the architecture of FAF_{A} influence knockoff performance? We study the influence using two choices of the blackbox FVF_{V} architecture: Resnet-34 and VGG-16 . Keeping these fixed, we vary architecture of the knockoff FAF_{A} by choosing from: Alexnet , VGG-16 , Resnet-{18, 34, 50, 101} , and Densenet-161 .

From Figure 10, we observe: (i) performance of the knockoff ordered by model complexity: Alexnet (lowest performance) is at one end of the spectrum while significantly more complex Resnet-101/Densenet-161 are at the other; (ii) performance transfers across model families: Resnet-34 achieves similar performance when stealing VGG-16 and vice versa; (iii) complexity helps: selecting a more complex model architecture of the knockoff is beneficial. This contrasts KD settings where the objective is to have a more compact student (knockoff) model.

3 Stealing Functionality of a Real-world Black-box Model

Now we validate model functionality stealing on a popular image analysis API. Such image recognition services are gaining popularity allowing users to obtain image-predictions for a variety of tasks at low costs ($1-2 per 1k queries). These image recognition APIs have also been used to evaluate other attacks e.g., adversarial examples . We focus on a facial characteristics API which given an image, returns attributes and confidences per face. Note that in this experiment, we have semantic information of blackbox output classes.

Collecting PAP_{A}. The API returns probability vectors per face in the image and thus, querying irrelevant images leads to a wasted result with no output information. Hence, we use two face image sets PAP_{A} for this experiment: CelebA (220k images) and OpenImages-Faces (98k images). We create the latter by cropping faces (plus margin) from images in the OpenImages dataset .

Evaluation. Unlike previous experiments, we cannot access victim’s test data. Hence, we create test sets for each image set by collecting and manually screening seed annotations from the API on ∼\sim5K images.

How does this translate to the real-world? We model two variants of the knockoff using the random strategy (adaptive is not used since no relevant auxiliary information of images are available). We present each variant using two choices of architecture FAF_{A}: a compact Resnet-34 and a complex Resnet-101. From Figure 11, we observe: (i) strong performance of the knockoffs achieving 0.760.76-0.82×0.82\times performance as that of the API on the test sets; (ii) the diverse nature OpenImages-Faces helps improve generalization resulting in 0.82×0.82\times accuracy of the API on both test-sets; (iii) the complexity of FAF_{A} does not play a significant role: both Resnet-34 and Resnet-101 show similar performance indicating a compact architecture is sufficient to capture discriminative features for this particular task.

We find model functionality stealing translates well to the real-world with knockoffs exhibiting a strong performance. The knockoff circumvents monetary and labour costs of: (a) collecting images for the task; (b) obtaining expert annotations; and (c) tuning a model. As a result, an inexpensive knockoff is trained which exhibits strong performance, using victim API queries amounting to only $30.

Conclusion

We investigated the problem of model functionality stealing where an adversary transfers the functionality of a victim model into a knockoff via blackbox access. In spite of minimal assumptions on the blackbox, we demonstrated the surprising effectiveness of our approach. Finally, we validated our approach on a popular image recognition API and found strong performance of knockoffs. We find functionality stealing poses a real-world threat that potentially undercuts an increasing number of deployed ML models.

Acknowledgement. This research was partially supported by the German Research Foundation (DFG CRC 1223). We thank Yang Zhang for helpful discussions.

References

Appendix A Contents

Aggregating OpenImages and OpenImages-Faces

Adaptive strategy: With/without hierarchy

Appendix B Extended Descriptions

In this section, we provide additional detailed descriptions and implementation details.

We supplement Section 5.1 by providing extended descriptions of the blackboxes listed in Table 1. Each blackbox FVF_{V} is trained on one particular image classification dataset.

Black-box 1: Caltech256 . Caltech-256 is a popular dataset for general object recognition gathered by downloading relevant examples from Google Images and manually screening for quality and errors. The dataset contains 30k images covering 256 common object categories.

Black-box 2: CUBS200 . A fine-grained bird-classifier is trained on the CUBS-200-2011 dataset. This dataset contains roughly 30 train and 30 test images for each of 200 species of birds. Due to the low intra-class variance, collecting and annotating images is challenging even for expert bird-watchers.

Black-box 3: Indoor67 . We introduce another fine-grained task of recognizing 67 types of indoor scenes. This dataset consists of 15.6k images collected from Google Images, Flickr, and LabelMe.

Black-box 4: Diabetic5 . Diabetic Retinopathy (DR) is a medical eye condition characterized by retinal damage due to diabetes. Cases are typically determined by trained clinicians who look for presence of lesions and vascular abnormalities in digital color photographs of the retina captured using specialized cameras. Recently, a dataset of such 35k retinal image scans was made available as a part of a Kaggle competition . Each image is annotated by a clinician on a scale of 0 (no DR) to 4 (proliferative DR). This highly-specialized biomedical dataset also presents challenges in the form of extreme imbalance (largest class contains 30×\times as the smallest one).

B.2 Overlap: Open-world

In this section, we supplement Section 5.2.1 in the main paper by providing more details on how overlap was calculated in the open-world scenarios. We manually compute overlap between labels of the blackbox (KK, e.g., 256 Caltech classes) and the adversary’s dataset (ZZ, e.g., 1k ILSVRC classes) as: 100×∣K∩Z∣/∣K∣100\times|K\cap Z|/|K|. We denote two labels k∈Kk\in K and z∈Zz\in Z to overlap if: (a) they have the same semantic meaning; or (b) zz is a type of kk e.g., zz = “maltese dog” and kk = “dog”. The exact numbers are provided in Table 3. We remark that this is a soft-lower bound. For instance, while ILSVRC contains “Hummingbird” and CUBS-200-2011 contains three distinct species of hummingbirds, this is not counted towards the overlap as the adversary lacks annotated data necessary to discriminate among the three species.

B.3 Dataset Aggregation

All datasets used in the paper (expect OpenImages) have been used in the form made publicly available by the authors. We use a subset of OpenImages due to storage constraints imposed by its massive size (9M images). The description to obtain these subsets are provided below.

OpenImages. We retrieve 2k images for each of the 600 OpenImages “boxable” categories, resulting in 554k unique images. ∼\sim19k images are removed for either being corrupt or representing Flickr’s placeholder for unavailable images. This results in a total of 535k unique images.

OpenImages-Faces. We download all images (422k) from OpenImages with label “/m/0dzct: Human face” using the OID tool . The bounding box annotations are used to crop faces (plus a margin of 25%) containing at least 180×\times180 pixels. We restrict to at most 5 faces per image to maintain diversity between train/test splits. This results in a total of 98k faces images.

B.4 Additional Implementation Details

In this section, we provide implementation details to supplement discussions in the main paper.

Training FV=F_{V}= Diabetic5. Training this victim model is identical to other blackboxes except for one aspect: weighted loss. Due to the extreme imbalance between classes of the dataset, we weigh each class as follows. Let nkn_{k} denote the number of images belonging to class kk and let nmin⁡=min⁡knkn_{\min}=\min_{k}n_{k}. We weigh the loss for each class kk as nmin⁡/nkn_{\min}/n_{k}. From our experiments with weighted loss, we found approximately 8% absolute improvement in overall accuracy on the test set. However, the training of knockoffs of all blackboxes are identical in all aspects, including a non-weighted loss irrespective of the victim blackbox targeted.

Creating ILSVRC Hierarchy. We represent the 1k labels of ILSVRC as a hierarchy Figure 4b in the form: root node “entity” →\rightarrow NN coarse nodes →\rightarrow 1k leaf nodes. We obtain NN (30 in our case) coarse labels as follows: (i) a 2048-d mean feature vector representation per 1k labels is obtained using an Imagenet-pretrained ResNet ; (ii) we cluster the 1k features into NN clusters using scikit-learn’s implementation of agglomerative clustering; (iii) we obtain semantic labels per cluster (i.e., coarse node) by finding the common parent in the Imagenet semantic hierarchy.

Adaptive Strategy. Recall from Section 6, we train the knockoff in two phases: (a) Online: during transfer set construction; followed by (b) Offline: the model is retrained using transfer set obtained thus far. In phase (a), we train FAF_{A} with SGD (with 0.5 momentum) with a learning rate of 0.0005 and batch size of 4 (i.e., 4 images sampled at each tt). In phase (b), we train the knockoff FAF_{A} from scratch on the transfer set using SGD (with 0.5 momentum) for 100 epochs with learning rate of 0.01 decayed by a factor of 0.1 every 60 epochs.

Appendix C Extensions of Existing Results

In this section, we present extensions of existing results discussed in the main paper.

Qualitative results to supplement Figure 6 are provided in Figures 13-16. Each row in the figures correspond to an output class of the blackbox whose images the knockoff has never encountered before. Images in the “transfer set” column were randomly sampled from ILSVRC . In contrast, images in the “test set” belong to the victim’s test set (Caltech256, CUBS-200-2011, etc.).

C.2 Sample Efficiency: Training Knockoffs on GT

We extend Figure 5 in the main paper to include training on the same ground-truth data used to train the blackboxes. This extension “PV(FV)P_{V}(F_{V})” is illustrated in Figure 12, displayed alongside KD approach. The figure represents the sample-efficiency of the first two rows of 2. Here we observe: (i) comparable performance in all but one case (Diabetic5, discussed shortly) indicating KD is an effective approach to train knockoffs; (ii) we find KD achieve better performance in Caltech256 and Diabetic5 due to regularizing effect of training on soft-labels on an imbalanced dataset.

C.3 Policies learnt by Adaptive

We inspected the policy π\pi learnt by the adaptive strategy in Section 6.1. In this section, we provide policies over all blackboxes in the closed- and open-world setting. Figures 17(a) and 17(c) display probabilities of each action z∈Zz\in Z at t=2500t=2500.

Since the distribution of rewards is non-stationary, we visualize the policy over time in Figure 17(b) for CUBS200 in a closed-world setup. From this figure, we observe an evolution where: (i) at early stages (t∈t\in), the approach samples (without replacement) images that overlaps with the victim’s train data; and (ii) at later stages (t∈t\in), since the overlapping images have been exhausted, the approach explores related images from other datasets e.g., “ostrich”, “jaguar”.

C.4 Reward Ablation

The reward ablation experiment 8 for the remaining datasets are provided in Figure 18. We make similar observations as before for Indoor67. However, since FV=F_{V}= Diabetic5 demonstrates confident predictions in all images (discussed in 563-567), we find little-to-no improvement for knockoffs of this victim model.

Appendix D Auxiliary Experiments

In this section, we present experiments to supplement existing results in the main paper.

We now discuss evaluation to supplement Section 5.2.1 and Section 6.

In 6.1 we highlighted strong performance of the knockoff even among classes that were never encountered (see Table 3 for exact numbers) during training. To elaborate, we split the blackbox output classes into “seen” and “unseen” categories and present mean per-class accuracies in Figure 19. Although we find better performance on classes seen while training the knockoff, performance of unseen classes is remarkably high, with the knockoff achieving >>70% performance in both cases.

D.2 Adaptive: With and without hierarchy

The adaptive strategy presented in Section 4.1.2 uses a hierarchy discussed in Section 5.2.2. As a result, we approached this as a hierarchical multi-armed bandit problem. Now, we present an alternate approach adaptive-flat, without the hierarchy. This is simply a multi-armed bandit problem with ∣Z∣|Z| arms (actions).

Figure 20 illustrates the performance of these approaches using PA=D2P_{A}=D^{2} (∣Z∣|Z| = 2129) and rewards {certainty, diversity, loss}. We observe adaptive consistently outperforms adaptive-flat. For instance, in CUBS200, adaptive is 2×\times more sample-efficient to reach accuracy of 50%. We find the hierarchy helps the adversary (agent) better navigate the large action space.

D.3 Semi-open World

The closed-world experiments (PA=D2P_{A}=D^{2}) presented in Section 6.1 and discussed in Section 5.2.1 assumed access to the image universe. Thereby, the overlap between PAP_{A} and PVP_{V} was 100%. Now, we present an intermediate overlap scenario semi-open world by parameterizing the overlap as: (i) τd\tau_{\text{d}}: The overlap between images PAP_{A} and PVP_{V} is 100×τd100\times\tau_{\text{d}}; and (ii) τk\tau_{\text{k}}: The overlap between labels KK and ZZ is 100×τk100\times\tau_{\text{k}}. In both these cases τd,τk∈(0,1]\tau_{\text{d}},\tau_{\text{k}}\in(0,1] represents the fraction of PAP_{A} used. τd=τk=1\tau_{\text{d}}=\tau_{\text{k}}=1 depicts the closed-world scenario discussed in Section 6.1.

From Figure 21, we observe: (i) the random strategy is unaffected in the semi-open world scenario, displaying comparable performance for all values of τd\tau_{\text{d}} and τk\tau_{\text{k}}; (ii) τd\tau_{\text{d}}: knockoff obtained using adaptive obtains strong performance even with low overlap e.g., a difference of at most 3% performance in Caltech256 even at τd=0.1\tau_{\text{d}}=0.1; (iii) τk\tau_{\text{k}}: although the adaptive strategy is minimally affected in few cases (e.g., CUBS200), we find the performance drop due to a pure exploitation (certainty) that is used. We observed recovery in performance by using all rewards indicating exploration goals (diversity, loss) are necessary when transitioning to an open-world scenario.