Copy, Right? A Testing Framework for Copyright Protection of Deep Learning Models

Jialuo Chen, Jingyi Wang, Tinglan Peng, Youcheng Sun, Peng Cheng, Shouling Ji, Xingjun Ma, Bo Li, Dawn Song

I Introduction

Deep learning models, e.g., deep neural networks (DNNs), have become the standard models for solving many complex real-world problems, such as image recognition , speech recognition , natural language processing , and autonomous driving . However, training large-scale DNN models is by no means trivial, which requires not only large-scale datasets but also significant computational resources. The training cost can grow rapidly with task complexity and model capacity. For instance, it can cost $1.6 million to train a BERT model on Wikipedia and Book corpora (15 GB) . It is thus of utmost importance to protect DNNs from unauthorized duplication or reproduction.

One concerning fact is that well-trained DNNs are often exposed to the public via remote services (APIs), cloud platforms (e.g., Amazon AWS, Google Cloud and Microsoft Azure), or open-source toolkits like OpenVINOhttps://github.com/openvinotoolkit/open_model_zoo. It gives rise to adversaries (e.g., a model “thief”) who attempt to steal the model in stealthy ways, causing copyright infringement and economic losses to the model owners. Recent studies have shown that stealing a DNN can be done very efficiently without leaving obvious traces . Arguably, unauthorized finetuning or pruning is the most straightforward way of model stealing, if the model parameters are publicly accessible (for research purposes only) or the adversary is an insider. Even when only the API is exposed, the adversary can still exploit advanced model extraction techniques to steal most functionalities of the hidden model. These attacks pose serious threats to the copyright of deep learning models, calling for effective protection methods.

A number of defense techniques have been proposed to protect the copyright of DNNs, where DNN watermarking is one major type of technique. DNN watermarking embeds a secret watermark (e.g., logo or signature) into the model by exploiting the over-parameterization property of DNNs . The ownership can then be verified when the same or similar watermark is extracted from a suspect model. The use of watermarks has an obvious advantage, i.e., the owner identity can be embedded and verified exactly, given that the watermark can be fully extracted. However, these methods still suffer from certain weaknesses. Arguably, the most concerning one is that they are invasive, i.e., they need to tamper with the training procedure to embed the watermark, which may compromise model utility or introduce new security threats into the model .

More recently, DNN fingerprinting has been proposed as a non-invasive alternative to watermarking. Lying at the design core of fingerprinting is uniqueness — the unique feature of a DNN model. Specifically, fingerprinting extracts a unique identifier (or fingerprint) from the owner model to differentiate it from other models. The ownership can be claimed if the identifier of the owner model matches with that of a suspect model. However, in the context of deep learning, a single fingerprinting feature/metric can hardly be sufficient or flexible enough to handle all the randomness in DNNs or against different types of model stealing and adaptive attacks (as we will show in our experiments). In other words, there exist many scenarios where a DNN model can easily lose its unique feature or property (i.e., fingerprint).

In this work, we propose a testing approach for DNN copyright protection. Instead of solely relying on one metric, we propose to actively test the “similarities” between a victim model and a suspect model from multiple angles. The core idea is to 1) carefully construct a set of test cases to comprehensively characterize the victim model, and 2) measure how similarly the two models behave on the test cases. Intuitively, if a suspect model is a stolen copy of the victim model, it will behave just like the victim model in certain ways. An extreme case is that the suspect is the exact duplicate of the victim model, and in this case, the two models will behave identically on these test cases. This testing view creates a dilemma for the adversary as better stealing will inevitably lead to higher similarities to the victim model. We further identify two major challenges for testing-based copyright protection: 1) how to define comprehensive testing metrics to fully characterize the similarities between two models, and 2) how to effectively generate test cases to amplify the similarities. The set of similarity scores can be viewed as a proof obligation that provides a chain of strong evidence to judge a stolen copy.

Following the above idea, we design and implement DeepJudge, a novel testing framework for DNN copyright protection. As illustrated in Fig. 1, DeepJudge is composed of three core components. First, we propose a set of multi-level testing metrics to fully characterize a DNN model from different angles. Second, we propose efficient test case generation algorithms to magnify the similarities (or differences) measured by the testing metrics between the two models. Finally, a ‘yes’/‘no’ (stolen copy) judgment will be made for the suspect model based on all similarity scores.

The advantages of DeepJudge include 1) non-invasive: it works directly on the trained models and does not tamper with the training process; 2) efficient: it can be done very efficiently with only a few seed examples and a quick scan of the models; 3) flexible: it can easily incorporate new testing metrics or test case generation methods to obtain more evidence and reliable judgement, and can be applied in both white-box and black-box scenarios with different testing metrics; 4) robust: it is fairly robust to adaptive attacks such as model extraction and defense-aware attacks. The above advantages make DeepJudge a practical, flexible, and extensible tool for copyright protection of deep learning models.

We have implemented DeepJudge as an open-source self-contained toolkit and evaluated DeepJudge on four benchmark datasets (i.e., MNIST, CIFAR-10, ImageNet and Speech Commands) with different DNN architectures, including both convolutional and recurrent neural networks. The results confirm the effectiveness of DeepJudge in providing strong evidence for identifying the stolen copies of a victim model. DeepJudge is also proven to be more robust to a set of adaptive attacks compared to existing defense techniques.

We propose a novel testing framework DeepJudge for copyright protection of deep learning models. DeepJudge determines whether one model is a copy of the other depending on the similarity scores obtained from a comprehensive set of testing metrics and test case generation algorithms.

We identify three typical scenarios of model copying including finetuned copy, pruned copy, and extracted copy; define positive and negative suspect models for each scenario; and consider both white-box and black-box protection settings. DeepJudge can produce reliable evidence and judgement to correctly identify the positive suspects across all scenarios and settings.

DeepJudge is a self-contained open-source tool for robust copyright protection of deep learning models and a strong complement to existing techniques. DeepJudge can be flexibly applied in different DNN copyright protection scenarios and is extensible to new testing metrics and test case generation algorithms.

II Background

A DNN classifier is a decision function f:X→Yf:X\rightarrow Y mapping an input x∈X{\bm{x}}\in X to a label y∈Y={1,2,⋯ ,C}y\in Y=\{1,2,\cdots,C\}, where CC is the total number of classes. It comprises of LL layers: {f1,f2,⋯ ,fL−1,fL},\{f^{1},f^{2},\cdots,f^{L-1},f^{L}\}, where f1f^{1} is the input layer, fLf^{L} is the probability output layer, and f2,⋯ ,fL−1f^{2},\cdots,f^{L-1} are the hidden layers. Each layer flf^{l} can be denoted by a collection of neurons: {nl,1,nl,2,⋯ ,nl,Nl},\{n_{l,1},n_{l,2},\cdots,n_{l,N_{l}}\}, where NlN_{l} is the total number of neurons at that layer. Each neuron is a computing unit that computes its output by applying a linear transformation followed by a non-linear operation to its input (i.e., output from the precedent layer). We use ϕl,i(x)\phi_{l,i}({\bm{x}}) to denote the function that returns the output of neuron nl,in_{l,i} for a given input x∈X{\bm{x}}\in X. Then, we have the output vector of layer flf^{l} (2≤l≤L2\leq l\leq L): fl(x)=⟨ϕl,1(x),ϕl,2(x),⋯ ,ϕl,Nl(x)⟩f^{l}({\bm{x}})=\left\langle\phi_{l,1}({\bm{x}}),\phi_{l,2}({\bm{x}}),\cdots,\phi_{l,N_{l}}({\bm{x}})\right\rangle. Finally, the output label f(x)f({\bm{x}}) is computed as f(x)=arg max⁡fL(x)f({\bm{x}})=\operatorname*{arg\,max}f^{L}({\bm{x}}).

II-B DNN Watermarking

A number of watermarking techniques have been proposed to protect the copyright of DNN models . Similar to traditional multimedia watermarking, DNN watermarking works in two steps: embedding and verification. In the embedding step, the owner embeds a secret watermark (e.g., a signature or a trigger set) into the model during the training process. Depending on how much knowledge of the model is available in the verification step, existing watermarking methods can be broadly categorized into two classes: a) white-box methods for the case when model parameters are available; and b) black-box methods when only predictions of the model can be acquired.

White-box watermarking embeds a pre-designed signature (e.g., a string of bits) into the parameter space of the model via certain regularization terms . The ownership could be claimed when the extracted signature from a suspect model is similar to that of the owner model. Black-box watermarking usually leverages backdoor attacks to implant a watermark pattern into the owner model by training the model with a set of backdoor examples (also known as the trigger set) relabeled to a secret class . The ownership can then be claimed when the defender queries the suspect model for examples attached with the watermark trigger and receives the secret class as predictions.

II-C DNN Fingerprinting

Recently, DNN fingerprinting techniques have been proposed to verify model ownership via two steps: fingerprint extraction and verification. According to the categorization rule for watermarking, fingerprinting methods are all black-box techniques. Moreover, they are non-invasive, which is in sharp contrast with watermarking techniques. Instead of modifying the training procedure to embed identities, fingerprinting directly extracts a unique feature or property of the owner model as its fingerprint (i.e., a unique identifier). The ownership can then be verified if the fingerprint of the owner model matches with that of the suspect model. For example, IPGuard leverages data points close to the classification boundary to fingerprint the boundary property of the owner model. A suspect model is determined to be a stolen copy of the owner model if it predicts the same labels for most boundary data points. proposes a Conferrable Ensemble Method (CEM) to craft conferrable (a subclass of transferable examples) adversarial examples to fingerprint the overlap between two models’ decision boundaries or adversarial subspaces. CEM fingerprinting demonstrates robustness to removal attacks including finetuning, pruning and extraction attacks, except several adapted attacks like adaptive transfer learning and adversarial training . It is the closest work to our DeepJudge. However, as a fingerprinting method, CEM targets uniqueness, while as a testing framework, our DeepJudge targets completeness, i.e., comprehensive characterization of a model with multi-level testing metrics and diverse test case generation methods. Note that CEM fingerprinting can be incorporated into our framework as a black-box metric.

III DNN Copyright Threat Model

We consider a typical attack-defense setting with two parties: the victim and the adversary. Here, the model owner is the victim who trains a DNN model (i.e., the victim model) using private resources. The adversary attempts to steal a copy of the victim model, which 1) mimics its functionality while 2) cannot be easily recognized as a copy. Following this setting, we identify three common threats to DNN copyright: 1) model finetuning, 2) model pruning, and 3) model extraction. The three threats are illustrated in the top row of Fig. 1.

Threat 1: Model Finetuning. In this case, we assume the adversary has full knowledge of the victim model, including model architecture and parameters, and has a small dataset to finetune the model . This occurs, for example, when the victim open-sourced the model for academic purposes only, but the adversary attempts to finetune the model to build commercial products.

Threat 2: Model Pruning. In this case, we also assume the adversary has full knowledge of the victim model’s architecture and parameters. Model pruning adversaries first prune the victim model using some pruning methods, then finetune the model using a small set of data .

Threat 3: Model Extraction. In this case, we assume the adversary can only query the victim model for predictions (i.e., the probability vector). The adversary may be aware of the architecture of the victim model but has no knowledge of the training data or model parameters. The goal of model extraction adversaries is to accurately steal the functionality of the victim model through the prediction API . To achieve this, the adversary first obtains an annotated dataset by querying the victim model for a set of auxiliary samples, then trains a copy of the victim model on the annotated dataset. The auxiliary samples can be selected from a public dataset or synthesized using some adaptive strategies .

IV Testing for DNN Copyright Protection

In this section, we present DeepJudge, the proposed testing framework that produces supporting evidence to determine whether a suspect model is a copy of a victim model. The victim model can be copied by model finetuning, pruning, or extraction, as discussed in Section III. We identify the following criteria for a reliable copyright protection method:

Fidelity. The protection or ownership verification process should not affect the utility of the owner model.

Effectiveness. The verification should have high precision and recall in identifying stolen model copies.

Efficiency. The verification process should be efficient, e.g., taking much less time than model training.

Robustness. The protection should be resilient to adaptive attacks.

DeepJudge is a testing framework designed to satisfy all the above criteria. In the following three subsections, we will first give an overview of DeepJudge, then introduce its multi-level testing metrics and test case generation algorithms.

As illustrated in the bottom row of Fig. 1, DeepJudge consists of two components and a final judgement step: i) test case generation, ii) a set of multi-level distance metrics for testing, and iii) a thresholding and voting based judgement mechanism. Alg. 1 depicts the complete procedure of DeepJudge with pseudocode. It takes the victim model O\mathcal{O}, a suspect model S\mathcal{S}, and a set of data D\mathcal{D} associated with the victim model as inputs and returns the values of the testing metrics as evidence as well as the final judgement. The set of data D\mathcal{D} can be provided by the owner from either the training or testing set of the victim model. At the test case generation step, it selects a set of seeds from the input dataset D\mathcal{D} (Line 1) and carefully generates a set of extreme test cases from the seeds (Line 2). Based on the test cases generated, DeepJudge computes the distance (dissimilarity) scores defined by the testing metrics between the suspect and victim models (Line 3). The final judgement of whether the suspect is a copy of the victim can be made via a thresholding and voting mechanism according to the dissimilarity scores between the victim and a set of negative suspect models (Line 4).

IV-B Multi-level Testing Metrics

We first introduce the testing metrics for two different settings respectively: white-box and black-box. 1) White-box Setting: In this setting, DeepJudge has full access to the internals (i.e., intermediate layer outputs) and the final probability vectors of the suspect model S\mathcal{S}. 2) Black-box Setting: In this setting, DeepJudge can only query the suspect model S\mathcal{S} to obtain the probability vectors or the predicted labels. In both settings, we assume the model owner is willing to provide full access to the victim model O\mathcal{O}, including the training and test datasets, and the training details if necessary.

The proposed testing metrics are summarized in Table I, with their suitable defense settings highlighted in the last column. DeepJudge advocates evidence-based ownership verification of DNNs via multi-level testing metrics that complement each other to produce more reliable judgement.

There is an abundant set of model properties that could be used to characterize the similarities between two models, such as the adversarial robustness property and the fairness property . Here, we consider the former and define the robustness distance to measure the adversarial robustness discrepancy between two models on the same set of test cases. We will test more properties in our future work.

Denote the function represented by the victim model O\mathcal{O} by ff, given an input xi{\bm{x}}_{i} and its ground truth label yiy_{i}, an adversarial example xi′{{\bm{x}}^{\prime}_{i}} can be crafted by slightly perturbing xi{\bm{x}}_{i} towards maximizing the classification error of ff. This process is known as the adversarial attack, and f(xi′)≠yif({\bm{x}}^{\prime}_{i})\neq y_{i} indicates a successful attack. Adversarial examples can be generated using any existing adversarial attack methods such as FGSM and PGD . Given a set of test cases, we can obtain its adversarial version T={x1′,x2′,⋯ }T=\{{\bm{x}}^{\prime}_{1},{\bm{x}}^{\prime}_{2},\cdots\}, where xi′{\bm{x}}^{\prime}_{i} denotes the adversarial example of xi{\bm{x}}_{i}. The robustness property of model ff can then be defined as its accuracy on TT:

Robustness Distance (RobD). Let f^\hat{f} be the suspect model, we define the robustness distance between ff and f^\hat{f} by the absolute difference between the two models’ robustness:

The intuition behind RobD is that model robustness is closely related to the decision boundary learned by the model through its unique optimization process, and should be considered as a type of fingerprint of the model. RobD requires minimal knowledge of the model (only its output labels).

IV-B2 Neuron-level metrics

Neuron-level metrics are suitable for white-box testing scenarios where the internal layers’ output of the model is accessible. Intuitively, the output of each neuron in a model follows its own statistical distribution, and the neuron outputs in different models should vary. Motivated by this, DeepJudge uses the output status of neurons to capture the difference between two models and defines the following two neuron-level metrics NOD and NAD.

Neuron Output Distance (NOD). For a particular neuron nl,in_{l,i} with ll being the layer index and ii being the neuron index within the layer, we denote the neuron output function of the owner’s victim model and the suspect copy model by ϕl,i\phi_{l,i} and ϕ^l,i\hat{\phi}_{l,i} respectively. NOD measures the average neuron output difference between the two models over a given set T={x1,x2,⋯ }T=\{{\bm{x}}_{1},{\bm{x}}_{2},\cdots\} of test cases:

Neuron Activation Distance (NAD). Inspired by the Neuron Coverage for testing deep learning models, NAD measures the difference in activation status (‘activated’ vs. ‘not activated’) between the neurons of two models. Specifically, for a given test case x∈T{\bm{x}}\in T, the neuron nl,in_{l,i} is determined to be ‘activated’ if its output value ϕl,i(x)\phi_{l,i}({\bm{x}}) is larger than a pre-specified threshold. The NAD between the two models with respect to neuron nl,in_{l,i} can then be calculated as:

where the step function S(ϕl,i(x))S(\phi_{l,i}({\bm{x}})) returns 11 if ϕl,i(x)\phi_{l,i}({\bm{x}}) is greater than a certain threshold, otherwise.

IV-B3 Layer-level metrics

The layer-wise metrics in DeepJudge take into account the output values of the entire layer in a DNN model. Compared with neuron-level metrics, layer-level metrics provide a full-scale view of the intermediate layer output difference between two models.

Layer Output Distance (LOD). Given a layer index ll, let flf^{l} and f^l\hat{f}^{l} represent the layer output functions of the victim model and the suspect model, respectively. LOD measures the LpL^{p}-norm distance between the two models’ layer outputs:

where ∣∣⋅∣∣p||\cdot||_{p} denotes the LpL^{p}-norm (p=2p=2 in our experiments).

Layer Activation Distance (LAD). LAD measures the average NAD of all neurons within the same layer:

where NlN_{l} is the total number of neurons at the ll-th layer, and ϕl,i\phi_{l,i} and ϕ^l,i\hat{\phi}_{l,i} are the neuron output functions from flf^{l} and f^l\hat{f}^{l}.

Jensen-Shanon Distance (JSD). JSD is a metric that measures the similarly of two probability distributions, and a small JSD value implies the two distributions are very similar. Let fLf^{L} and f^L\hat{f}^{L} denote the output functions (output layer) of the victim model and the suspect model, respectively. Here, we apply JSD to the output layer as follows:

where q=(fL(x)+f^L(x))/2q=(f^{L}({\bm{x}})+\hat{f}^{L}({\bm{x}}))/2 and K(⋅,⋅)K(\cdot,\cdot) is the Kullback-Leibler divergence. JSD quantifies the similarity between two models’ output distributions, and is particularly more powerful against model extraction attacks where the suspect model is extracted based on the probability vectors (distributions) returned by the victim model.

IV-C Test Case Generation

To fully exercise the testing metrics defined above, we need to magnify the similarities between a positive suspect and the victim model, while minimizing the similarities of a negative suspect to the victim model. In DeepJudge, this is achieved by smart test case generation methods. Meanwhile, test case generation should respect the model accessibility in different defense settings, i.e., black-box vs. white-box.

When only the input and output of a suspect model are accessible, we populate the test set TT using adversarial inputs generated by existing adversarial attack methods on the victim model. We consider three widely used adversarial attack methods, including Fast Gradient Sign Method (FGSM) , Projected Gradient Descent (PGD) , and Carlini & Wagner’s (CW) attack , where FGSM and PGD are L∞L^{\infty}-bounded adversarial methods, and CW is an L2L^{2}-bounded attack method. This gives us more diverse test cases with both L∞L^{\infty}- and L2L^{2}-norm perturbed adversarial test cases. The detailed description and exact parameters used for adversarial test case generation are provided in Appendix A-D.

Fig. 2 illustrates the rationale behind using adversarial examples as test cases. Finetuned and pruned model copies are directly derived from the victim model, thus they share similar decision boundaries (purple line) as the victim model. However, the negative suspect models are trained from scratch on different data or with different initializations, thus having minimum or no overlapping with the victim model’s decision boundary. By subverting the model’s predictions, adversarial examples cross the decision boundary from one side to the other (we use untargeted adversarial examples). Although the extracted models by model extraction attacks are trained from scratch by the adversary, the training relies on the probability vectors returned by the victim model, which contains information about the decision boundary. This implies that the extracted model will gradually mimic the decision boundary of the victim model. From this perspective, the decision boundary (or robustness) based testing imposes a dilemma to model extraction adversaries: the better the extraction, the more similar the extracted model to the victim model, and the easier it to be identified by our decision boundary based testing.

IV-C2 White-box setting

In this case, the internals of the suspect model are accessible, thus a more fine-grained approach for test case generation becomes feasible. As shown in Fig. 3, given a seed input and a specified layer, DeepJudge generates one test case for each neuron, and the corner case of the neuron’s activation is of our particular interest.

The test generation algorithm is described in Alg. 2. It takes the owner’s victim model O\mathcal{O} and a set of selected seeds SeedsSeeds as input, and returns the set of generated test cases TT. TT is initialized to be empty (Line 1). The main content of the algorithm is a nested loop (Lines 2-13), in which for each neuron nl,in_{l,i} required by the metrics in Section IV-B, the algorithm searches for an input that activates the neuron’s output ϕl,i(x′)\phi_{l,i}({\bm{x}}^{\prime}) more than a threshold value. At each outer loop iteration, an input is sampled from the SeedsSeeds (Lines 3-4). It is then iteratively perturbed in the inner loop following the neuron activation’s gradient update (Lines 6-7), until an input x′{\bm{x}}^{\prime} that can satisfy the threshold condition is found and is added into the test suite TT (Lines 8-11) or when the maximum number of iterations is reached. The parameter lrlr (Line 7) is used to control that the search space of the input with perturbation is close enough to its seed input. Finally, the generated test suite TT is returned.

We discuss how to configure the specific threshold kk for a neuron nl,in_{l,i} used in Alg. 2. Since the statistics may vary across different neurons, we pre-compute the threshold kk based on the training data and the owner model, which is the maximum value (upper bound) of the corresponding neuron output over all training samples. The final threshold value is then adjusted by a hyper-parameter mm to be used in Alg. 2 for more reasonable and adaptive thresholds. Note that the thresholds for all interested neurons can be calculated once by populating layer by layer across the model.

IV-D Final Judgement

The judgment mechanism of DeepJudge has two steps: thresholding and voting. The thresholding step determines a proper threshold for each testing metric based on the statistics of a set of negative suspect models (see Section V-A3 for more details). The voting step examines a suspect model against each testing metric, and gives it a positive vote if its distance to the victim model is lower than the threshold of that metric. The lower a measured metric value, the more likely the suspect model is a copy of the victim, according to this metric. The final judgment can then be made based on the votes: the suspect model will be identified as a positive suspect if it receives more positive votes, and a negative suspect otherwise.

For each testing metric λ\lambda, we set the threshold adaptively using an ε\varepsilon-difference strategy. Specifically, we use one-tailed T-test to calculate the lower bound LBλLB_{\lambda} based on the statistics of the negative suspect models at the 99% confidence level. If the measured difference λ(O,S,T)\lambda(\mathcal{O},\mathcal{S},T) is lower than LBλLB_{\lambda}, S\mathcal{S} will be a copy of O\mathcal{O} with high probability. The threshold for each metric λ\lambda is defined as: τλ=αλ⋅LBλ\tau_{\lambda}=\alpha_{\lambda}\cdot LB_{\lambda}, where αλ\alpha_{\lambda} is a user-specified relaxing parameter controlling the sensitivity of the judgement. As αλ\alpha_{\lambda} decreases, the false positive rate (the possibility of misclassifying a negative suspect as a stolen copy) will also increase. We empirically set α=0.9\alpha=0.9 for black-box metrics and α=0.6\alpha=0.6 for white-box metrics respectively, depending on the negative statistics.

DeepJudge makes the final judgement by voting below:

IV-E DeepJudge vs. Watermarking & Fingerprinting

Here, we briefly discuss why our testing approach is more favorable in certain settings and how it complements existing defense techniques. Table II summarizes the differences of DeepJudge to existing watermarking and fingerprinting methods, from three aspects: 1) whether the method is non-invasive (i.e., independent of model training); 2) whether it is particularly designed for or evaluated in different defense settings (i.e., white-box vs. black-box); and 3) whether the method is evaluated against different attacks (i.e., finetuning, pruning and extraction). DeepJudge is training-independent, able to be flexibly applied in either white-box or black-box settings, and evaluated (also proven to be robust) against all three types of common copyright attacks including model finetuning, pruning and extraction, with empirical evaluations and comparisons deferred to Section V-B3.

Watermarking is invasive (training-dependent), whereas fingerprinting and testing are non-invasive (training independent). The effectiveness of watermarking depends on how well the owner model memorizes the watermark and how robust the memorization is to different attacks. While watermarking can be robust to finetuning or pruning attacks , it is particularly vulnerable to the emerging model extraction attack (see Section V-C2). This is because model extraction attacks only extract the key functionality of the model, however, watermarks are often task-irrelevant. Despite the above weaknesses, watermarking is the only technique that can embed the owner identity/signature into the model, which is beyond the functionalities of fingerprinting or testing.

Fingerprinting shares certain similarities with testing. However, they differ in their goals. Fingerprinting aims for “uniqueness”, i.e., a unique fingerprint of the model, while testing aims for “completeness”, i.e., to test as many dimensions as possible to characterize not only the unique but also the common properties of the model. Arguably, effective fingerprints are also valid black-box testing metrics. But as a testing framework, our DeepJudge is not restricted to a particular metric or test case generation method. Our experiments in Section VI show that a single metric or fingerprint is not sufficient to handle the diverse and adaptive model stealing attacks. In Section VI-B, we will also show that our DeepJudge can survive those adaptive attacks that break fingerprinting by dynamically changing the test case generation strategy. We anticipate a long-running arms race in deep learning copyright protection between model owners and adversaries, where watermarking, fingerprinting and testing methods are all important for a comprehensive defense.

V Experiments

We have implemented DeepJudge as a self-contained toolkit in PythonThe tool and all the data in the experiment are publicly available via https://github.com/Testing4AI/DeepJudge. In the following, we first evaluate the performance of DeepJudge against model finetuning and model pruning (Section V-B), which are two threat scenarios extensively studied by watermarking methods . We then examine DeepJudge against more challenging model extraction attacks in Section V-C. Finally, we test the robustness of DeepJudge under adaptive attacks in Section VI. Overall, we evaluated DeepJudge with 11 attack methods, 3 baselines, and over 300 deep learning models trained on 4 datasets.

We run the experiments on three image classification datasets (i.e., MNIST , CIFAR-10 and ImageNet ) and one audio recognition dataset (i.e., SpeechCommands ). The models used for the four datasets are summarized in Table III, including three convolutional architectures and one recurrent neural network. For each dataset, we divide the training data into two subsets. The first subset (50% of the training examples) is used to train the victim model. More detailed experimental settings can be found in Appendix A-A.

V-A2 Positive suspect models

Positive suspect models are derived from the victim models via finetuning, pruning, or model extraction. These models are considered as stolen copies of the owner’s victim model. DeepJudge should provide evidence for the victim to claim ownership.

V-A3 Negative suspect models

Negative suspect models have the same architecture as the victim models but are trained independently using either the remaining 50% of training data or the same data but with different random initializations. The negative suspect models serve as the control group to show that DeepJudge will not claim wrong ownership. These models are also used to compute the testing thresholds (τ\tau). The same training pipeline and the setting are used to train the negative suspect models. Specifically, “Neg-1” are trained with different random initializations while “Neg-2” are trained using a separate dataset (the other 50% of training samples).

V-A4 Seed selection

Seed selection prepares the SeedsSeeds examples used to generate the test cases. Here, we apply the sampling strategy used in DeepGini to select a set of high-confidence seeds from the test dataset (details are in Appendix A-B). The intuition is that high-confidence seeds are well-learned by the victim model, thus carrying more unique features of the victim model. More adaptive seed selection strategies are explored in the adaptive attack section VI-B1.

V-A5 Adversarial example generation

We use three classic attacks including FGSM , PGD and CW to generate adversarial test cases as introduced in Section IV-C1.

V-B Defending Against Model Finetuning & Pruning

As model finetuning and pruning threats are similar in processing the victim model (see Section III), we discuss them together here. These two are also the most extensively studied threats in prior watermarking works .

Given a victim model and a small set of data in the same task domain, we consider the following four commonly used model finetuning & pruning strategies: a) Finetune the last layer (FT-LL). Update the parameters of the last layer while freezing all other layers. b) Finetune all layers (FT-AL). Update the parameters of the entire model. c) Retrain all layers (RT-AL). Re-initialize the parameters of the last layer then update the parameters of the entire model. d) Parameter pruning (P-r%). Prune rr percentage of the parameters that have the smallest absolute values, then finetune the pruned model to restore the accuracy. We test both low (rr=20%20\%) and high (rr=60%60\%) pruning rates. Typical data-augmentations are also used to strengthen the attacks. More details of these attacks are in Appendix A-C.

V-B2 Effectiveness of DeepJudge

The results are presented separately for black-box vs. white-box settings.

Black-box Testing. In this setting, only the output probabilities of the suspect model are accessible. Here, DeepJudge uses the two black-box metrics: RobD and JSD. For both metrics, the smaller the value, the more similar the suspect model is to the victim model. Table IV reports the results of DeepJudge on the four datasets. Note that we randomly repeat the experiment 6 times for each finetuning or pruning attack and 12 times for independent training (as more negative suspect models will result in a more accurate judging threshold). Then, we report the average and standard deviation (in the form of a±ba\pm b) in each entry of Table IV. Clearly, all positive suspect models are more similar to the victim model with significantly smaller RobD and JSD values than negative suspect models. Specifically, a low RobD value indicates that the adversarial examples generated on the victim model have a high transferability to the suspect model, i.e., its decision boundary is closer to the victim model. In contrast, the RobD values of the negative suspect models are much larger than that of the positives, which matches our intuition in Fig. 2.

To further confirm the effectiveness of the proposed metrics, we show the ROC curve for a total of 54 models (30 positive suspect models and 24 negative suspect models) for RobD and JSD in Figure 4. The AUC values are 1 for both metrics. Note that we omit the plots for the following white-box testing as the AUC values for all metrics are also 1.

White-box Testing. In this setting, all intermediate-layer outputs of the suspect model are accessible. DeepJudge can thus use the four white-box metrics (i.e., NOD, NAD, LOD, and LAD) to test the models. Table V reports the results on the four datasets. Similar to the two black-box metrics, the smaller the white-box metrics, the more likely the suspect model is a stolen copy. As shown in Table V, there is a fundamental difference between the two sets (positive vs. negative) of suspect models according to each of the four metrics. That is, the two sets of models are completely separable, leading to highly accurate detection of the positive copies. It is not surprising as white-box testing can collect more fine-grained information from the suspect models. In both the black-box and white-box settings, the voting in DeepJudge overwhelmingly supports the correct final judgement (the ‘Copy?’ column).

Combined Visualization. To better understand the power of DeepJudge, we combine the black-box and white-box testing results for each suspect model into a single radar chart in Fig. 5. Each dimension of the radar chart corresponds to a similarity score given by one testing metric. For better visual effect, we normalize the values of the testing metrics into the range $$, and the larger the normalized value, the more similar the suspect model to the victim. Thus, the filled area could be viewed as the accumulated supporting evidence by DeepJudge metrics for determining whether the suspect model is a stolen copy. Clearly, DeepJudge is able to accurately distinguish positive suspects from negative ones. Among the positive suspect models, the areas of RT-AL and P-60% are noticeably smaller than the other two, meaning they are harder to detect. This is because these two attacks make the most parameter modifications to the victim model. Comparing the metrics, activation-based metrics (e.g., NAD) demonstrate better performance than output-based metrics (e.g., NOD), while white-box metrics are stronger than black-box metrics, especially against strong attacks like RT-AL. In Appendix A-D, we also analyze the influencing factors including adversarial test case generation and layer selection (for computing the testing metrics) via several calibration experiments. An analysis of how different levels of finetuning or pruning affect DeepJudge is presented in Appendix A-H.

Time Cost of DeepJudge. The time cost of generating test cases using 1k seeds is provided in appendix Table IX. For the black-box setting, we report the cost of PGD-based generation, while for the white-box setting, we report that of Algorithm 2. It shows that the time cost of white-box generation is slightly higher but is still very efficient in practice. The maximum time cost occurs on the SpeechCommands dataset for white-box generation, which is ∼1.2\sim 1.2 hours. This time cost is regarded as efficient since test case generation is a one-time effort, and the additional time cost of scanning a suspect model with the test cases is almost negligible.

V-B3 Comparison with existing techniques

We compare DeepJudge with three state-of-the-art copyright defense methods against model finetuning and pruning attacks. More details of these defense methods can be found in Appendix A-E.

Black-box: Comparison to Watermarking and Fingerprinting. DNNWatermarking is a black-box watermarking method based on backdoors, and IPGuard is a black-box fingerprinting method based on targeted adversarial attacks. Here, we compare these two baselines with DeepJudge in the black-box setting. For DNNWatermarking, we train the watermarked model (i.e., victim model) using additionally patched samples from scratch to embed the watermarks, and the TSA (Trigger Set Accuracy) of the suspect model is calculated for ownership verification. IPGuard first generates targeted adversarial examples for the watermarked model then calculates the MR (Matching Rate) (between the victim and the suspect) for verification. For DeepJudge, we only apply the RobD (robustness distance) metric here for a fair comparison.

The left subfigure of Fig. 6 visualizes the results. DeepJudge demonstrates the best overall performance in this black-box setting. DNNWatermarking and IPGuard fail to identify the positive suspect models duplicated by FT-AL, RT-AL, P-20% and P-60%. Their scores (TSA and MR) drop drastically against these four attacks. This basically means that the embedded watermarks are completely removed, or the fingerprint can no longer be verified. While for the RobD metric of DeepJudge, the gap remains huge between the negative and positive suspects, demonstrating much better effectiveness to diverse finetuning and pruning attacks.

White-box: Comparison to Watermarking. EmbeddingWatermark is a white-box watermarking method based on signatures. It requires access to model parameters for signature extraction. We train the victim model with the embedding regularizer from scratch to embed a 128-bits signature. The BER (Bit Error Rate) is calculated and used to measure the verification performance. The right subfigure of Fig. 6 visualizes the comparison results to two white-box DeepJudge metrics NOD and NAD. The three metrics demonstrate a comparable performance with NAD wins on 4 out of the 5 positive suspects. Note that the huge gap between the positives and negatives indicates that all metrics can correctly identify the positive suspects. Here, a single metric of DeepJudge was able to achieve the same level of protection as EmbeddingWatermark.

V-C Defending Against Model Extraction

Model extraction (also known as model stealing) is considered to be a more challenging threat to DNN copyright. In this part, we evaluate DeepJudge against model extraction attacks, which has not been thoroughly studied in prior work.

We consider model extraction with two different types of supporting data: auxiliary or synthetic (see Section III). We consider the following state-of-the-art model extraction attacks: a) JBA (Jacobian-Based Augmentation ) samples a set of seeds from the test dataset, then applies Jacobian-based data augmentation to synthesize more data from the seeds. b) Knockoff (Knockoff Nets ) works with an auxiliary dataset that shares similar attributes as the original training data used to train the victim model. c) ESA (ES Attack ) requires no additional data but a huge amount of queries. ESA utilizes an adaptive gradient-based optimization algorithm to synthesize data from random noise. ESA could be applied in scenarios where it is hard to access the task domain data, such as personal health data. With the extracted data, the adversary can train a new model from scratch, assuming knowledge of the victim model’s architecture. The new model is considered as a successful stealing if its performance matches with the victim model.

V-C2 Failure of watermarking

Our experiments in Section V-B show the effectiveness and robustness of watermarking to finetuning and pruning attacks. Unfortunately, here we show that the embedded watermarks can be removed by model extraction attacks. We show the results of DNNWatermarking and EmbeddingWatermark in Fig. 12. The extracted models by different extraction attacks all differ greatly from the victim model according to either TSA (from DNNWatermarking) or BER (from EmbeddingWatermark). For example, the TSA value for the victim model is 100%, however, the TSA values for the three extracted copies are all below 1%. This basically means that the original watermarks are all erased in the extracted models. It will inevitably lead to failed ownership claims. This is somewhat not too surprising as watermarks are task-irrelevant contents and not the focus of model extraction.

V-C3 Effectiveness of DeepJudge

Table VI summarizes the results of DeepJudge, which successfully identifies all positive suspect models, except when the stolen copies (by JBA) have extremely poor performance with 15%, 44% and 55% lower accuracy than the corresponding victim model. We note that model extraction does not always work, and poorly performed extractions are less likely to pose a real threat. We also observe that DeepJudge works better when the extraction is better, which therefore counters the ultimate perfect matching goal of model extraction attacks.

Compared to model finetuning or pruning, the average RobD and JSD values on extracted models are relatively larger, meaning that the decision boundaries of extracted models are more different from that of the victim model. The reason is that extracted models are often trained from a random point, while finetuning only slightly shifts the original boundary of the victim model, as depicted in Fig. 2. As such, model extraction is more stealthy and more challenging for ownership verification. Nonetheless, the two metrics RobD and JSD, can still reveal the unique similarities (smaller values) of the extracted models to the victim model: the better the extraction (higher accuracy of the extracted model), the lower the RobD and JSD values. This indicates that the extracted model behaves more similarly to the victim as its decision boundary gradually approaching that of the victim, and also highlights the unique advantage of DeepJudge against model extraction attacks. Note that JBA attack can only extract 50% of the original accuracy on either CIFAR-10 or SpeechCommands, which should not be considered as successful extractions.

In Fig. 7, we further show the evolution of the RobD and JSD values throughout the entire extraction process of Knockoff, ESA and JBA attacks. We find that both RobD (orange line) and JSD (red line) values decrease as the extraction progresses, again, except for JBA. This confirms our speculation that, when tested by DeepJudge, a better extracted model will expose more similarities to its victim. By contrast, we also study how these two values change during the training process of the negative model in Fig. 7, which shows that the independently trained negative suspect models tend to vary more from the victim model and produce higher RobD and JSD values.

VI Robustness to Adaptive Attackers

In this section, we explore potential adaptive attacks to DeepJudge based on the adversary’s knowledge of DeepJudge: 1) the adversary knows the testing metrics and the test cases, or 2) the adversary only knows the testing metrics. Contrast evaluation of watermarking & fingerprinting against similar adaptive attacks are in Appendix A-G.

In this threat model, the adversary has full knowledge of DeepJudge including the testing metrics Λ\Lambda and the secret test cases TT. We also assume the adversary has a subset of clean data. In DeepJudge, we have two test settings, i.e., white-box testing and black-box testing. The two testings differ in the testing metrics and the generated test cases (see examples in Fig. 17). The black-box test cases are labeled. Therefore, the adversary can mix TT into its clean subset to finetune the stolen model to have large testing distances (i.e., black-box testing metrics RobD and JSD) while maintaining good classification performance. This will fool DeepJudge to identify the stolen model to be significantly different from the victim model. This adaptive attack against black-box testing is denoted by Adapt-B. Since the white-box test cases are unlabeled, the adversary can use the predicted labels (by the victim model) as ground-truth and finetunes the stolen model following a similar procedure as Adapt-B. This attack against white-box testing is denoted by Adapt-W. Note that the suffix ‘-B/-W’ marks the target testing setting to attack, while both attacks are white-box adaptive attacks knowing all the information.

The results of DeepJudge using the exposed test cases TT are reported in Table VII. It shows that: 1) DeepJudge is robust to Adapt-W, which fails to maximize the output distance and activation distance simultaneously nor maintaining the original classification accuracy; 2) though DeepJudge is not robust to Adapt-B when the test cases are exposed with labels, it can easily recover the performance with new test cases generated with different seeds (see the ROC curves on the exposed and new test cases in Fig. 8); and 3) DeepJudge can still correctly identify the stolen copies by Adapt-B when combining black-box and white-box testings (the final judgements are all correct). Comparing the non-trivial effort of retraining/finetuning a model to the efficient generation of new test cases, DeepJudge holds a clear advantage in the arms race against finetuning-based adaptive attacks.

It is noteworthy that Adapt-W did not break all white-box metrics of DeepJudge, since the mechanism of white-box testing is robust. Specifically, black-box testing characterizes the behaviors of the output layer, while white-box testing characterizes the internal behaviors of more shallow layers. Due to the over-parameterization property of DNNs, it is relatively easy to fine-tune the model to overfit the set of black-box test cases, subverting the results of the black-box metrics. However, in white-box testing, changing the activation status of all hidden neurons on the set of white-box test cases is almost impossible without completely retraining the model. Therefore, white-box testing is inherently more robust to adaptive attacks, especially when the test cases are exposed.

VI-B Knowing Only the Testing Metrics

In this threat model, the adversary can still adapt in different ways. We consider two adaptive attacks: adversarial training targeting on black-box testing and a general transfer learning attack on white-box testing, respectively.

Since our black-box testing mainly relies on probing the decision boundary difference using adversarial test cases, the adversary may utilize adversarial training to improve the robustness of the stolen copy. Given the PGD parameters and a subset of clean data (20% of the original training data), the adversary iteratively trains the stolen model to smooth the model decision boundaries following . This type of adaptive attack is denoted by Adv-Train. As Table VII shows, it can indeed circumvent our black-box testing, with a sacrifice of ∼10%\sim 10\% performance (a phenomenon known as accuracy-robustness trade-off ). However, interestingly, if we replace the high-confidence seeds used in DeepJudge with low-confidence seeds, DeepJudge becomes effective again (as shown in Fig. 9). One possible reason is that, compared to high-confidence seeds, these low-confidence seeds are natural boundary (hard) examples that are close to the decision boundary, thus can generate more test cases to cross the adversarially smoothed decision boundary within certain perturbation budget. Examples of high/low confidence test seeds are provided in Fig. 15. It is also worth mentioning that our white-box testing still performs well in this case. Overall, DeepJudge is robust to Adv-Train or at least can be made robust by efficiently updating the seeds.

VI-B2 Transfer learning

The adversary may transfer the stolen copy of the victim model to a new dataset. The adversary exploits the main structure of the victim model as a backbone and adds more layers to it. Here, we test a vanilla transfer learning (VTL) strategy from the 10-class CIFAR-10 to a 5-class SVHN . The last layer of the CIFAR-10 victim model is first replaced by a new classification layer. We then fine-tune all layers on the subset of SVHN data. Note that, in this setting, the black-box metrics are no longer feasible since the suspect model has different output dimensions to the victim model, however, the white-box metrics can still be applied since the shallow layers are kept. The results are reported in Table VII. Remarkably, DeepJudge succeeds in identifying transfer learning attacks with distinctively low testing distances and an AUC=1AUC=1.

In one recent work , it was observed that the knowledge of the victim model could be transferred to the stolen models. Dataset Inference (DI) technique was then proposed to probe whether the victim’s knowledge (i.e., private training data) is preserved in the suspect model. We believe such knowledge-level testing metrics could also be incorporated into DeepJudge to make it more comprehensive. An analysis of how different levels of transfer learning could affect DeepJudge can be found in Appendix A-H.

VII Conclusion

In this work, we proposed DeepJudge, a novel testing framework for copyright protection of deep learning models. The core of DeepJudge is a family of multi-level testing metrics that characterize different aspects of similarities between the victim model and a suspect model. Efficient and flexible test case generation methods are also developed in DeepJudge to help boost the discriminating power of the testing metrics. Compared to watermarking methods, DeepJudge does not need to tamper with the model training process. Compared to fingerprinting methods, it can defend more diverse attacks and is more resistant to adaptive attacks. DeepJudge is applicable in both black-box and white-box settings against model finetuning, pruning and extraction attacks. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness and efficiency of DeepJudge. We have implemented DeepJudge as a self-contained open-source toolkit. As a generic testing framework, new testing metrics or test case generation methods can be effortlessly incorporated into DeepJudge to help defend future threats to deep learning copyright protection.

Acknowledgement

We are grateful to the anonymous reviewers and shepherd for their valuable comments. This research was supported by the Key R&D Program of Zhejiang (2022C01018) and the NSFC Program (62102359, 61833015).

References

Appendix A Appendix

We use four benchmark datasets from two domains for the evaluation:

MNIST . This is a handwritten digits (from 0 to 9) dataset, consisting of 70,000 images with size 28×28×128\times 28\times 1, of which 60,000 and 10,000 are training and test data.

CIFAR-10 . This is a 10-class image classification dataset, consisting of 60,000 images with size 32×32×332\times 32\times 3, of which 50,000 and 10,000 are training and testing data.

ImageNet . This is a large-scale image dataset containing more than 1.2 million training images of 1,000 categories. It is more challenging due to the higher image resolution 224×224×3224\times 224\times 3. We randomly sample 100 classes to construct a subset of ImageNet, of which 120,000 are training data and 30,000 are testing data.

Speech Commands . This is an audio dataset of 10 single spoken words, consisting of about 40,000 training samples and 4,000 testing samples. We pre-processed the data to obtain a Mel Spectrogram . Each audio sample is transformed into an array of size 120×85120\times 85.

To explore the scalability of DeepJudge, various model structures are tested as in Table III. LeNet-5, ResNet-20 and VGG-16 are standard CNN structures, while LSTM(128) is an RNN structure: an LSTM layer with 128 hidden units, followed by three fully-connected layers (128/64/10).

A-B Seed Selection Strategy

Seed selection is important for generating high-quality test cases. We use DeepGini to measure the certainty of each candidate sample. Given the victim model ff and a testing dataset D\mathcal{D}, we first calculate the Certainty Score (CS) for each seed x∈D{\bm{x}}\in\mathcal{D} as: CS(fL,x)=∑iCfiL(x)2,CS(f^{L},{\bm{x}})=\sum_{i}^{C}f^{L}_{i}({\bm{x}})^{2}, then we rank the seed list by the certainty score, and the first part of the seeds of the highest scores (i.e., most certainties) will be chosen for the following generation process. Here, we assume to have two seeds {x1,x2}\{{\bm{x}}_{1},{\bm{x}}_{2}\} with CS(fL,x1)>CS(fL,x2)CS(f^{L},{\bm{x}}_{1})>CS(f^{L},{\bm{x}}_{2}), that means the victim model ff is more confident at x1{\bm{x}}_{1}, which also means that x1{\bm{x}}_{1} is farther from the decision boundary and easier for classification (see examples in Fig. 15).

A-C Data-augmentation

During the finetuning and pruning processes, typical data-augmentation techniques are used to strengthen the attacks except for the SpeechCommands dataset, including random rotation (10∘10^{\circ}), random width- and height-shift (both 0.10.1).

A-D Test Case Generation Details and Calibrations

Specifically, we consider three adversarial attacks for generating black-box test cases (see Section IV-C1).

FGSM perturbs a normal example x{\bm{x}} by one single step of size ϵ\epsilon to maximize the model’s prediction error with respect to the groundtruth label yy:x′=x+ϵ⋅sign⁡(▽xL(fL(x),y)){\bm{x}}^{\prime}={\bm{x}}+\epsilon\cdot\operatorname{sign}(\bigtriangledown_{{\bm{x}}}\mathcal{L}(f^{L}({\bm{x}}),y)), where sign⁡(⋅)\operatorname{sign}(\cdot) is the sign function, L\mathcal{L} is the cross entropy (CE) loss, and ▽xL\bigtriangledown_{{\bm{x}}}\mathcal{L} is the gradient of the loss to the input.

PGD is an iterative version of FGSM but with smaller step size: xk=Πϵ(xk−1+α⋅sign⁡(▽xL(fL(xk−1),y))){\bm{x}}^{k}=\Pi_{\epsilon}({\bm{x}}^{k-1}+\alpha\cdot\operatorname{sign}(\bigtriangledown_{{\bm{x}}}\mathcal{L}(f^{L}({\bm{x}}^{k-1}),y))), where xk{\bm{x}}^{k} is the adversarial example obtained at the kk-th perturbation step, α\alpha is the step size and Πϵ\Pi_{\epsilon} is a projection (clipping) operation that projects the perturbation back onto the ϵ\epsilon-ball centered around x{\bm{x}} if it goes beyond.

CW generates adversarial examples by solving the optimization problem: x′=min⁡x′∥x′−x∥22−c⋅L(fL(x′),y){\bm{x}}^{\prime}=\underset{{\bm{x}}^{\prime}}{\min}\left\lVert{\bm{x}}^{\prime}-{\bm{x}}\right\rVert_{2}^{2}-c\cdot\mathcal{L}(f^{L}({\bm{x}}^{\prime}),y), where cc is a hyperparameter balancing the two terms and the pixel values of adversarial example x′{\bm{x}}^{\prime} are bounded to be within a legitimate range, e.g., $$ for 0-1 normalized input.

The hyper-parameters used for the generation algorithms on different datasets are summarized in Table VIII. Here, we take CIFAR-10 dataset as an example and analyze the influencing factors of the test case generation process.

Adversarial Examples. PGD is the default choice for generating adversarial examples in the black-box setting. Here, we further compare PGD with two other methods, FGSM and CW. We use the same selected seeds for the generation. Table X shows the results of the RobD metric. We observe that the gap in RobD values between the positive and negative suspect models is very small when CW is used, which fails to distinguish the two types of models. One reason is that CW attack optimizes adversarial examples for minimal perturbations, which is more sensitive (less robust) to model modifications. FGSM can be regarded as a one-step PGD, which usually has a larger average perturbation than PGD. When the perturbation increases, the RobD value of negative suspect models would decrease since adversarial examples with larger perturbations tend to have better transferability . It is similar to PGD3ϵ when the perturbation bound increases. In general, the absolute RobD gap between the positive and negative suspects tested with PGD-generated test cases is larger than that of FGSM and CW. Moreover, PGD is relatively cheaper to calculate than CW, i.e., the time cost of PGD is 100×100\times lower than CW. Overall, PGD is more suitable for fingerprinting the decision boundary with untargeted adversarial examples, as shown in Fig. 2. We will explore more effective metrics and test case generation methods with diverse granularity in future work.

Remark 5: Different generation strategies and parameters can impact DeepJudge differently. Overall, PGD is a better choice for characterizing the model’s decision boundary. Layer Selection. Layer selection is important when applying DeepJudge in the white-box setting with NOD and NAD. Here, we evaluate how the choice of layers affects the performance of DeepJudge. For comparison, we choose a shallow layer and a deep layer of the victim model, and re-generate the test cases respectively for each layer. Fig. 10 shows the results of NOD and NAD metrics. In general, the NOD/NAD difference between the positive and negative suspect models becomes much larger at the shallow layer. The reason is that the shallow layers of a network usually learn the low-level features , and they tend to stay the same or at least similar during model finetuning. Particularly, the performance on RT-AL degrades the most when the deep layer is selected, since the parameters of the last layer are re-initialized. Thus, choosing the shallow layers to compute the NOD and NAD metrics could help the robustness of DeepJudge. Moreover, the time cost of generating and testing with the shallow layer is 10×\times less than the deep layer, since most of the back-propagation computations are eliminated.

A-E Defense Baselines

embeds backdoors into the model. In our experiments, we select 500 samples from the training dataset, of which the ground truth labels are “automobile”. Then we patch an “apple” logo at the bottom right corner of each sample and change their labels to “cat” (see Fig. 11). These trigger examples (i.e., trigger set) are mixed into the clean training dataset to train a watermarked model from scratch. The initial TSA of the watermarked model is 100.0% (on a separate trigger set).

A-E2 Signature-based watermarking (White-box)

embeds a TT-bit vector (i.e., the watermark) b∈{0,1}Tb\in\{0,1\}^{T} into one of the convolutional layers, by adding an additional parameter regularizer into the loss function: E(w)=E0(w)+λER(w),E(w)=E_{0}(w)+\lambda E_{R}(w), where E0(w)E_{0}(w) is the original task loss function, ER(w)E_{R}(w) is the regularizer that imposes a certain restriction on the model parameters ww, and λ\lambda is a hyper-parameter. In our experiments, λ\lambda is set to 0.010.01, and we embed a 128-bit watermark (generated by the random strategy) into the second convolutional block (Conv-2 group) as recommended in . The initial BER of the watermarked model is 3.13%.

A-E3 Fingerprinting (Black-box)

IPGuard proposes a type of adversarial attack that targets on generating adversarial examples x′x^{\prime} around the classification boundaries of the victim model, and the matching rate (MR) of these key samples is calculated for the verification similar to . We generate a set of 1,000 adversarial examples following and the initial MR of the victim model on the generated key samples is 100.0%.

A-F Model Extraction Attacks

Jacobian-Based Augmentation. The seeds used for augmentation are all sampled from the testing dataset. We sample 150 seeds for extracting the MNIST victim model, 500 seeds for SpeechCommands, 1,000 seeds for CIFAR-10, and use all other default settings .

Knockoff Nets. We use the Fashion-MNIST dataset for extracting the MNIST victim model, an independent speech dataset for SpeechCommands, and CIFAR-100 for CIFAR-10. We use other default hyper-parameter settings of .

ES Attack. We use the OPT-SYN algorithm to heuristically synthesize the surrogate data. We set the stealing epoch to 50 for MNIST and 400 for CIFAR-10. We failed to extract the SpeechCommands Victim model since the validation accuracy could not exceed 20%. All other hyper-parameters are the same as in .

∗Functionality-equivalent Extraction. Besides the above three extraction attacks, we are also aware of the functionality-equivalent extraction attacks that attempt to obtain a precise functional approximation of the victim model. For instance, proposed a differential attack that could steal the parameters of the victim model up to floating-point precision without the knowledge of training data. We remark that defending this type of attack is a trivial task for DeepJudge as there will be no difference between the extracted model and the victim model in an ideal approximation.

Note that Black-box model extraction is still underexplored, and more extraction attacks may appear in the future. This poses a continuous challenge for deep learning copyright protection. We hope that DeepJudge could evolve with the adversaries by incorporating more advanced testing metrics and test case generation methods, and provide a possibility to fight against this continuing model stealing threat.

A-G Adaptive Attacks for Watermarking & Fingerprinting

In addition to Section VI, here we conduct an extra evaluation of existing watermarking and fingerprinting methods under similar adaptive attack settings.

Adaptive attacks. Adv-Train and VTL are the two adaptive attacks in Table VII, while the Adapt-X attack is specifically designed for each method as follows:

Adapt-X for . Since the embedded watermark (signature) is known, the adversary copies (steals) the victim model then fine-tunes it on a small subset of clean examples while maximizing the embedding loss ER(w)E_{R}(w) on the signature.

Adapt-X for . Since the embedded watermark (backdoor) is known, the adversary can follow a similar approach as above to steal the victim model and remove the backdoor watermark with a few backdoor-patched but correctly-labeled examples.

Adapt-X for . Similar to our Adapt-B for DeepJudge, the adversary copies the victim model then fine-tunes it on a small subset of clean and correctly-labeled fingerprint examples to circumvent fingerprinting.

As the results in Table XI show, all three methods are completely broken by the adaptive attacks. DeepJudge is the only method that can survive these attacks and was partially compromised but not fully broken (the final judgments are still correct, as shown in the ‘Copy?’ column of Table VII). This implies that a single metric of watermarking or fingerprinting is not sufficient enough to combat adaptive attacks. By contrast, a testing framework with comprehensive testing metrics and test case generation methods may have the required flexibility to address this challenge. For example, Adv-Train may break the black-box testing of DeepJudge but cannot break the white-box testing (see Section VI-B). Moreover, DeepJudge can quickly recover its performance by switching to a new set of seeds (see Fig. 9).

A-H How Different Levels of Finetuning, Pruning and Transfer Learning Affect DeepJudge?

There is a spectrum of building a new model with access to a victim model, from different ways of finetuning to transfer learning. Different levels of modifications to the victim model would accordingly influence the testing of DeepJudge in different ways. Intuitively, a larger modification would lead to more dissimilarity between the victim and suspect models and a larger metric distance.

Here, we test different proportions of training samples and learning rates used for finetuning, proportions of pruned weights (pruning ratios) for pruning, and proportions of samples used for transfer learning (w.r.t. the setting described in Section VI-B2). Fig. 16 shows the metrics’ values at different levels of finetuning, pruning and transfer learning. At a high level, black-box metrics (i.e., RobD and JSD) have higher normalized distances than white-box metrics (i.e., NOD and NAD) on average. This implies that the model’s decision boundary is more sensitive to almost all levels of modifications. For finetuning, the two black-box metrics (yellow and orange bars) increase significantly with the amount of finetuning samples or amplified learning rate, whereas the two white-box metrics are relatively stable. Note that ‘4x’ (4 times the default learning rate) causes a significant drop (∼20%\sim 20\%) in the model accuracy. For pruning, all metrics including the white-box metrics increase with the amount of pruned weights at a much higher rate than finetuning with different sample sizes. This indicates that pruning has more impact on the model than finetuning and will greatly distort the model’s internal activations (measured by the two white-box metrics NOD and NAD). Transfer learning has much higher metric values (only white-box metrics are applicable here) than finetuning. This is because the victim model’s functionality has been greatly altered by transferring to a new data distribution. However, it seems that the modification caused by transfer learning does not accumulate with more samples, resulting in similar metric values even with 40% more samples.

In this work, we follow the principle that any derivations from the victim model other than independent training should be treated as having a certain level of copying. However, it can be hard to judge what degree of similarity (or level of modification) should be considered as “real copying” in real-world scenarios. In DeepJudge, we introduced a good range of testing metrics, hoping to provide more comprehensive evidence for making the final judgement. Moreover, the final judgement mechanism of DeepJudge (Section IV-D) can be flexibly adjusted to suit different application needs.

A-I Additional Figures