DAWN: Dynamic Adversarial Watermarking of Neural Networks
Sebastian Szyller, Buse Gul Atli, Samuel Marchal, N. Asokan
Introduction
Recent progress in machine learning (ML) has led to a dramatic surge in the use of ML models for a wide variety of applications. Major enterprises like Google, Apple, and Facebook have already deployed ML models in their products (TechWorld, 2018). ML-related businesses are expected to generate trillions of dollars in revenue in the near future (Forbes, 2019). The process of collecting training data and training ML models is the basis of the business advantage of model owners. Hence, protecting the intellectual property (IP) embodied in ML models is necessary.
One approach for IP protection of ML models is watermarking. Recent work (Merrer et al., 2017; Adi et al., 2018; Zhang et al., 2018) has shown how digital watermarks can be embedded into deep neural network models (DNNs) during training. Watermarks consist of a set of inputs, the trigger set, with incorrectly assigned labels. A legitimate model owner can use the trigger set, along with a large training set with correct labels, to train a watermarked model and distribute it to his customers. If he later encounters a model he suspects to be a copy of his own, he can demonstrate ownership by using the trigger set as inputs to the suspected model. These watermarking schemes allow legitimate model owners to detect theft or misuse of their models.
Instead of distributing ML models to customers, an increasingly popular alternative business paradigm is to allow customers to use models via prediction APIs. But one can mount a model extraction (Tramèr et al., 2016) attack via such APIs by sending a sequence of API queries with different inputs and using the resulting predictions to train a surrogate model with similar functionality as the queried model. Model extraction attacks are effective even against complex DNN models (Juuti et al., 2019; Orekondy et al., 2019), and are difficult to prevent (Juuti et al., 2019). Existing watermarking techniques, which rely on model owners to embed watermarks during training, are ineffective against model extraction since it is the adversary who trains the surrogate model.
In this paper we introduce DAWN (Dynamic Adversarial Watermarking of Neural Networks), a new watermarking approach intended to deter IP theft via model extraction. DAWN is designed to be deployed within the prediction API of a model. It dynamically watermarks a tiny fraction of queries from a client by changing the prediction responses for them. The watermarked queries serve as the trigger set if an adversarial client trains a surrogate model using the responses to its queries. The model owner can use the trigger set to demonstrate IP ownership of the extracted surrogate model as in prior DNN watermarking solutions (Adi et al., 2018; Zhang et al., 2018). DAWN differs from them in that it is the adversary (model thief), rather than the defender (original owner) who trains the watermarked model. This raises two new challenges: (1) defenders must choose trigger sets from among queries sent by clients and cannot choose optimal trigger sets from the whole input space; (2) adversaries can select the training data or manipulate the training process to resist the embedding of watermarks. DAWN addresses both these challenges.
DAWN watermarks are client-specific: DAWN not only infers whether a given model is a surrogate but, in case of model extraction, also identifies the client whose queries were used to train the surrogate. DAWN is parametrized so that changed predictions needed for watermarking are sufficiently rare as to not degrade the utility of the original model for legitimate API clients.
present DAWN, the first approach for dynamic, selective watermarking for DNN models at their prediction APIs for deterring IP theft via model extraction (Sect. 4),
empirically assess it (Sect. 5) using several DNN models and datasets showing that DAWN is robust to adversarial manipulations and resilient to evasion (Sect. 6 and 8), and
show that DAWN is resistant to two state-of-the-art extraction attacks, reliably demonstrating ownership (with confidence >) with negligible impact on model utility (0.03-0.5% decrease in accuracy) (Sect. 7).
Code to reproduce our experiments is available on GitHub github.com/ssg-research/dawn-dynamic-adversarial-watermarking-of-neural-networks .
Background
In model extraction (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018; Pal et al., 2020), an adversary wants to “steal” a DNN model of a victim by making a series of prediction requests to and obtaining predictions . and are used by to train a surrogate model . ’s goal is to have as close as possible to . All model extraction attacks (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018) operate in a black-box setting: has access to a prediction API, uses the set to iteratively refine the accuracy of . Depending on the adversary model, ’s capabilities can be divided into three categories: model knowledge, data access, and querying strategy.
Model knowledge. does not know the exact architecture of or the hyperparameters or the training process. However, given the purpose of the API (e.g., image recognition) and expected complexity of the task, may attempt to guess the architecture of the model (Papernot et al., 2017; Juuti et al., 2019). On the other hand, if is complex, can use a publicly available, high capacity model pre-trained with a very large benchmark datasets (Orekondy et al., 2019). While the above methods focus on DNNs, there are alternatives targeting simpler models: logistic regression, decision trees, shallow neural networks (Tramèr et al., 2016).
Data access. ’s main limitation is the lack of access to natural data that comes from the same distribution as the data used to train . may use data that comes from the same domain as ’s training data but from a different distribution (Correia-Silva et al., 2018). If does not exactly know the distribution or the domain, it may use widely available natural data (Orekondy et al., 2019; Pal et al., 2020) to mount the attack. Alternatively, it may use only synthetic samples (Tramèr et al., 2016) or a mix of a small number of natural samples augmented by synthetic samples (Papernot et al., 2017; Juuti et al., 2019).
Querying strategy. All model stealing attacks (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018; Pal et al., 2020) consist of alternating phases of querying , followed by training the surrogate model using the obtained predictions. queries with all its data and then trains the surrogate model (Correia-Silva et al., 2018; Orekondy et al., 2019). Alternatively, if relies primarily on synthetic data (Papernot et al., 2017; Juuti et al., 2019), it deliberately crafts inputs that would help it train .
2. Watermarking DNN models
Digital watermarking is a technique used to covertly embed a marker, the watermark, in an object (image, audio, etc.) which can be used to demonstrate ownership of the object. Watermarking of DNN models leverages the massive overcapacity of DNNs and their ability to fit data with arbitrary labels (Zhang et al., 2017). DNNs have a large number of parameters, many of which have little significance for their primary classification task. These parameters can be used to carry additional information beyond what is required for its primary classification task. This property is exploited by backdooring attacks, which consist in training a DNN model that deliberately outputs incorrect predictions for some selected inputs (Chen et al., 2017; Gu et al., 2017).
The trigger set and the outputs of the backdoor function for its elements compose the watermark: . Let be a DNN model that copies . The watermark can be used to demonstrate ownership of . It only requires to expose a prediction API which can be used to query all samples in the trigger set . A sufficient number of predictions such that demonstrates that is a copy of the watermarked model .
Problem Statement
The adversary mounts a model extraction attack against a victim model using queries to its prediction API. ’s goal is model functionality stealing (Orekondy et al., 2019): train a surrogate model that performs well on a classification task for which was designed. If then ’s goal is that , which can be considered successful if . A secondary goal is to minimize the number of queries to necessary for to train .
has full control over the samples it chooses to query with. These can be natural (Orekondy et al., 2019) or synthetic (Juuti et al., 2019; Papernot et al., 2017; Tramèr et al., 2016). obtains a prediction for each query in the form of probability vectors or single classes . uses queried samples and their predictions to train , a DNN. It chooses the DNN model architecture, training hyperparameters and training process. Requiring to be a DNN is justified by the observations in prior work on model extraction attacks (Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017) that needs to have equal or larger capacity than in order for model extraction to be successful. DNNs have the greatest capacity among ML models (Zhang et al., 2017).
2. Assumptions
We assume that for a given input , has no a priori expectation regarding the prediction . treats as the ground truth label for . expects that multiple queries of the same input must return the same prediction .
Our focus is on who makes available via a prediction API since it has the greatest impact on ’s business advantage. We do not consider an who keeps for private use. This is similar to media watermarking schemes where access to allegedly stolen media is a pre-requisite for ownership demonstration (Petitcolas et al., 1999).
3. DAWN Goals and Overview
On one hand, model extraction attacks against DNNs have been proven difficult to defend against (Juuti et al., 2019). On the other hand, existing watermarking techniques (Merrer et al., 2017; Adi et al., 2018; Darvish Rouhani et al., 2019; Chen et al., ; Li et al., ) are vulnerable to model extraction attacks (Zhang et al., 2018). To address these limitations, we design a solution to identify and prove the ownership of DNN models stolen through a prediction API.
Our solution, DAWN (Dynamic Adversarial Watermarking of Neural Networks), is an additional component added in front of a model prediction API (Fig. 1). DAWN dynamically embeds a watermark in responses to queries made by a API client. This watermark is composed of inputs for which we return incorrect predictions . uses all the responses including these mislabeled samples to train . will remember those samples as a backdoor (Chen et al., 2017) that represents the watermark (as in traditional DNN watermarking techniques). If exposes a public prediction API, a judge can run a verification process (verify), which confirms is a surrogate of . Verify checks that for sufficient number of inputs , we have . DAWN embeds a watermark into a subset of queries it receives so that any trained using these responses will retain the watermark.
4. System requirements
We define the following requirements for the watermark that DAWN embeds in during an extraction attack. W1-W3 were introduced in (Adi et al., 2018) while W4 is a new requirement specific to DAWN.
Unremovability: is unable to remove the watermark from without significantly decreasing its accuracy, rendering it “unusable”. If is free of the watermark, then .
Reliability: If verify outputs “true” for a watermark on a model , then is a surrogate of , with high confidence. On the other hand, if is not a surrogate, cannot generate a watermark such that verify outputs “true” (non-trivial ownership).
Non-ownership piracy: cannot produce a watermark for a model that was already watermarked by , such that it can cast ’s ownership into doubt.
Linkability: If verify outputs “true” for a model , the watermark used for verification can be linked to a specific API client whose queries were used to train .
We identify additional requirements X1-X3:
Utility: Incorrect predictions returned by DAWN do not significantly degrade the prediction service provided to legitimate API clients: .
Indistinguishability: cannot distinguish incorrect predictions from correct victim model predictions .
Collusion resistance: Watermark unremovability (W1), linkability (W4) and indistinguishability (X2) must remain valid even if the extraction attack is distributed among several API clients.
5. Relation to other attacks
Dynamic Adversarial Watermarks
We first present the method for generating and embedding an adversarial watermark. Then we describe the process for proving ownership of a model using the watermark.
We define watermarking an input as returning an incorrect prediction instead of the correct prediction . The collection of all watermarked inputs composes the trigger set that will be a backdoor to any trained using responses from including . Consequently, inputs and their corresponding prediction classes compose the watermark to the surrogate model . We define two functions:
: should the response to be watermarked?
: what is the (backdoored watermark) response?
must not be able to predict or distinguish between and . The same query, regardless of the API client, must always get the same output. Both functions must be deterministic random functions specific to to fulfill these properties.
We use the result of a keyed cryptographic hash function as a source for randomness. We compute using SHA-256, where is a model-specific secret key generated by DAWN and is an input to . If is a matrix of dimension , it is flattened to a 1-dimensional vector. The result of the hash is split in two parts and , respectively used in and . These numbers are independent and provide a sufficient source for randomness for each function.
is a boolean function. We define as the fraction of inputs to be watermarked out of inputs submitted by an API client. will define the size of the trigger set . Then:
The expectation that returns and thus to watermark a sample is uniformly equal to . It is worth noting that DAWN does not differentiate adversaries from benign API clients. Consequently, any API client obtains a rate of incorrect predictions. must be defined to meet a trade-off. A large increases the reliability of ownership demonstration and prevents trivial ownership demonstration W2 as later discussed in Sect. 4.3. A small maximizes utility X1 by minimizing the number of incorrect predictions returned to benign API clients.
1.2. Backdoor function
does not need to permute all positions of but only those with highest probabilities for the purpose of backdooring. A large number of classes typically have a 0 probability value when is large. Considering that the number of positions to permute is small, we use the Fisher-Yates shuffle algorithm (Fisher et al., 1949) to implement . We use as the key that determines the permutations performed during the Fisher-Yates shuffle algorithm. A 128-bits key allows for list permutation of up to 34 positions (34 prediction probabilities) in a secure manner.
1.3. Indistinguishability
Outputs must be indistinguishable from X2. This requirement is partially addressed by our assumption that has no expectation regarding predictions obtained from (Sect. 3.2). Nevertheless, our watermarking function is configured by a hash of the input . A subtle modification to produces a different hash and consequently, a different result . If receives different predictions for and for a small , it can discard both and from its training set to avoid the watermark.
For each input , we obtain its latent representation based on as . This ensures that as long as ’s prediction is resilient to perceptual modifications (e.g. translation, illumination), so is . Next, we smoothen by binarizing it based on the median value of each of its features. The median of each feature value is obtained by querying using ’s training set, recording corresponding and taking the median. Using 5000 samples and their intermediate representation of length 100, we get 100 feature vectors of length 5000 and thus, 100 median values. We evaluate in Sect. 6.2.
2. Watermark embedding
uses the set of inputs and the corresponding predictions returned by DAWN-protected prediction API of to train . Approximately samples from constitute the trigger set consisting of incorrect predictions . Given that has enough capacity (large enough number of parameters), it will be able to remember a certain amount of training data having arbitrarily incorrect labels (Zhang et al., 2017). This phenomenon is called overfitting and it can be prevented using regularization (Bishop, 2006). But it is not effective for DNNs with a large capacity (Zhang et al., 2017). This is the rationale for the existence of DNN backdoors (Liu et al., 2018) and for DAWN. We expect our watermark to be embedded as a backdoor in as a natural effect of training a model with high capacity. If the watermark is not embedded, we expect ’s accuracy on the primary task to be too low to make it usable (W1).
Different adversaries will have different datasets . Consequently, the trigger sets selected by DAWN will also be different. Different surrogate models will embed distinctive watermarks. Each watermark thus links to the API client identifier. DAWN meets the linkability requirement W4.
3. Watermark verification
We present the verify function used by to prove a model is a surrogate of . Verify tests if a given watermark is embedded in a model suspected to be a surrogate of . We first define that computes the ratio of different results between the backdoor function and the suspected surrogate model for all inputs in the trigger set.
The watermark verification succeeds, i.e., verify returns “true”, if and only if , where is a tolerated error rate that must be defined. This means we must have at most samples where and differ in order to declare is a surrogate of . The choice for the value of is a trade-off between correctness and completeness for watermark verification (reliability W2). Assume we want to use a pre-generated watermark to verify if an arbitrary model is a surrogate. For simplicity, we assume a uniform probability of matching the prediction of a watermarked input , where is the number of classes of . The probability for trivial watermark verification success, given a trigger set of size and an error rate , can be computed using the cumulative binomial distribution function as follows.
This probability is the average success rate of wanting to frame for model stealing using an arbitrary watermark. Figure 2 depicts the decrease of this success rate as we increase the watermark size. We see that the verification function can accommodate a large error rate () while preventing trivial success in verification using a small watermark (). The error rate must be defined proportionally to the number of classes . Large error rates can be used for models with a large number of classes. For instance, we can set for a model with classes, limiting the adversary success rate to less than for a watermark of size 70.
The success rate in trivial verification is the complement of the confidence for reliable watermark verification, and for reliable demonstration of ownership by transition . The choice of defines the minimum watermark size given a targeted confidence. Recall that this size must also be small to ensure utility of the model to protect X1. The tolerated error must necessarily be lower than the probability of random class match: . Also, must be larger than where is the accuracy of the watermarked surrogate model on the trigger set.
The success of watermark verification is not sufficient to declare ownership of a surrogate model . can increase its success in trivial watermark verification from random using several means. For instance, knowing and , can find inputs for which and use pairs as a watermark that would successfully pass watermark verification. Thus demonstrating ownership requires a careful process to ensure that the probability for matching an incorrect prediction class remains random, ensuring that the probability for trivial watermark verification follows Eq. 3.
4. Demonstrating ownership
We present the process for a model owner to demonstrate ownership of a surrogate model watermarked by DAWN. It only requires the suspected surrogate model to expose a prediction API. This process uses a judge who is trusted to (a) ensure confidentiality of all data submitted as input to the process and (b) correctly execute and report the results of the specified verify. It also uses a time-stamped public bulletin board, e.g., a blockchain, in which information can be published to provide proof of anteriority. can be implemented using an trusted execution environment (TEE) (Ekberg et al., 2014).
publishes cryptographic commitments of the following elements in the public bulletin board:
for each API client , one registered watermark .
The commitment can be instantiated using a cryptographic hash function , e.g., SHA-3. Each watermark should be linked to the corresponding model, e.g., by associating with each registered watermark.
Several updated versions of the registered watermark can be published for each API client, as they make more queries to the prediction API and their watermarks grow. The verification of any one of these watermarks is sufficient to demonstrate ownership of the model. We define the following rules for reliable demonstration of ownership W2 that prevents ownership piracy W3:
is valid only if published later than .
can refute is a surrogate model only if has been published.
can only demonstrate that is a surrogate of if is published later than (or not published at all).
in case of contention, the model having its commitment first published is deemed to be the original.
4.2. Verification process
When suspects a model is a surrogate of trained by an API client , it provides a pointer to the prediction API of to . It also provides the following secret information using a confidential communication channel: the API client watermark and . does the following to check if is a surrogate of . If any step fails, the ownership of is not considered to have been demonstrated. If all succeed, gives the verdict that is a surrogate of .
compute and use it as a pointer to retrieve the registered watermark from the public bulletin.
compute and verify , where is extracted from the registered watermark.
retrieve from the public bulletin and verify it was published before .
query to ’s prediction API and verify that
.
input to and verify .
If ’s owner () wants to contest the verdict, it must provide the original model to using a confidential communication channel. assesses that the provided model and the API model are the same by verifying . Then, computes and retrieves it from the public bulletin. If was published before , concludes that is an original model.
Experimental setup
We evaluate DAWN using four image recognition datasets that were used in prior work to evaluate DNN extraction attacks. MNIST (LeCun et al., 2010) (60,000 train and 10,000 test samples, 10 classes) and GTSRB (Stallkamp et al., 2011) (39,209 train and 12,630 test samples, 43 classes) are respectively a handwritten-digit and traffic-sign dataset used to showcase the extraction of low capacity DNN models (Juuti et al., 2019; Papernot et al., 2017). CIFAR10 (Krizhevsky, 2009) (50,000 train and 10,000 test samples, 10 classes) and Caltech256 (Griffin et al., 2007) (23,703 train and 6,904 test samples, 256 classes) contain images depicting miscellaneous objects that were used to showcase the extraction of high capacity DNN models (Orekondy et al., 2019; Correia-Silva et al., 2018).
We also selected a random subset of 100,000 samples from ImageNet dataset (Deng et al., 2009) (1000 classes), which contains images of natural and man-made objects. We use it to evaluate the embedding of different types of watermarks and to perform a model extraction attack that requires such samples (Orekondy et al., 2019).
1.2. Models
We select two kinds of DNN models to evaluate the embedding of a watermark: low-capacity models having less than 10M parameters, and high-capacity models having over 20M parameters. These models are presented in Table 2.
In order to accurately reconstruct model extraction attacks, we use the same model architectures and training process as in (Juuti et al., 2019) for low-capacity models and as in (Orekondy et al., 2019) for high-capacity models. Similarly to prior work (Orekondy et al., 2019), we use ResNet34 (He et al., 2016) architecture pre-trained on ImageNet as a basis for high-capacity models. We fine-tuned Caltech-RN34, GTSRB-RN34 and CIFAR10-RN34 models using Caltech256, GTSRB and CIFAR10 datasets respectively We chose to reproduce only the Caltech-RN34 experiment from (Orekondy et al., 2019) because of its best performance. We used CIFAR10 and GTSRB to conduct supplementary experiments with high capacity models as they allow us to juxtapose results of experiments with low and high capacity models on the same datasets.. We also trained DenseNet121 (Huang et al., 2017) models to perform additional experiments due to the absence of dropout layers in ResNet34 models. All models were trained using Adam optimizer with learning rate of 0.001 that was decreased over time to 0.0005 (after 100 epochs for ResNet34 models and half-way for the other), except for Caltech-RN34. For Caltech-RN34, we used SGD optimizer with an initial learning rate of 0.1 that was decreased by a factor of 10 every 60 epochs over 250 epochs. We used a batch size of 16 for fine-tuning ResNet34 and DenseNet121 based models.
2. Watermarking Procedure
Inputs from ’s dataset are submitted to the DAWN-enhanced prediction API of which returns correct or incorrect predictions according to the result of the watermarking function . For the experiments in Section 6.2 (evaluating the effectiveness of the mapping function ), we use the embedding from as . Experiments in Section 6.1 and Section 7 do not depend on the choice of . Therefore, for the sake of simplicity, we use the identity function as in these experiments.
We simulate who uses the whole set , which includes samples with incorrect labels, to train its surrogate model . trains without being aware of the watermarked samples in .
3. Evaluation Metrics
We use two metrics to evaluate the success of ’s goal and ’s goal respectively. ’s goal is to train a surrogate model that has maximum accuracy on ’s primary classification task. We evaluate this by computing the test accuracy of the surrogate model on the test set of each dataset.
’s goal is to maximize the embedding of the watermark in any surrogate model built from responses from such that its surrogacy can be reliably demonstrated. We evaluate this by computing the watermark accuracy of the surrogate model on the trigger set of watermarked inputs.
DAWN aims to maximize regardless of . aims to maximize while minimizing . In our experiments, we calculate both metrics every 5 epochs in order to evaluate their progress during the training process.
Robustness of watermarking
We assess ’s ability to prevent the embedding of a watermark in a surrogate model, i.e., to violate the unremovability requirement W1. Prior work evaluated unremovability after a watermarked model is trained showing that backoor-based watermarks are resilient to model pruning and adversarial fine tuning (Adi et al., 2018; Merrer et al., 2017; Zhang et al., 2018). DAWN also embeds backdoor-based watermarks resilient to removal using post-training manipulations. Thus, we focus on adversarial manipulations during training by evaluating several solutions that could prevent watermark embedding. We then evaluate the ability for to identify watermarked inputs using the trained surrogate model, i.e., to violate the indistinguishability requirement X2.
We take an ideal model extraction attack scenario where is a perfect oracle. has access to a large dataset of natural samples from the same distribution as training data: we use the whole training set from each dataset (Sect. 5.1) for . We use a large watermark of fixed size in all following experiments. Embedding a large watermark is challenging since the model must learn many isolated errors (mislabeled inputs). We take as an upper bound to the watermark size and a worst case scenario for DAWN watermark embedding.
We evaluate the impact of two parameters on embedding a watermark during DNN training. The first parameter is the capacity of . can limit this capacity such that the model could only learn the primary classification task and cannot learn the watermark. The second parameter is the use of regularization. Regularization accommodates classification errors on the training data, which is considered as noise. The watermark consists of incorrectly labeled inputs which can potentially be discarded using regularization.
We evaluate the impact of model capacity and regularization on watermark accuracy and test accuracy of . We trained several surrogate models having low and high capacity. was randomly selected from the respective training sets. We used plain training and two regularization methods, namely weight decay (Krogh and Hertz, 1992) with decaying factor and dropout (DO=X) (Srivastava et al., 2014) with probability X=. We selected values optimal for : such that they maximize the difference .
Table 3(a) and 3(b) present the results of this experiment for DNN models with low and high capacity respectively. We report and results at three training stages providing (1) best watermark accuracy (best for ), (2) best test accuracy (best for ) and (3) when training is completed. Overall, we observe that and are high for most settings. Using plain training, is mostly higher than and often close to 100%. The ownership of all these surrogate models can be reliably demonstrated using a low tolerated error rate, e.g., .
Model capacity. High-capacity models can provide higher watermark and test accuracy than low-capacity models as highlighted by comparing results for GTSRB and CIFAR10 in both tables. While is low for some low-capacity models, e.g., MNIST-3L, MNIST-5L (DO), their test accuracy is similarly low and close to random . This shows that reducing the model capacity can prevent the embedding of the watermark. However, decreasing to a level where it cannot be used to reliably prove ownership makes unusable. and are closely tied when manipulating the model capacity and thus this is not a useful strategy to circumvent DAWN.
Regularization. Regularization is useful for decreasing the watermark accuracy in a few cases. Weight decay is useful for low-capacity GTSRB-5L and CIFAR10-9L models. Dropout is useful for low-capacity MNIST-5L and CIFAR10-9L models, and for high-capacity Caltech-DN121 model. Dropout completely prevents the embedding of the watermark into MNIST-5L model as depicted by . However, is also significantly reduced, by 50% at best, making potentially unusable. In all remaining cases, is reduced down to 20-35%, while preserving high test accuracy similar to models trained with non-watermarked datasets. While is low, the watermark can still successfully demonstrate ownership by increasing the tolerated error rate to, e.g., . Considering the large watermark size of 250, this demonstration would still be reliable despite the high tolerated error rate as evaluated in Sect. 4.3.
It is worth noting that no regularization method is effective at removing the watermark from high capacity GTSRB-RN34 and CIFAR10-RN34 models. The likely reason is that ResNet34 architecture has significant overcapacity for the primary task of classifying these datasets. Regularization cannot limit this capacity to an extent where the watermark would not be embedded. This means needs sufficient knowledge of to select an appropriate model architecture for . It must have sufficient capacity to learn the primary classification task of the victim model while preventing watermark embedding. In model extraction attacks, has black-box access to , which forces to use with sufficient capacity to maximize the attack success (Orekondy et al., 2019). In this setting, regularization is not useful to circumvent DAWN.
Finally, while regularization can be useful, needs relevant test data and ground truth to optimize the regularization parameters (e.g., decaying factor ). In all extraction attacks (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018; Pal et al., 2020) the availability of relevant data is the main limitation. All this data is typically used for training the surrogate model and none is used for test purposes, which prevents optimization of regularization parameters and early stopping.
2. Mapping Function
can try to identify watermarked inputs and remove them from prior to training in order to prevent watermark embedding. Because DAWN relies on a hash to decide if an input is watermarked, can query multiple perturbed versions of inputs in and discard those that return different predictions. The mapping function presented in Sect. 4.1.3 is meant to prevent this evasion. We evaluate the effectiveness of by querying 10 perturbed versions of each of the 10,000 samples in , which includes . For each query, we check whether they get consistent mapping and classification . Table 4 reports the results of this experiment for various perturbation size for the MNIST dataset. We distinguish cases where 1 (same ) or (diff ); 2 (same ) or (diff ). Same and same means keeps a watermarked sample in ( succeeds). Same and different means discards a watermarked sample from ( fails). Different means wrongfully discards a sample from regardless of ( is too large and changes ’s prediction). We see succeeds to provide a consistent mapping in over 85% cases for , meaning 85% of the is preserved in . As perturbations increase in size, returns an increasing rate of inconsistent mapping, but this rate is similar to the one of changed predictions from . Thus, we conclude is resilient to perturbations and DAWN can effectively watermark .
Protecting against model extraction attacks
We evaluate DAWN’s effectiveness at watermarking surrogate DNN models constructed using two model extraction attacks: 1) the PRADA attack (Juuti et al., 2019) achieves state-of-the-art performance in extracting low-capacity DNN models primarily using synthetic data and we launch it against MNIST-5L, GTSRB-5L and CIFAR10-9L; 2) the KnockOff attack (Orekondy et al., 2019) extracts high-capacity DNN models using only natural data and we launch it against GTSRB-RN34, CIFAR10-RN34 and Caltech-RN34. The test accuracy of each extracted with these respective attacks is reported in Tab. 6.
We demonstrate how to setup DAWN to protect a given victim model . We evaluate the successful embedding of watermarks in several surrogate models as well as their utility considering a circumvention strategy.
Watermarking decision: DAWN degrades utility by a factor equal to due to incorrect predictions for watermarked inputs. The value of is specific to . Given a desired level of confidence for reliable ownership demonstration equal to (cf. Eq. 3), a tolerated error rate and the number of classes for , we can compute the minimum size for the watermark using Eq. 3. Given that can estimate the minimum number of queries required by to train a usable surrogate model for , we can compute . This ratio ensures that if can successfully train a usable surrogate model , then will embed a watermark large enough to reliably demonstrate its ownership .
The probability for successful trivial watermark verification is valid for testing a single watermark. This probability increases by a factor equal to the number of tested watermarks. DAWN creates and registers client-specific watermarks. must estimate the number of API clients to calculate the actual probability for trivial demonstration of ownership considering that all registered watermarks should be tested. When verifying a watermark, the judge counts the number of registered watermarks for in the public bulletin. computes the real probability for successful trivial watermark verification accordingly and decides if a demonstration of ownership is reliable or not according to this final confidence.
Utility for legitimate clients: Suppose we want a confidence for reliable demonstration of ownership equal to . has a prediction API with 1M API clients (1M watermarks are registered for ). We need to be able to test all registered watermarks while achieving our targeted confidence. We choose a tolerated error rate . Table 5 reports the computed watermark ratio required to protect six models against model extraction. We see must always be lower than 0.5% to reach confidence for any victim model. ’s accuracy is thus degraded in a negligible manner that does not impact its utility. DAWN meets the reliability W2 and utility X1 requirements.
Overhead: Storing 1M watermarks would require at most a few TBs (cf. Tab. 5). Watermark verification consists in obtaining predictions from a purported surrogate model. It is operated by who gets predictions at no monetary cost. Thus, demonstration of ownership is only a matter of time and getting one prediction from our most complex model (Caltech-RN34) takes 9ms (on Tesla P100 GPU). Verifying one watermark for this model takes 0.25s (27 queries) and verifying 100,000 watermarks takes 7 hours using a single GPU. can initially verify all watermarks with a lower confidence to reduce this time (by testing only a subset of each watermark). Only successful verification would later undergo a verification of the full watermark. Testing the same 100,000 watermarks with targeted confidence (instead of ) requires 1h15 (5 samples per watermark). This time can further be reduced by parallelizing predictions on several GPUs. DAWN’s verification process is more computationally expensive due to the requirement of testing all watermarks to account for Sybils. However, unlike prior watermarking schemes, DAWN is effective against model extraction attacks.
2. Effectiveness against real extraction attacks
We want to show that any surrogate of a victim model protected by DAWN will embed a watermark that allows for reliable demonstration of ownership. We evaluate the effectiveness of DAWN against two landmark model extraction attacks namely PRADA (Juuti et al., 2019) and KnockOff (Orekondy et al., 2019).
Low-capacity models expose a prediction API that returns prediction classes required for the PRADA attack. High-capacity models return the full probability vector . Each victim model is protected by DAWN using the setting presented in Sect. 7.1. This setting enables to demonstrate ownership of each surrogate model with confidence using a tolerated error rate . For demonstration of ownership to be successful, the surrogate model must pass the watermark verification test . In our setting, it means that DAWN successfully defends against an extraction attack if the watermark accuracy for is larger than 50%, i.e., .
Table 6 presents the result of this experiment. We see all surrogate models have a watermark accuracy , which means is successful in demonstrating their ownership. DAWN successfully defends against the PRADA and KnockOff attacks for all tested models while incurring little decrease in ’s utility (evaluated in Sect. 7.1). We have shown DAWN effectively embeds a watermark in surrogate models stolen using extraction attacks. In Table 6, note that DAWN significantly decreases the surrogate model test accuracy ( for ) for MNIST-5L while it has little impact on the same for other datasets. Drastic reduction in is not a concern from the defender’s perspective - in fact it can, by itself, serve as a deterrence for against model extraction. In all cases, adequate watermark accuracy serves as a deterrence.
3. Resilience to distributed extraction attack
Distributing a model extraction attack across several API clients means several adversaries query a subset from the whole set used to train the surrogate model . Recall that DAWN is a deterministic mechanism Sect. 4.1. The watermarking and backdoor functions are deterministic and specific to . Their results only depend on the input queried to . The responses to , and its corresponding trigger set, remain the same regardless of which client(s) query the prediction API. Thus, is labeled in the same manner and it includes the same trigger set whether it is queried by one or by multiple API clients. Thus, the watermark in trained using will remain indistinguishable X2 and unremovable W1 even if multiple clients collude.
Note that in the case of colluding clients, each adversary has a subset of the whole trigger set . When verifying ownership, the judge will have several successful watermark verifications : one for each adversary who colluded to build the surrogate model . The verification of each sub-watermark has the same expectation for success as the verification of the whole watermark . will conclude that each API client whose watermark is successfully verified is a perpetrator of the distributed extraction attack used to build the surrogate model . Linkability W4 remains valid in case of collusion.
In a distributed attack, the watermark associated to each colluding client is smaller than in a centralized attack. To verify ownership with a same reliability, we must increase the watermark size and consequently by a factor equal to the number of colluding clients. We assume the number of real colluding clients is limited, e.g., a few tens. Nevertheless, it is possible to mount a Sybil attack in which several API accounts are created by a single adversary. The API account registration process must require providing information that maximizes difficulty of creating trusted accounts, e.g., verified phone number or credit card, to mitigate this threat. Also, Sybils-detection techniques exist (Tran et al., 2009; Wang et al., 2013) and it is possible to link Sybils accounts by examining querying patterns and IP addresses for instance (Stringhini et al., 2015). For example, to protect Caltech-RN34, we could increase to reliably verify the watermark of 35 colluders while maintaining the utility loss below 1%. Consequently, the higher the number of classes, the greater the reliability of watermark verification (c.f. Eq. 3) and we can tolerate more Sybils. For a classifier with 10,000 classes and utility loss below 1% we can reliably verify the watermark of 87 colluders.
Watermark Removal
Several techniques can identify if a DNN model has a backdoor (Chen et al., 2019; Guo et al., 2019; Wang et al., 2019). Most techniques like Neural Cleanse (Wang et al., 2019) and TABOR (Guo et al., 2019) can only detect backdoors for which the trigger is a static pattern added to original inputs (e.g., yellow square added to an image). In contrast, our trigger set is composed of unmodified samples having only incorrect labels. Consequently, techniques like Neural Cleanse and TABOR are ineffective at detecting DAWN watermark. In this section, we evaluate the resilience of DAWN watermarks to removal using six attacks: (1) double-extraction of a second order surrogate model , (2) fine-tuning (Kornblith et al., 2018), (3) pruning (Blalock et al., 2020), (4) training with noise, (5) adding noise during the inference, and (6) recognizing queries from training data.
Double extraction and fine-tuning. A watermark may be removed by performing an extraction attack against to obtain a second order surrogate model . has full control over : its prediction API is not protected by DAWN and does not intentionally return incorrect prediction. If uses a disjoint set of queries to extract a surrogate from , then may not embed the watermark, preventing the demonstration of its ownership by . Instead of starting the second extraction from scratch, it can use as the starting point for and fine-tune it. We call this stealing+fine-tuning.
We observed in Tab. 6 that surrogate models have a lower accuracy than victim models because model extraction incurs a necessary decrease in surrogate model accuracy. We evaluate the extent of the decrease in and if launches two successive extraction attacks instead of one or steals+fine-tunes the model to obtain : the first against and the second against . We evaluate these evasion techniques using the PRADA and KnockOff attacks.
The two successive extraction attacks and stealing+fine-tuning are performed in the same conditions as the first attack. The only difference is that uses half the seed samples for each PRADA attack (5 per class for MNIST-5L and GTSRB-5L, 500 per class for CIFAR10-9L) and runs an additional duplication round to query the same number of inputs. The number of seed samples is a limited adversarial capability in PRADA (Juuti et al., 2019), so we grant with the same capability for single and double extraction attack. For each KnockOff attack, uses a different set of 100,000 inputs from ImageNet. For the second extraction attack (against ), we query the same number of inputs as for the first one. While this number can be increased, we empirically observed that the test accuracy of reaches its maximum and stagnates before the PRADA and KnockOff attacks finish, i.e., more queries do not improve of .
As it can be observed in Tab. 7 and 8, double extraction attack and stealing+fine-tuning effectively remove the watermark from the second order surrogate model . The watermark accuracy is low enough (3-22%) to fail demonstration of ownership for , which empirically confirms that prior DNN watermarking techniques are not resilient to this class of model extraction attacks (Zhang et al., 2018).
While removing the watermark, the extraction of the second order surrogate model using these attacks also increases the degradation in test accuracy by 20% to 80% for compared to . Considering a powerful having unlimited access to natural data (e.g., KnockOff adversary model) the extraction of the first order surrogate model incurs little accuracy degradation and so does the extraction of the second order surrogate model. The final model stolen using KnockOff attack has its watermark removed and preserves its utility (from -1pp to -10pp compared to ). DAWN cannot protect model extraction attacks where has unlimited access to natural data. However, ’s access to data is limited in many scenarios, e.g., access to medical imaging that are privacy sensitive, and highly specialized models may not return meaningful predictions to random images. In such a scenario, the KnockOff attack may not be effective and the PRADA attack is more effective. We see that for the PRADA attack the test accuracy of decreases sharply during each attack that removes the watermark (3 top rows in Tab. 7 and 8) . In most cases, the final test accuracy of is less than half of (from -35pp to -62pp) and we consider that makes too inaccurate to be useful. DAWN can effectively protect a model against extraction attack that uses a limited amount of data - it destroys the utility of the model deprived of the watermark.
Double-extraction and fine-tuning can remove the watermark while preserving test accuracy given that has unlimited access to natural data. However, DAWN defeats both attacks when has limited access to data, which is the case for most model extraction attack scenarios (cf. Sect. 2.1). ’s access to data is limited in many scenarios, e.g., medical imaging classifiers, where these removal attacks would be ineffective.
Pruning. Alternatively, may attempt to prune the model by setting random weights of the model to zero. In our experiments we prune weights uniformly randomly. We show that for large values of pruning is effective at removing the watermark but it sacrifices model’s utility and renders it useless (c.f. Table 9). Furthermore, we observe that and do not fall proportionally i.e. there is no guarantee that sacrificing X% results in the same drop in . Also, our experiments show that determining appropriate without knowing is challenging as there is no consistent drop in accuracy for a particular value of across all models.
Training with noise. may attempt to weaken the embedding of the watermark by adding some noise to the samples before they start training. ’s goal is not to expose the model to the samples that would be eventually used for the verification. We show that for the values up to , is not affected (c.f. Table 10). Beyond that in all but one case, either remains above the effectiveness threshold, or the drop in becomes unacceptable. This approach is thus not effective at successfully stealing the model while circumventing DAWN.
Inference with noise. Instead of training with noisy samples, can add noise to all samples during the inference in attempt to avoid verification. However, this will reduce utility for ’s clients. In Tab. 11, we show the decrease in and corresponding for various amounts of noise . We show that in almost all cases remains high or drops below the acceptable utility level (purple, underline). In few cases (red, dashed underline) remains high while decreases to . However, the value that benefits the most is not consistent across the models. Hence, for a particular value of that preserves , has no guarantee that watermark verification would fail.
Recognizing queries from training data. Alternatively, having deployed and being aware of DAWN can try to prevent watermark verification performed by . can check if queries to belong to used to steal , and return different predictions for them. We evaluated that searching for exact matches can be done efficiently using a hash table (28-44ms additional overhead per query for our datasets). Hashing the query is the most time consuming part of this search.
To prevent this evasion by , can slightly perturb samples in before submitting them for verification to . In Tab. 11, we show the resilience of watermark verification to various amounts of noise () added to each image in . We show that we can maintain for all models up to or even more for some models. can no longer perform a simple lookup in a hash table to identify queries from if is perturbed. must compute the distance to every sample in and find the nearest neighbor to the query. We demonstrate that such search incurs a substantial computational overhead (c.f. Tab. 12). Searching in 100,000 ImageNet samples can take over 7s per query (on a server-grade machine with a Xeon CPU and 64 GB of RAM), which is too long to be acceptable for deployment. Also, a small distance to an element in does not mean that queried image was in fact part of . has to set up a threshold for rejecting the queries that will affect the utility of the model, on top of the described computational overhead.
Discussion
Unremovability W1. We extensively evaluated (Sect 6.1) that manipulation of the training process of either does not prevent the embedding of the watermark or if it does, it significantly degrades ’s test accuracy. Proper use of regularization can effectively mitigate the watermark embedding but it requires to be granted more capabilities e.g. increased access to relevant data. We also showed (Sect. 8) that two successive extraction attacks can remove a watermark from . However, it also decreases the test accuracy to an extent that makes unusable. Finally, prior work (Adi et al., 2018; Merrer et al., 2017; Zhang et al., 2018) has shown that manipulations after training such as pruning and adversarial fine-tuning are ineffective against backdoor-based DNN watermarks. Our evaluation confirmed that watermark cannot be removed using pruning or fine-tuning (Sect. 8) without sacrificing the utility of the model. We can conclude that DAWN watermarking meets unremovability requirement.
Indistinguishability X2. We defined model-specific watermarking and backdoor functions (Sect. 4.1) that always return the same same output (correct or incorrect) for the same input. We also introduced a solution for mapping inputs with minor differences to similar predictions (Sect 4.1.3).
Reliability W2 and utility X1. The watermark registration and verification protocol that we introduce (Sect. 4.4) ensures that the success of in demonstrating ownership of an arbitrary model is negligible. We showed how to set up DAWN in order to reliably demonstrate ownership of several surrogate models stolen using two state-of-the-art model extraction attacks with high confidence equal to (Sect. 7). DAWN effectively watermarked every surrogate model while causing a negligible decrease of ’s utility (0.03-0.5%). DAWN allows for reliable ownership demonstration while preserving the utility.
Non-ownership piracy W3 is guaranteed by our watermark registration and verification protocol (Sect. 4.4). In case of contention, the first registered model is deemed the original.
Linkablity W4. DAWN selects watermarked inputs from API client queries and registers one watermark per API client. Different clients make different queries and they will consequently have different watermarks. Given that we meet the reliability requirement W2, a single watermark will succeed in proving is a surrogate of . Watermarks are API client-specific which makes a surrogate model linkable to an API client.
Collusion resistance X3. DAWN relies on deterministic functions for watermarking () and backdooring () that are specific to but independent of the client sending a query. Thus, the watermark remains indistinguishable despite collusion X2. DAWN is resilient to a Sybil attack (bounded to a certain number of Sybils) and assure successful verification by (Sect. 7.3).
2. Limitations
can attempt to prevent ownership demonstration for by ensuring that watermark verification (Eq. 2) fails. This entails reducing the watermark accuracy by training using only a subset of the trigger set . Since watermarked inputs are indistinguishable (X2) and uniformly distributed in , cannot selectively discard them. Nevertheless, can discard of the whole , which statistically, will result in of watermarked input being discarded. This should reduce by . If is high enough, the resulting can be brought down low enough for watermark verification to systematically fail.
While this strategy is effective, it deprives from a large part of . This decreases and consequently the utility of (Sec. 8). Alternatively, must collect a set larger and make more queries to to compensate for later discarded training inputs. We already discussed in Sect 2.1 that access to relevant data is the main limitation for . The secondary goal of is to limit the number of queries to (cf. Sect. 3.1). This evasion strategy requires more adversarial capabilities (access to data) and it compromises one adversary goal (minimum number of queries). Consequently, even if effective, we do not consider it a realistic evasion strategy.
Another potential limitation of DAWN is circumvention of the mapping function (Sec. 4.1.3). If mapping is too aggressive, may probe the input space and try to identify subspaces that are grouped together. However, this is not guaranteed to work because the behavior of the model on synthetic samples is undefined (Goodfellow et al., 2014). impacts only the watermarking decision and not the returned label - cannot interact directly with the mapping function. On the other hand, if the tolerated modification is too small, might identify watermarked queries by submitting several samples with minor modifications and taking the majority vote of the label. We evaluated this attack in Sect. 6.2 showing that mappings from are as consistent as predictions from . Semantic-preserving modifications to image queries (e.g., translation, rotation, change in color intensity, etc,) could be used to improve this attack. However, by using an embedding from to implement , both functions should be as resilient to semantic-preserving modifications.
can attempt to weaken the embedding of the watermark by adding a small amount of noise to its training samples before starting the training or to all queries during the inference time (Sect. 8). Although cannot know the optimal value that minimizes accuracy loss while rendering watermark verification ineffective, they can choose a loss budget and incur that loss fully (e.g. 10 pp in our examples) - that implies that in Table 11 would have succeeded in two out of the six cases. How to strengthen DAWN against such an adversary that is ready to incur the maximal allowable accuracy loss is still an open problem.
Related Work
Watermarking DNN models. The first watermarking technique for DNNs (Uchida et al., 2017) explicitly embeds additional information into the weights of a DNN after it is trained. Verifying the watermark requires white-box access to the model in order to analyze the weights. A limitation of this approach is that the watermark can be easily removed by minimally retraining the watermarked model.
Alternative approaches (Merrer et al., 2017; Adi et al., 2018; Zhang et al., 2018; Darvish Rouhani et al., 2019; Jia et al., 2020) that are more robust have been proposed, where the watermark can be verified in a black-box setting. These are based on backdooring and they allow for watermark extraction using only a prediction API, as discussed in Sect. 2.2. These approaches use both a carefully selected trigger set and a specific training process chosen by the model owner. The first proposal for such approach (Merrer et al., 2017) consist in modifying the original model boundary using adversarial retraining (Madry et al., 2017) in order to make the model unique. The watermark is composed of synthetically generated adversarial samples (Goodfellow et al., 2014) that are close to the decision boundary. The impact of selecting a particular distribution for a watermark has been evaluated in (Zhang et al., 2018). It shows that selecting a trigger set from the same distribution as the training data (albeit with minor synthetic modifications) or from a different distribution, does not affect the accuracy of the model for its primary classification task or on its training time, while the watermark gets perfectly embedded. Finally, more formal foundations and theoretical guarantees for backdoor-based DNN watermarking have been provided in (Adi et al., 2018). This work empirically assesses that the removability of a DNN watermark is highly dependent on the training process of the watermarked model (training from scratch vs. re-training).
Defenses against model extraction. It was suggested that the distribution of queries made during an extraction attack is different from benign queries (Juuti et al., 2019). Hence, model extraction can be detected using density estimation methods, namely by assessing the ability for queries to fit a Gaussian distribution or not. However, this technique protects only against attacks using synthetic queries and is not effective against, e.g., the KnockOff attack. Other detection methods analyse subsequent queries close to the classes’ decision boundaries (Quiring et al., 2018; Zheng et al., 2019) or queries exploring abnormally large region of the input space (Kesarwani et al., 2018). Both methods are effective but detect only extraction attacks against decision trees. They are ineffective against complex models like DNNs. Altering predictions returned to API clients can mitigate model extraction attacks. Predictions can be restricted to classes (Tramèr et al., 2016) or adversarially modified to degrade the performance of the surrogate model (Lee et al., 2018; Orekondy et al., 2020). However, some extraction attacks (Juuti et al., 2019) circumvent such defenses because they remain effective using just prediction classes.
Prior defenses to model extraction are designed to protect only simple models (Quiring et al., 2018; Kesarwani et al., 2018) or to prevent only specific extraction attacks (Lee et al., 2018; Zheng et al., 2019). It is arguable if a generic defense would ever be effective at detecting/preventing model extraction. Consequently, with DAWN we take a different approach where we assume a surrogate model can be extracted. Then we propose a generic defense to identify surrogate DNN models that have been extracted from any victim model using any extraction attack.
References
Appendix A Datasets and Models
Table 13 presents the characteristics of the datasets we used in our experiments. These are divided into a training and a testing set. Images were resized to fit the corresponding model architectures used in prior work. Table 14 presents model architectures used for conducting experiments with low capacity models - the perfect-knowledge attacker in Sect. 6.1 and reproduction of the PRADA (Juuti et al., 2019) attack in Sect. 7.2.
Appendix B Detecting watermarked inputs
We assess if watermarked inputs can be identified such that could remove them from the before training the surrogate model.
This defense consists in first training a DNN model with the whole training dataset. Then, training data is predicted using the trained model and we record the activations of the last hidden layer of the DNN model. These activations are projected to three dimensions using Independent Component Analysis (ICA) and clustered into two clusters using k-means. These clusters are expected to group benign training data and poisoned data (watermarked inputs) respectively. The intuition for this approach is that incorrectly labeled inputs (watermark) trigger different activations than correctly labeled inputs in the trained DNN model. The size and silhouette score (Rousseeuw, 1987) of the two clusters are analyzed to conclude (1) if there is backdoor in the model and (2) which training inputs compose the backdoor. According to authors, a low silhouette score (0.1/0.15) and a high difference in relative cluster size is expected if the model embeds a watermark. The smallest cluster should contain the watermarked inputs.
To evaluate this defense against DAWN, we trained two sets of DNN models, plain models using a correctly labeled training set only, and watermarked models each embedding a watermark of size . We applied the watermark detection process discussed above on these models and report results in Tab. 15. Clustering is not able to isolate watermarked inputs into a single cluster; the main part of watermarked inputs belongs to large clusters. The recommendation (Chen et al., 2019) to discard small clusters from training would deprive from a large number of correctly labeled samples while a large part of the watermark would be preserved. Using this approach, the defense wrongly discards 26.4% of clean data from on average while detecting only 67 out of 250 watermarked samples (26.8% of ). The silhouette score is not useful for detecting the watermark either since watermarked and plain models have close scores that are all above the recommended detection threshold (0.1/0.15) (Chen et al., 2019). Our plain models are detected as embedding a watermark using this defense. We conclude that this defense is ineffective at detecting watermarks generated by DAWN.
We assume the reason for this ineffectiveness is due to the nature of our watermark, selected from the same distribution as the training set. In contrast to prior DNN watermarking solutions (Merrer et al., 2017; Adi et al., 2018; Zhang et al., 2018), our watermarked inputs do not come from a single manifold distant from the training data manifold. Consequently the model does not learn a “single” activation that generalizes to the whole watermark but rather learns individual exceptions for each watermarked input. The activations of watermarked inputs are thus different from each other and they are scattered among the activations of the remaining training data (correctly labeled). This can be observed in Fig. 3 - left, where we see that watermarked inputs are scattered among correctly labeled inputs. This explains why generated clusters cannot isolate watermarked inputs from correctly labeled inputs (Fig. 3 - right).