DAWN: Dynamic Adversarial Watermarking of Neural Networks

Sebastian Szyller, Buse Gul Atli, Samuel Marchal, N. Asokan

Introduction

Recent progress in machine learning (ML) has led to a dramatic surge in the use of ML models for a wide variety of applications. Major enterprises like Google, Apple, and Facebook have already deployed ML models in their products (TechWorld, 2018). ML-related businesses are expected to generate trillions of dollars in revenue in the near future (Forbes, 2019). The process of collecting training data and training ML models is the basis of the business advantage of model owners. Hence, protecting the intellectual property (IP) embodied in ML models is necessary.

One approach for IP protection of ML models is watermarking. Recent work (Merrer et al., 2017; Adi et al., 2018; Zhang et al., 2018) has shown how digital watermarks can be embedded into deep neural network models (DNNs) during training. Watermarks consist of a set of inputs, the trigger set, with incorrectly assigned labels. A legitimate model owner can use the trigger set, along with a large training set with correct labels, to train a watermarked model and distribute it to his customers. If he later encounters a model he suspects to be a copy of his own, he can demonstrate ownership by using the trigger set as inputs to the suspected model. These watermarking schemes allow legitimate model owners to detect theft or misuse of their models.

Instead of distributing ML models to customers, an increasingly popular alternative business paradigm is to allow customers to use models via prediction APIs. But one can mount a model extraction (Tramèr et al., 2016) attack via such APIs by sending a sequence of API queries with different inputs and using the resulting predictions to train a surrogate model with similar functionality as the queried model. Model extraction attacks are effective even against complex DNN models (Juuti et al., 2019; Orekondy et al., 2019), and are difficult to prevent (Juuti et al., 2019). Existing watermarking techniques, which rely on model owners to embed watermarks during training, are ineffective against model extraction since it is the adversary who trains the surrogate model.

In this paper we introduce DAWN (Dynamic Adversarial Watermarking of Neural Networks), a new watermarking approach intended to deter IP theft via model extraction. DAWN is designed to be deployed within the prediction API of a model. It dynamically watermarks a tiny fraction of queries from a client by changing the prediction responses for them. The watermarked queries serve as the trigger set if an adversarial client trains a surrogate model using the responses to its queries. The model owner can use the trigger set to demonstrate IP ownership of the extracted surrogate model as in prior DNN watermarking solutions (Adi et al., 2018; Zhang et al., 2018). DAWN differs from them in that it is the adversary (model thief), rather than the defender (original owner) who trains the watermarked model. This raises two new challenges: (1) defenders must choose trigger sets from among queries sent by clients and cannot choose optimal trigger sets from the whole input space; (2) adversaries can select the training data or manipulate the training process to resist the embedding of watermarks. DAWN addresses both these challenges.

DAWN watermarks are client-specific: DAWN not only infers whether a given model is a surrogate but, in case of model extraction, also identifies the client whose queries were used to train the surrogate. DAWN is parametrized so that changed predictions needed for watermarking are sufficiently rare as to not degrade the utility of the original model for legitimate API clients.

present DAWN, the first approach for dynamic, selective watermarking for DNN models at their prediction APIs for deterring IP theft via model extraction (Sect. 4),

empirically assess it (Sect. 5) using several DNN models and datasets showing that DAWN is robust to adversarial manipulations and resilient to evasion (Sect. 6 and 8), and

show that DAWN is resistant to two state-of-the-art extraction attacks, reliably demonstrating ownership (with confidence >1−2−641-2^{-64}) with negligible impact on model utility (0.03-0.5% decrease in accuracy) (Sect. 7).

Code to reproduce our experiments is available on GitHub github.com/ssg-research/dawn-dynamic-adversarial-watermarking-of-neural-networks .

Background

In model extraction (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018; Pal et al., 2020), an adversary A\mathcal{A} wants to “steal” a DNN model FVF_{\mathcal{V}} of a victim V\mathcal{V} by making a series of prediction requests UU to FVF_{\mathcal{V}} and obtaining predictions FV(U)F_{\mathcal{V}}(U). UU and FV(U)F_{\mathcal{V}}(U) are used by A\mathcal{A} to train a surrogate model FAF_{\mathcal{A}}. A\mathcal{A}’s goal is to have Acc(FA)Acc(F_{\mathcal{A}}) as close as possible to Acc(FV)Acc(F_{\mathcal{V}}). All model extraction attacks (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018) operate in a black-box setting: A\mathcal{A} has access to a prediction API, A\mathcal{A} uses the set <U,FV(U)><U,F_{\mathcal{V}}(U)> to iteratively refine the accuracy of FAF_{\mathcal{A}}. Depending on the adversary model, A\mathcal{A}’s capabilities can be divided into three categories: model knowledge, data access, and querying strategy.

Model knowledge. A\mathcal{A} does not know the exact architecture of FVF_{\mathcal{V}} or the hyperparameters or the training process. However, given the purpose of the API (e.g., image recognition) and expected complexity of the task, A\mathcal{A} may attempt to guess the architecture of the model (Papernot et al., 2017; Juuti et al., 2019). On the other hand, if FVF_{\mathcal{V}} is complex, A\mathcal{A} can use a publicly available, high capacity model pre-trained with a very large benchmark datasets (Orekondy et al., 2019). While the above methods focus on DNNs, there are alternatives targeting simpler models: logistic regression, decision trees, shallow neural networks (Tramèr et al., 2016).

Data access. A\mathcal{A}’s main limitation is the lack of access to natural data that comes from the same distribution as the data used to train FVF_{\mathcal{V}}. A\mathcal{A} may use data that comes from the same domain as V\mathcal{V}’s training data but from a different distribution (Correia-Silva et al., 2018). If A\mathcal{A} does not exactly know the distribution or the domain, it may use widely available natural data (Orekondy et al., 2019; Pal et al., 2020) to mount the attack. Alternatively, it may use only synthetic samples (Tramèr et al., 2016) or a mix of a small number of natural samples augmented by synthetic samples (Papernot et al., 2017; Juuti et al., 2019).

Querying strategy. All model stealing attacks (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018; Pal et al., 2020) consist of alternating phases of A\mathcal{A} querying FVF_{\mathcal{V}}, followed by training the surrogate model FAF_{\mathcal{A}} using the obtained predictions. A\mathcal{A} queries FVF_{\mathcal{V}} with all its data and then trains the surrogate model (Correia-Silva et al., 2018; Orekondy et al., 2019). Alternatively, if A\mathcal{A} relies primarily on synthetic data (Papernot et al., 2017; Juuti et al., 2019), it deliberately crafts inputs that would help it train FAF_{\mathcal{A}}.

2. Watermarking DNN models

Digital watermarking is a technique used to covertly embed a marker, the watermark, in an object (image, audio, etc.) which can be used to demonstrate ownership of the object. Watermarking of DNN models leverages the massive overcapacity of DNNs and their ability to fit data with arbitrary labels (Zhang et al., 2017). DNNs have a large number of parameters, many of which have little significance for their primary classification task. These parameters can be used to carry additional information beyond what is required for its primary classification task. This property is exploited by backdooring attacks, which consist in training a DNN model that deliberately outputs incorrect predictions for some selected inputs (Chen et al., 2017; Gu et al., 2017).

The trigger set TT and the outputs of the backdoor function for its elements B^(T)\hat{B}(T) compose the watermark: (T,B^(T))(T,\hat{B}(T)). Let F′F^{\prime} be a DNN model that copies FF. The watermark can be used to demonstrate ownership of F′F^{\prime}. It only requires F′F^{\prime} to expose a prediction API which can be used to query all samples in the trigger set x∈Tx\in T. A sufficient number of predictions F^′(x)\hat{F}^{\prime}(x) such that F^′(x)=B^(x)\hat{F}^{\prime}(x)=\hat{B}(x) demonstrates that F′F^{\prime} is a copy of the watermarked model FF.

Problem Statement

The adversary A\mathcal{A} mounts a model extraction attack against a victim model FVF_{\mathcal{V}} using queries to its prediction API. A\mathcal{A}’s goal is model functionality stealing (Orekondy et al., 2019): train a surrogate model FAF_{\mathcal{A}} that performs well on a classification task for which FVF_{\mathcal{V}} was designed. If F^V∼Of\hat{F}_{\mathcal{V}}\sim O_{f} then A\mathcal{A}’s goal is that F^A∼Of\hat{F}_{\mathcal{A}}\sim O_{f}, which can be considered successful if Acc(FA)∼Acc(FV)Acc(F_{\mathcal{A}})\sim Acc(F_{\mathcal{V}}). A secondary goal is to minimize the number of queries to FVF_{\mathcal{V}} necessary for A\mathcal{A} to train FAF_{\mathcal{A}}.

A\mathcal{A} has full control over the samples DAD_{\mathcal{A}} it chooses to query FVF_{\mathcal{V}} with. These can be natural (Orekondy et al., 2019) or synthetic (Juuti et al., 2019; Papernot et al., 2017; Tramèr et al., 2016). A\mathcal{A} obtains a prediction for each query in the form of probability vectors FV(x)F_{\mathcal{V}}(x) or single classes F^V(x),∀x∈DA\hat{F}_{\mathcal{V}}(x),\forall x\in D_{\mathcal{A}}. A\mathcal{A} uses queried samples and their predictions to train FAF_{\mathcal{A}}, a DNN. It chooses the DNN model architecture, training hyperparameters and training process. Requiring FAF_{\mathcal{A}} to be a DNN is justified by the observations in prior work on model extraction attacks (Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017) that FAF_{\mathcal{A}} needs to have equal or larger capacity than FVF_{\mathcal{V}} in order for model extraction to be successful. DNNs have the greatest capacity among ML models (Zhang et al., 2017).

2. Assumptions

We assume that for a given input x∈DAx\in D_{\mathcal{A}}, A\mathcal{A} has no a priori expectation regarding the prediction FV(x)F_{\mathcal{V}}(x). A\mathcal{A} treats y=FV(x)y=F_{\mathcal{V}}(x) as the ground truth label for x∈DAx\in D_{\mathcal{A}}. A\mathcal{A} expects that multiple queries of the same input xx must return the same prediction yy.

Our focus is on A\mathcal{A} who makes FAF_{\mathcal{A}} available via a prediction API since it has the greatest impact on V\mathcal{V}’s business advantage. We do not consider an A\mathcal{A} who keeps FAF_{\mathcal{A}} for private use. This is similar to media watermarking schemes where access to allegedly stolen media is a pre-requisite for ownership demonstration (Petitcolas et al., 1999).

3. DAWN Goals and Overview

On one hand, model extraction attacks against DNNs have been proven difficult to defend against (Juuti et al., 2019). On the other hand, existing watermarking techniques (Merrer et al., 2017; Adi et al., 2018; Darvish Rouhani et al., 2019; Chen et al., ; Li et al., ) are vulnerable to model extraction attacks (Zhang et al., 2018). To address these limitations, we design a solution to identify and prove the ownership of DNN models stolen through a prediction API.

Our solution, DAWN (Dynamic Adversarial Watermarking of Neural Networks), is an additional component added in front of a model prediction API (Fig. 1). DAWN dynamically embeds a watermark in responses to queries made by a API client. This watermark is composed of inputs xi∈Tx_{i}\in T for which we return incorrect predictions B(xi)≠FV(xi)B(x_{i})\neq F_{\mathcal{V}}(x_{i}). A\mathcal{A} uses all the responses including these mislabeled samples (xi,B(xi))(x_{i},B(x_{i})) to train FAF_{\mathcal{A}}. FAF_{\mathcal{A}} will remember those samples as a backdoor (Chen et al., 2017) that represents the watermark (as in traditional DNN watermarking techniques). If FAF_{\mathcal{A}} exposes a public prediction API, a judge J\mathcal{J} can run a verification process (verify), which confirms FAF_{\mathcal{A}} is a surrogate of FVF_{\mathcal{V}}. Verify checks that for sufficient number of inputs xi∈Tx_{i}\in T, we have F^A(xi)=B^(xi)≠F^V(xi)\hat{F}_{\mathcal{A}}(x_{i})=\hat{B}(x_{i})\neq\hat{F}_{\mathcal{V}}(x_{i}). DAWN embeds a watermark into a subset of queries it receives so that any FAF_{\mathcal{A}} trained using these responses will retain the watermark.

4. System requirements

We define the following requirements for the watermark that DAWN embeds in FAF_{\mathcal{A}} during an extraction attack. W1-W3 were introduced in (Adi et al., 2018) while W4 is a new requirement specific to DAWN.

Unremovability: A\mathcal{A} is unable to remove the watermark from FAF_{\mathcal{A}} without significantly decreasing its accuracy, rendering it “unusable”. If FVF_{\mathcal{V}} is free of the watermark, then Acc(FA)≪Acc(FV)Acc(F_{\mathcal{A}})\ll Acc(F_{\mathcal{V}}).

Reliability: If verify outputs “true” for a watermark (T,B^V(T))(T,\hat{B}_{\mathcal{V}}(T)) on a model F′F^{\prime}, then F′F^{\prime} is a surrogate of FVF_{\mathcal{V}}, with high confidence. On the other hand, if F′F^{\prime} is not a surrogate, A\mathcal{A} cannot generate a watermark (T,B^(T))(T,\hat{B}(T)) such that verify outputs “true” (non-trivial ownership).

Non-ownership piracy: A\mathcal{A} cannot produce a watermark for a model that was already watermarked by V\mathcal{V}, such that it can cast V\mathcal{V}’s ownership into doubt.

Linkability: If verify outputs “true” for a model FAF_{\mathcal{A}}, the watermark used for verification (T,B^(T))(T,\hat{B}(T)) can be linked to a specific API client whose queries were used to train FAF_{\mathcal{A}}.

We identify additional requirements X1-X3:

Utility: Incorrect predictions returned by DAWN do not significantly degrade the prediction service provided to legitimate API clients: Acc(\textscDAWN+FV)∼Acc(FV)Acc(\textsc{DAWN}+F_{\mathcal{V}})\sim Acc(F_{\mathcal{V}}).

Indistinguishability: A\mathcal{A} cannot distinguish incorrect predictions B(x)B(x) from correct victim model predictions FV(x)F_{\mathcal{V}}(x).

Collusion resistance: Watermark unremovability (W1), linkability (W4) and indistinguishability (X2) must remain valid even if the extraction attack is distributed among several API clients.

5. Relation to other attacks

Dynamic Adversarial Watermarks

We first present the method for generating and embedding an adversarial watermark. Then we describe the process for proving ownership of a model using the watermark.

We define watermarking an input xx as returning an incorrect prediction BV(x)B_{\mathcal{V}}(x) instead of the correct prediction FV(x)F_{\mathcal{V}}(x). The collection of all watermarked inputs composes the trigger set TAT_{\mathcal{A}} that will be a backdoor to any FAF_{\mathcal{A}} trained using responses from FVF_{\mathcal{V}} including TAT_{\mathcal{A}}. Consequently, inputs x∈TAx\in T_{\mathcal{A}} and their corresponding prediction classes B^V(x)\hat{B}_{\mathcal{V}}(x) compose the watermark to the surrogate model (TA,B^V(TA))(T_{\mathcal{A}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}})). We define two functions:

WV(x)W_{\mathcal{V}}(x): should the response to xx be watermarked?

BV(x)B_{\mathcal{V}}(x): what is the (backdoored watermark) response?

A\mathcal{A} must not be able to predict WV(x)W_{\mathcal{V}}(x) or distinguish between BV(x)B_{\mathcal{V}}(x) and FV(x)F_{\mathcal{V}}(x). The same query, regardless of the API client, must always get the same output. Both functions must be deterministic random functions specific to FVF_{\mathcal{V}} to fulfill these properties.

We use the result of a keyed cryptographic hash function as a source for randomness. We compute HMAC(Kw,x)\text{HMAC}(K_{w},x) using SHA-256, where KwK_{w} is a model-specific secret key generated by DAWN and xx is an input to FVF_{\mathcal{V}}. If xx is a matrix of dimension d>1d>1, it is flattened to a 1-dimensional vector. The result of the hash is split in two parts HMAC(Kw,x)\text{HMAC}(K_{w},x) and HMAC(Kw,x)\text{HMAC}(K_{w},x), respectively used in WVW_{\mathcal{V}} and BVB_{\mathcal{V}}. These numbers are independent and provide a sufficient source for randomness for each function.

WV(x)W_{\mathcal{V}}(x) is a boolean function. We define rwr_{w} as the fraction of inputs to be watermarked out of NN inputs submitted by an API client. rwr_{w} will define the size of the trigger set ∣TA∣=⌊rw×N⌋|T_{\mathcal{A}}|=\lfloor r_{w}\times N\rfloor. Then:

The expectation that WVW_{\mathcal{V}} returns 11 and thus to watermark a sample is uniformly equal to rwr_{w}. It is worth noting that DAWN does not differentiate adversaries from benign API clients. Consequently, any API client obtains a rate rwr_{w} of incorrect predictions. rwr_{w} must be defined to meet a trade-off. A large rwr_{w} increases the reliability of ownership demonstration and prevents trivial ownership demonstration W2 as later discussed in Sect. 4.3. A small rwr_{w} maximizes utility X1 by minimizing the number of incorrect predictions returned to benign API clients.

1.2. Backdoor function

π(Kπ,FV(x))\pi(K_{\pi},F_{\mathcal{V}}(x)) does not need to permute all mm positions of FV(x)F_{\mathcal{V}}(x) but only those with highest probabilities for the purpose of backdooring. A large number of classes typically have a 0 probability value when mm is large. Considering that the number of positions to permute is small, we use the Fisher-Yates shuffle algorithm (Fisher et al., 1949) to implement π\pi. We use Kπ=HMAC(Kw,x)[128,255]K_{\pi}=\text{HMAC}(K_{w},x)\left[128,255\right] as the key that determines the permutations performed during the Fisher-Yates shuffle algorithm. A 128-bits key allows for list permutation of up to 34 positions (34 prediction probabilities) in a secure manner.

1.3. Indistinguishability

Outputs BV(x)B_{\mathcal{V}}(x) must be indistinguishable from FV(x)F_{\mathcal{V}}(x) X2. This requirement is partially addressed by our assumption that A\mathcal{A} has no expectation regarding predictions obtained from FVF_{\mathcal{V}} (Sect. 3.2). Nevertheless, our watermarking function WVW_{\mathcal{V}} is configured by a hash of the input xx. A subtle modification δ\delta to xx produces a different hash and consequently, a different result WV(x)≠WV(x+δ)W_{\mathcal{V}}(x)\neq W_{\mathcal{V}}(x+\delta). If A\mathcal{A} receives different predictions for xx and x+δx+\delta for a small δ\delta, it can discard both xx and x+δx+\delta from its training set to avoid the watermark.

For each input xx, we obtain its latent representation LVL_{\mathcal{V}} based on FVF_{\mathcal{V}} as xL=LV(FV,x)x_{L}=L_{\mathcal{V}}(F_{\mathcal{V}},x). This ensures that as long as FVF_{\mathcal{V}}’s prediction is resilient to perceptual modifications (e.g. translation, illumination), so is MVM_{\mathcal{V}}. Next, we smoothen xLx_{L} by binarizing it based on the median value of each of its features. The median of each feature value is obtained by querying FVF_{\mathcal{V}} using V\mathcal{V}’s training set, recording corresponding xLx_{L} and taking the median. Using 5000 samples and their intermediate representation of length 100, we get 100 feature vectors of length 5000 and thus, 100 median values. We evaluate MVM_{\mathcal{V}} in Sect. 6.2.

2. Watermark embedding

A\mathcal{A} uses the set of inputs DAD_{\mathcal{A}} and the corresponding predictions returned by DAWN-protected prediction API of FVF_{\mathcal{V}} to train FAF_{\mathcal{A}}. Approximately ⌊rw×∣DA∣⌋\lfloor r_{w}\times|D_{\mathcal{A}}|\rfloor samples from DAD_{\mathcal{A}} constitute the trigger set TAT_{\mathcal{A}} consisting of incorrect predictions BV(x)B_{\mathcal{V}}(x). Given that FAF_{\mathcal{A}} has enough capacity (large enough number of parameters), it will be able to remember a certain amount of training data having arbitrarily incorrect labels (Zhang et al., 2017). This phenomenon is called overfitting and it can be prevented using regularization (Bishop, 2006). But it is not effective for DNNs with a large capacity (Zhang et al., 2017). This is the rationale for the existence of DNN backdoors (Liu et al., 2018) and for DAWN. We expect our watermark (TA,B^V(TA))(T_{\mathcal{A}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}})) to be embedded as a backdoor in FAF_{\mathcal{A}} as a natural effect of training a model FAF_{\mathcal{A}} with high capacity. If the watermark is not embedded, we expect FAF_{\mathcal{A}}’s accuracy on the primary task to be too low to make it usable (W1).

Different adversaries Ai\mathcal{A}_{i} will have different datasets DAiD_{\mathcal{A}_{i}}. Consequently, the trigger sets TAiT_{\mathcal{A}_{i}} selected by DAWN will also be different. Different surrogate models FAiF_{\mathcal{A}_{i}} will embed distinctive watermarks. Each watermark thus links to the API client identifier. DAWN meets the linkability requirement W4.

3. Watermark verification

We present the verify function used by J\mathcal{J} to prove a model F′F^{\prime} is a surrogate of FF. Verify tests if a given watermark (T,B^(T))(T,\hat{B}(T)) is embedded in a model F′F^{\prime} suspected to be a surrogate of FF. We first define L(T,B^(T),F′)L(T,\hat{B}(T),F^{\prime}) that computes the ratio of different results between the backdoor function B^(x)\hat{B}(x) and the suspected surrogate model F^′(x)\hat{F}^{\prime}(x) for all inputs in the trigger set.

The watermark verification succeeds, i.e., verify returns “true”, if and only if L(T,B^(T),F′)<eL(T,\hat{B}(T),F^{\prime})<e, where ee is a tolerated error rate that must be defined. This means we must have at most ⌊e×∣T∣⌋\lfloor e\times|T|\rfloor samples where B^(x)\hat{B}(x) and F^′(x)\hat{F}^{\prime}(x) differ in order to declare F′F^{\prime} is a surrogate of FF. The choice for the value of ee is a trade-off between correctness and completeness for watermark verification (reliability W2). Assume we want to use a pre-generated watermark (T,B^(T))(T,\hat{B}(T)) to verify if an arbitrary model F′F^{\prime} is a surrogate. For simplicity, we assume a uniform probability of matching the prediction of a watermarked input P(B^(x)=F^′(x))=1/mP(\hat{B}(x)=\hat{F}^{\prime}(x))=1/m, where mm is the number of classes of F′F^{\prime}. The probability for trivial watermark verification success, given a trigger set of size ∣T∣|T| and an error rate ee, can be computed using the cumulative binomial distribution function as follows.

This probability is the average success rate of A\mathcal{A} wanting to frame V\mathcal{V} for model stealing using an arbitrary watermark. Figure 2 depicts the decrease of this success rate as we increase the watermark size. We see that the verification function can accommodate a large error rate (e>0.5e>0.5) while preventing trivial success in verification using a small watermark (∣T∣≈50|T|\approx 50). The error rate ee must be defined proportionally to the number of classes mm. Large error rates can be used for models with a large number of classes. For instance, we can set e=0.8e=0.8 for a model with m=256m=256 classes, limiting the adversary success rate to less than 2−642^{-64} for a watermark of size 70.

The success rate in trivial verification is the complement of the confidence for reliable watermark verification, and for reliable demonstration of ownership by transition 1−P(L<e)1-P(L<e). The choice of ee defines the minimum watermark size given a targeted confidence. Recall that this size must also be small to ensure utility of the model to protect X1. The tolerated error must necessarily be lower than the probability of random class match: e<(1−m)/me<(1-m)/m. Also, ee must be larger than ϵ\epsilon where Acc(FA)=1−ϵAcc(F_{\mathcal{A}})=1-\epsilon is the accuracy of the watermarked surrogate model FAF_{\mathcal{A}} on the trigger set.

The success of watermark verification is not sufficient to declare ownership of a surrogate model F′F^{\prime}. A\mathcal{A} can increase its success in trivial watermark verification from random using several means. For instance, knowing FF and F′F^{\prime}, A\mathcal{A} can find inputs xx for which F(x)≠F′(x)F(x)\neq F^{\prime}(x) and use pairs (x,F′(x))(x,F^{\prime}(x)) as a watermark that would successfully pass watermark verification. Thus demonstrating ownership requires a careful process to ensure that the probability for matching an incorrect prediction class remains random, ensuring that the probability for trivial watermark verification follows Eq. 3.

4. Demonstrating ownership

We present the process for a model owner V\mathcal{V} to demonstrate ownership of a surrogate model watermarked by DAWN. It only requires the suspected surrogate model FAF_{\mathcal{A}} to expose a prediction API. This process uses a judge J\mathcal{J} who is trusted to (a) ensure confidentiality of all data submitted as input to the process and (b) correctly execute and report the results of the specified verify. It also uses a time-stamped public bulletin board, e.g., a blockchain, in which information can be published to provide proof of anteriority. J\mathcal{J} can be implemented using an trusted execution environment (TEE) (Ekberg et al., 2014).

V\mathcal{V} publishes cryptographic commitments of the following elements in the public bulletin board:

for each API client ii, one registered watermark (TAi,B^V(TAi))(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}})).

The commitment can be instantiated using a cryptographic hash function H( )H(\>), e.g., SHA-3. Each watermark should be linked to the corresponding model, e.g., by associating H(FV)H(F_{\mathcal{V}}) with each registered watermark.

Several updated versions of the registered watermark can be published for each API client, as they make more queries to the prediction API and their watermarks grow. The verification of any one of these watermarks is sufficient to demonstrate ownership of the model. We define the following rules for reliable demonstration of ownership W2 that prevents ownership piracy W3:

(H(TAi,B^V(TAi)),H(FV))\left(H(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}})),H(F_{\mathcal{V}})\right) is valid only if published later than H(FV)H(F_{\mathcal{V}}).

A\mathcal{A} can refute FAF_{\mathcal{A}} is a surrogate model only if H(FA)H(F_{\mathcal{A}}) has been published.

(H(TAi,B^V(TAi)),H(FV))\left(H(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}})),H(F_{\mathcal{V}})\right) can only demonstrate that FAF_{\mathcal{A}} is a surrogate of FVF_{\mathcal{V}} if H(FA)H(F_{\mathcal{A}}) is published later than H(FV)H(F_{\mathcal{V}}) (or not published at all).

in case of contention, the model having its commitment first published is deemed to be the original.

4.2. Verification process

When V\mathcal{V} suspects a model FAF_{\mathcal{A}} is a surrogate of FVF_{\mathcal{V}} trained by an API client ii, it provides a pointer to the prediction API of FAF_{\mathcal{A}} to J\mathcal{J}. It also provides the following secret information using a confidential communication channel: the API client ii watermark (TAi,B^V(TAi))(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}})) and FVF_{\mathcal{V}}. J\mathcal{J} does the following to check if FAF_{\mathcal{A}} is a surrogate of FVF_{\mathcal{V}}. If any step fails, the ownership of FAF_{\mathcal{A}} is not considered to have been demonstrated. If all succeed, J\mathcal{J} gives the verdict that FAF_{\mathcal{A}} is a surrogate of FVF_{\mathcal{V}}.

compute H(TAi,B^V(TAi))H(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}})) and use it as a pointer to retrieve the registered watermark (H(TAi,B^V(TAi)),H(FV′))\left(H(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}})),H(F_{\mathcal{V}}^{\prime})\right) from the public bulletin.

compute H(FV)H(F_{\mathcal{V}}) and verify H(FV)=H(FV′)H(F_{\mathcal{V}})=H(F_{\mathcal{V}}^{\prime}), where H(FV′)H(F_{\mathcal{V}}^{\prime}) is extracted from the registered watermark.

retrieve H(FV)H(F_{\mathcal{V}}) from the public bulletin and verify it was published before (H(TAi,B^V(TAi)),H(FV′))\left(H(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}})),H(F_{\mathcal{V}}^{\prime})\right).

query TAiT_{\mathcal{A}_{i}} to FAF_{\mathcal{A}}’s prediction API and verify that

L(TAi,B^V(TAi),FA)<eL(T_{\mathcal{A}_{i}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}_{i}}),F_{\mathcal{A}})<e.

input TAiT_{\mathcal{A}_{i}} to FVF_{\mathcal{V}} and verify B^V(x)≠F^V(x),∀x∈TAi\hat{B}_{\mathcal{V}}(x)\neq\hat{F}_{\mathcal{V}}(x),\forall x\in T_{\mathcal{A}_{i}}.

If FAF_{\mathcal{A}}’s owner (A\mathcal{A}) wants to contest the verdict, it must provide the original model FA′F^{\prime}_{\mathcal{A}} to J\mathcal{J} using a confidential communication channel. J\mathcal{J} assesses that the provided model and the API model are the same FA=FA′F_{\mathcal{A}}=F^{\prime}_{\mathcal{A}} by verifying FA(x)=FA′(x),∀x∈TAiF_{\mathcal{A}}(x)=F^{\prime}_{\mathcal{A}}(x),\forall x\in T_{\mathcal{A}_{i}}. Then, J\mathcal{J} computes H(FA′)H(F^{\prime}_{\mathcal{A}}) and retrieves it from the public bulletin. If H(FA′)H(F^{\prime}_{\mathcal{A}}) was published before H(FV)H(F_{\mathcal{V}}), J\mathcal{J} concludes that FA′=FAF^{\prime}_{\mathcal{A}}=F_{\mathcal{A}} is an original model.

Experimental setup

We evaluate DAWN using four image recognition datasets that were used in prior work to evaluate DNN extraction attacks. MNIST (LeCun et al., 2010) (60,000 train and 10,000 test samples, 10 classes) and GTSRB (Stallkamp et al., 2011) (39,209 train and 12,630 test samples, 43 classes) are respectively a handwritten-digit and traffic-sign dataset used to showcase the extraction of low capacity DNN models (Juuti et al., 2019; Papernot et al., 2017). CIFAR10 (Krizhevsky, 2009) (50,000 train and 10,000 test samples, 10 classes) and Caltech256 (Griffin et al., 2007) (23,703 train and 6,904 test samples, 256 classes) contain images depicting miscellaneous objects that were used to showcase the extraction of high capacity DNN models (Orekondy et al., 2019; Correia-Silva et al., 2018).

We also selected a random subset of 100,000 samples from ImageNet dataset (Deng et al., 2009) (1000 classes), which contains images of natural and man-made objects. We use it to evaluate the embedding of different types of watermarks and to perform a model extraction attack that requires such samples (Orekondy et al., 2019).

1.2. Models

We select two kinds of DNN models to evaluate the embedding of a watermark: low-capacity models having less than 10M parameters, and high-capacity models having over 20M parameters. These models are presented in Table 2.

In order to accurately reconstruct model extraction attacks, we use the same model architectures and training process as in (Juuti et al., 2019) for low-capacity models and as in (Orekondy et al., 2019) for high-capacity models. Similarly to prior work (Orekondy et al., 2019), we use ResNet34 (He et al., 2016) architecture pre-trained on ImageNet as a basis for high-capacity models. We fine-tuned Caltech-RN34, GTSRB-RN34 and CIFAR10-RN34 models using Caltech256, GTSRB and CIFAR10 datasets respectively We chose to reproduce only the Caltech-RN34 experiment from (Orekondy et al., 2019) because of its best performance. We used CIFAR10 and GTSRB to conduct supplementary experiments with high capacity models as they allow us to juxtapose results of experiments with low and high capacity models on the same datasets.. We also trained DenseNet121 (Huang et al., 2017) models to perform additional experiments due to the absence of dropout layers in ResNet34 models. All models were trained using Adam optimizer with learning rate of 0.001 that was decreased over time to 0.0005 (after 100 epochs for ResNet34 models and half-way for the other), except for Caltech-RN34. For Caltech-RN34, we used SGD optimizer with an initial learning rate of 0.1 that was decreased by a factor of 10 every 60 epochs over 250 epochs. We used a batch size of 16 for fine-tuning ResNet34 and DenseNet121 based models.

2. Watermarking Procedure

Inputs from A\mathcal{A}’s dataset DAD_{\mathcal{A}} are submitted to the DAWN-enhanced prediction API of FVF_{\mathcal{V}} which returns correct FV(x)F_{\mathcal{V}}(x) or incorrect predictions BV(x)B_{\mathcal{V}}(x) according to the result of the watermarking function WV(x)W_{\mathcal{V}}(x). For the experiments in Section 6.2 (evaluating the effectiveness of the mapping function MVM_{\mathcal{V}}), we use the embedding from FVF_{\mathcal{V}} as MVM_{\mathcal{V}}. Experiments in Section 6.1 and Section 7 do not depend on the choice of MVM_{\mathcal{V}}. Therefore, for the sake of simplicity, we use the identity function as MVM_{\mathcal{V}} in these experiments.

We simulate A\mathcal{A} who uses the whole set DAD_{\mathcal{A}}, which includes ∣TA∣|T_{\mathcal{A}}| samples with incorrect labels, to train its surrogate model FAF_{\mathcal{A}}. A\mathcal{A} trains FAF_{\mathcal{A}} without being aware of the watermarked samples in DAD_{\mathcal{A}}.

3. Evaluation Metrics

We use two metrics to evaluate the success of A\mathcal{A}’s goal and V\mathcal{V}’s goal respectively. A\mathcal{A}’s goal is to train a surrogate model FAF_{\mathcal{A}} that has maximum accuracy on FVF_{\mathcal{V}}’s primary classification task. We evaluate this by computing the test accuracy of the surrogate model Acctest(FA)Acc_{test}(F_{\mathcal{A}}) on the test set TestTest of each dataset.

V\mathcal{V}’s goal is to maximize the embedding of the watermark in any surrogate model built from responses from FVF_{\mathcal{V}} such that its surrogacy can be reliably demonstrated. We evaluate this by computing the watermark accuracy of the surrogate model Accwm(FA)Acc_{wm}(F_{\mathcal{A}}) on the trigger set TAT_{\mathcal{A}} of watermarked inputs.

DAWN aims to maximize Accwm(FA)Acc_{wm}(F_{\mathcal{A}}) regardless of Acctest(FA)Acc_{test}(F_{\mathcal{A}}). A\mathcal{A} aims to maximize Acctest(FA)Acc_{test}(F_{\mathcal{A}}) while minimizing Accwm(FA)Acc_{wm}(F_{\mathcal{A}}). In our experiments, we calculate both metrics every 5 epochs in order to evaluate their progress during the training process.

Robustness of watermarking

We assess A\mathcal{A}’s ability to prevent the embedding of a watermark in a surrogate model, i.e., to violate the unremovability requirement W1. Prior work evaluated unremovability after a watermarked model is trained showing that backoor-based watermarks are resilient to model pruning and adversarial fine tuning (Adi et al., 2018; Merrer et al., 2017; Zhang et al., 2018). DAWN also embeds backdoor-based watermarks resilient to removal using post-training manipulations. Thus, we focus on adversarial manipulations during training by evaluating several solutions that could prevent watermark embedding. We then evaluate the ability for A\mathcal{A} to identify watermarked inputs using the trained surrogate model, i.e., to violate the indistinguishability requirement X2.

We take an ideal model extraction attack scenario where F^V=Of\hat{F}_{\mathcal{V}}=O_{f} is a perfect oracle. A\mathcal{A} has access to a large dataset DAD_{\mathcal{A}} of natural samples from the same distribution as V\mathcal{V} training data: we use the whole training set from each dataset (Sect. 5.1) for DAD_{\mathcal{A}}. We use a large watermark of fixed size ∣TA∣=250|T_{\mathcal{A}}|=250 in all following experiments. Embedding a large watermark is challenging since the model must learn many isolated errors (mislabeled inputs). We take ∣TA∣=250|T_{\mathcal{A}}|=250 as an upper bound to the watermark size and a worst case scenario for DAWN watermark embedding.

We evaluate the impact of two parameters on embedding a watermark during DNN training. The first parameter is the capacity of FAF_{\mathcal{A}}. A\mathcal{A} can limit this capacity such that the model could only learn the primary classification task and cannot learn the watermark. The second parameter is the use of regularization. Regularization accommodates classification errors on the training data, which is considered as noise. The watermark consists of incorrectly labeled inputs which can potentially be discarded using regularization.

We evaluate the impact of model capacity and regularization on watermark accuracy AccwmAcc_{wm} and test accuracy AcctestAcc_{test} of FAF_{\mathcal{A}}. We trained several surrogate models having low and high capacity. TAT_{\mathcal{A}} was randomly selected from the respective training sets. We used plain training and two regularization methods, namely weight decay (Krogh and Hertz, 1992) with decaying factor λ\lambda and dropout (DO=X) (Srivastava et al., 2014) with probability X={0.3,0.5}\left\{0.3,0.5\right\}. We selected λ\lambda values optimal for A\mathcal{A}: such that they maximize the difference Acctest−AccwmAcc_{test}-Acc_{wm}.

Table 3(a) and 3(b) present the results of this experiment for DNN models with low and high capacity respectively. We report AccwmAcc_{wm} and AcctestAcc_{test} results at three training stages providing (1) best watermark accuracy (best for V\mathcal{V}), (2) best test accuracy (best for A\mathcal{A}) and (3) when training is completed. Overall, we observe that AcctestAcc_{test} and AccwmAcc_{wm} are high for most settings. Using plain training, AccwmAcc_{wm} is mostly higher than AcctestAcc_{test} and often close to 100%. The ownership of all these surrogate models can be reliably demonstrated using a low tolerated error rate, e.g., e=0.3e=0.3.

Model capacity. High-capacity models can provide higher watermark and test accuracy than low-capacity models as highlighted by comparing results for GTSRB and CIFAR10 in both tables. While AccwmAcc_{wm} is low for some low-capacity models, e.g., MNIST-3L, MNIST-5L (DO), their test accuracy is similarly low and close to random Accwm∼Acctest∼10%Acc_{wm}\sim Acc_{test}\sim 10\%. This shows that reducing the model capacity can prevent the embedding of the watermark. However, decreasing AccwmAcc_{wm} to a level where it cannot be used to reliably prove ownership makes FAF_{\mathcal{A}} unusable. AccwmAcc_{wm} and AcctestAcc_{test} are closely tied when manipulating the model capacity and thus this is not a useful strategy to circumvent DAWN.

Regularization. Regularization is useful for decreasing the watermark accuracy in a few cases. Weight decay is useful for low-capacity GTSRB-5L and CIFAR10-9L models. Dropout is useful for low-capacity MNIST-5L and CIFAR10-9L models, and for high-capacity Caltech-DN121 model. Dropout completely prevents the embedding of the watermark into MNIST-5L model as depicted by Accwm∼10%Acc_{wm}\sim 10\%. However, AcctestAcc_{test} is also significantly reduced, by 50% at best, making FAF_{\mathcal{A}} potentially unusable. In all remaining cases, AccwmAcc_{wm} is reduced down to 20-35%, while preserving high test accuracy similar to models trained with non-watermarked datasets. While AccwmAcc_{wm} is low, the watermark can still successfully demonstrate ownership by increasing the tolerated error rate to, e.g., e=0.8>1−Accwme=0.8>1-Acc_{wm}. Considering the large watermark size of 250, this demonstration would still be reliable despite the high tolerated error rate as evaluated in Sect. 4.3.

It is worth noting that no regularization method is effective at removing the watermark from high capacity GTSRB-RN34 and CIFAR10-RN34 models. The likely reason is that ResNet34 architecture has significant overcapacity for the primary task of classifying these datasets. Regularization cannot limit this capacity to an extent where the watermark would not be embedded. This means A\mathcal{A} needs sufficient knowledge of FVF_{\mathcal{V}} to select an appropriate model architecture for FAF_{\mathcal{A}}. It must have sufficient capacity to learn the primary classification task of the victim model while preventing watermark embedding. In model extraction attacks, A\mathcal{A} has black-box access to FVF_{\mathcal{V}}, which forces to use FAF_{\mathcal{A}} with sufficient capacity to maximize the attack success (Orekondy et al., 2019). In this setting, regularization is not useful to circumvent DAWN.

Finally, while regularization can be useful, A\mathcal{A} needs relevant test data and ground truth to optimize the regularization parameters (e.g., decaying factor λ\lambda). In all extraction attacks (Tramèr et al., 2016; Juuti et al., 2019; Orekondy et al., 2019; Papernot et al., 2017; Correia-Silva et al., 2018; Pal et al., 2020) the availability of relevant data is the main limitation. All this data is typically used for training the surrogate model and none is used for test purposes, which prevents optimization of regularization parameters and early stopping.

2. Mapping Function

A\mathcal{A} can try to identify watermarked inputs and remove them from DAD_{\mathcal{A}} prior to training in order to prevent watermark embedding. Because DAWN relies on a hash to decide if an input is watermarked, A\mathcal{A} can query multiple perturbed versions of inputs in DAD_{\mathcal{A}} and discard those that return different predictions. The mapping function MVM_{\mathcal{V}} presented in Sect. 4.1.3 is meant to prevent this evasion. We evaluate the effectiveness of MVM_{\mathcal{V}} by querying 10 perturbed versions of each of the 10,000 samples in DAD_{\mathcal{A}}, which includes ∣TA∣=121|T_{\mathcal{A}}|=121. For each query, we check whether they get consistent mapping MVM_{\mathcal{V}} and classification F^V\hat{F}_{\mathcal{V}}. Table 4 reports the results of this experiment for various perturbation size δ\delta for the MNIST dataset. We distinguish cases where 1 F^V(x)=F^V(x+δ)\hat{F}_{\mathcal{V}}(x)=\hat{F}_{\mathcal{V}}(x+\delta) (same F^V\hat{F}_{\mathcal{V}}) or F^V(x)≠F^V(x+δ)\hat{F}_{\mathcal{V}}(x)\neq\hat{F}_{\mathcal{V}}(x+\delta) (diff F^V\hat{F}_{\mathcal{V}}); 2 MV(x)=MV(x+δ)M_{\mathcal{V}}(x)=M_{\mathcal{V}}(x+\delta) (same MVM_{\mathcal{V}}) or MV(x)≠MV(x+δ)M_{\mathcal{V}}(x)\neq M_{\mathcal{V}}(x+\delta) (diff MVM_{\mathcal{V}}). Same F^V\hat{F}_{\mathcal{V}} and same MVM_{\mathcal{V}} means A\mathcal{A} keeps a watermarked sample in DAD_{\mathcal{A}} (MVM_{\mathcal{V}} succeeds). Same F^V\hat{F}_{\mathcal{V}} and different MVM_{\mathcal{V}} means A\mathcal{A} discards a watermarked sample from DAD_{\mathcal{A}} (MVM_{\mathcal{V}} fails). Different F^V\hat{F}_{\mathcal{V}} means A\mathcal{A} wrongfully discards a sample from DAD_{\mathcal{A}} regardless of MVM_{\mathcal{V}} (δ\delta is too large and changes FVF_{\mathcal{V}}’s prediction). We see MVM_{\mathcal{V}} succeeds to provide a consistent mapping in over 85% cases for δ≤0.1\delta\leq 0.1, meaning 85% of the TAT_{\mathcal{A}} is preserved in DAD_{\mathcal{A}}. As perturbations δ\delta increase in size, MVM_{\mathcal{V}} returns an increasing rate of inconsistent mapping, but this rate is similar to the one of changed predictions from FVF_{\mathcal{V}}. Thus, we conclude MVM_{\mathcal{V}} is resilient to perturbations and DAWN can effectively watermark FAF_{\mathcal{A}}.

Protecting against model extraction attacks

We evaluate DAWN’s effectiveness at watermarking surrogate DNN models constructed using two model extraction attacks: 1) the PRADA attack (Juuti et al., 2019) achieves state-of-the-art performance in extracting low-capacity DNN models primarily using synthetic data and we launch it against MNIST-5L, GTSRB-5L and CIFAR10-9L; 2) the KnockOff attack (Orekondy et al., 2019) extracts high-capacity DNN models using only natural data and we launch it against GTSRB-RN34, CIFAR10-RN34 and Caltech-RN34. The test accuracy of each FAF_{\mathcal{A}} extracted with these respective attacks is reported in Tab. 6.

We demonstrate how to setup DAWN to protect a given victim model FVF_{\mathcal{V}}. We evaluate the successful embedding of watermarks in several surrogate models FAF_{\mathcal{A}} as well as their utility considering a circumvention strategy.

Watermarking decision: DAWN degrades FVF_{\mathcal{V}} utility by a factor equal to rw×Acc(FV)r_{w}\times Acc(F_{\mathcal{V}}) due to incorrect predictions for watermarked inputs. The value of rwr_{w} is specific to FVF_{\mathcal{V}}. Given a desired level of confidence for reliable ownership demonstration equal to 1−P(L<e)1-P(L<e) (cf. Eq. 3), a tolerated error rate ee and the number of classes mm for FVF_{\mathcal{V}}, we can compute the minimum size for the watermark ∣TA∣|T_{\mathcal{A}}| using Eq. 3. Given that V\mathcal{V} can estimate the minimum number of queries NN required by A\mathcal{A} to train a usable surrogate model for FVF_{\mathcal{V}}, we can compute rw=N/∣TA∣r_{w}=N/|T_{\mathcal{A}}|. This ratio ensures that if A\mathcal{A} can successfully train a usable surrogate model FAF_{\mathcal{A}}, then FAF_{\mathcal{A}} will embed a watermark large enough to reliably demonstrate its ownership .

The probability for successful trivial watermark verification P(L<e)P(L<e) is valid for testing a single watermark. This probability increases by a factor equal to the number of tested watermarks. DAWN creates and registers client-specific watermarks. V\mathcal{V} must estimate the number of API clients to calculate the actual probability for trivial demonstration of ownership considering that all registered watermarks should be tested. When verifying a watermark, the judge J\mathcal{J} counts the number of registered watermarks for FVF_{\mathcal{V}} in the public bulletin. J\mathcal{J} computes the real probability for successful trivial watermark verification accordingly and decides if a demonstration of ownership is reliable or not according to this final confidence.

Utility for legitimate clients: Suppose we want a confidence for reliable demonstration of ownership equal to 1−2−641-2^{-64}. FVF_{\mathcal{V}} has a prediction API with 1M API clients (1M watermarks are registered for FVF_{\mathcal{V}}). We need P(L<e)<10−6×2−64=5.4×10−26P(L<e)<10^{-6}\times 2^{-64}=5.4\times 10^{-26} to be able to test all registered watermarks while achieving our targeted confidence. We choose a tolerated error rate e=0.5e=0.5. Table 5 reports the computed watermark ratio rwr_{w} required to protect six models against model extraction. We see rwr_{w} must always be lower than 0.5% to reach 1−2−641-2^{-64} confidence for any victim model. FVF_{\mathcal{V}}’s accuracy is thus degraded in a negligible manner that does not impact its utility. DAWN meets the reliability W2 and utility X1 requirements.

Overhead: Storing 1M watermarks would require at most a few TBs (cf. Tab. 5). Watermark verification consists in obtaining predictions from a purported surrogate model. It is operated by J\mathcal{J} who gets predictions at no monetary cost. Thus, demonstration of ownership is only a matter of time and getting one prediction from our most complex model (Caltech-RN34) takes 9ms (on Tesla P100 GPU). Verifying one watermark for this model takes 0.25s (27 queries) and verifying 100,000 watermarks takes 7 hours using a single GPU. J\mathcal{J} can initially verify all watermarks with a lower confidence to reduce this time (by testing only a subset of each watermark). Only successful verification would later undergo a verification of the full watermark. Testing the same 100,000 watermarks with 1−2−161-2^{-16} targeted confidence (instead of 1−2−641-2^{-64}) requires 1h15 (5 samples per watermark). This time can further be reduced by parallelizing predictions on several GPUs. DAWN’s verification process is more computationally expensive due to the requirement of testing all watermarks to account for Sybils. However, unlike prior watermarking schemes, DAWN is effective against model extraction attacks.

2. Effectiveness against real extraction attacks

We want to show that any surrogate FAF_{\mathcal{A}} of a victim model FVF_{\mathcal{V}} protected by DAWN will embed a watermark that allows for reliable demonstration of ownership. We evaluate the effectiveness of DAWN against two landmark model extraction attacks namely PRADA (Juuti et al., 2019) and KnockOff (Orekondy et al., 2019).

Low-capacity models expose a prediction API that returns prediction classes F^V\hat{F}_{\mathcal{V}} required for the PRADA attack. High-capacity models return the full probability vector FVF_{\mathcal{V}}. Each victim model is protected by DAWN using the setting presented in Sect. 7.1. This setting enables V\mathcal{V} to demonstrate ownership of each surrogate model with confidence 1−2−641-2^{-64} using a tolerated error rate e=0.5e=0.5. For demonstration of ownership to be successful, the surrogate model FAF_{\mathcal{A}} must pass the watermark verification test L(TA,B^V(TA),FA)<eL(T_{\mathcal{A}},\hat{B}_{\mathcal{V}}(T_{\mathcal{A}}),F_{\mathcal{A}})<e. In our setting, it means that DAWN successfully defends against an extraction attack if the watermark accuracy for FAF_{\mathcal{A}} is larger than 50%, i.e., Accwm(FA)>1−eAcc_{wm}(F_{\mathcal{A}})>1-e.

Table 6 presents the result of this experiment. We see all surrogate models have a watermark accuracy Accwm≥50%Acc_{wm}\geq 50\%, which means V\mathcal{V} is successful in demonstrating their ownership. DAWN successfully defends against the PRADA and KnockOff attacks for all tested models while incurring little decrease in FVF_{\mathcal{V}}’s utility (evaluated in Sect. 7.1). We have shown DAWN effectively embeds a watermark in surrogate models FAF_{\mathcal{A}} stolen using extraction attacks. In Table 6, note that DAWN significantly decreases the surrogate model test accuracy (AcctestAcc_{test} for FAF_{\mathcal{A}}) for MNIST-5L while it has little impact on the same for other datasets. Drastic reduction in AcctestAcc_{test} is not a concern from the defender’s perspective - in fact it can, by itself, serve as a deterrence for A\mathcal{A} against model extraction. In all cases, adequate watermark accuracy Accwm>0.5Acc_{wm}>0.5 serves as a deterrence.

3. Resilience to distributed extraction attack

Distributing a model extraction attack across several API clients means several adversaries Ai\mathcal{A}_{i} query a subset DAiD_{\mathcal{A}_{i}} from the whole set DAD_{\mathcal{A}} used to train the surrogate model FAF_{\mathcal{A}}. Recall that DAWN is a deterministic mechanism Sect. 4.1. The watermarking WVW_{\mathcal{V}} and backdoor BVB_{\mathcal{V}} functions are deterministic and specific to FVF_{\mathcal{V}}. Their results only depend on the input queried to FVF_{\mathcal{V}}. The responses to DAD_{\mathcal{A}}, and its corresponding trigger set, remain the same regardless of which client(s) query the prediction API. Thus, DAD_{\mathcal{A}} is labeled in the same manner and it includes the same trigger set TAT_{\mathcal{A}} whether it is queried by one or by multiple API clients. Thus, the watermark in FAF_{\mathcal{A}} trained using DAD_{\mathcal{A}} will remain indistinguishable X2 and unremovable W1 even if multiple clients collude.

Note that in the case of colluding clients, each adversary Ai\mathcal{A}_{i} has a subset TAiT_{\mathcal{A}_{i}} of the whole trigger set TAT_{\mathcal{A}}. When verifying ownership, the judge J\mathcal{J} will have several successful watermark verifications L(TAi,BV(TAi),FA)<eL(T_{\mathcal{A}_{i}},B_{\mathcal{V}}(T_{\mathcal{A}_{i}}),F_{\mathcal{A}})<e: one for each adversary Ai\mathcal{A}_{i} who colluded to build the surrogate model FAF_{\mathcal{A}}. The verification of each sub-watermark (TAi,BV(TAi))(T_{\mathcal{A}_{i}},B_{\mathcal{V}}(T_{\mathcal{A}_{i}})) has the same expectation for success as the verification of the whole watermark (TA,BV(TA))(T_{\mathcal{A}},B_{\mathcal{V}}(T_{\mathcal{A}})). J\mathcal{J} will conclude that each API client ii whose watermark is successfully verified is a perpetrator of the distributed extraction attack used to build the surrogate model FAF_{\mathcal{A}}. Linkability W4 remains valid in case of collusion.

In a distributed attack, the watermark associated to each colluding client is smaller than in a centralized attack. To verify ownership with a same reliability, we must increase the watermark size and consequently rwr_{w} by a factor equal to the number of colluding clients. We assume the number of real colluding clients is limited, e.g., a few tens. Nevertheless, it is possible to mount a Sybil attack in which several API accounts are created by a single adversary. The API account registration process must require providing information that maximizes difficulty of creating trusted accounts, e.g., verified phone number or credit card, to mitigate this threat. Also, Sybils-detection techniques exist (Tran et al., 2009; Wang et al., 2013) and it is possible to link Sybils accounts by examining querying patterns and IP addresses for instance (Stringhini et al., 2015). For example, to protect Caltech-RN34, we could increase rwr_{w} to reliably verify the watermark of 35 colluders while maintaining the utility loss below 1%. Consequently, the higher the number of classes, the greater the reliability of watermark verification (c.f. Eq. 3) and we can tolerate more Sybils. For a classifier with 10,000 classes and utility loss below 1% we can reliably verify the watermark of 87 colluders.

Watermark Removal

Several techniques can identify if a DNN model has a backdoor (Chen et al., 2019; Guo et al., 2019; Wang et al., 2019). Most techniques like Neural Cleanse (Wang et al., 2019) and TABOR (Guo et al., 2019) can only detect backdoors for which the trigger is a static pattern added to original inputs (e.g., yellow square added to an image). In contrast, our trigger set is composed of unmodified samples having only incorrect labels. Consequently, techniques like Neural Cleanse and TABOR are ineffective at detecting DAWN watermark. In this section, we evaluate the resilience of DAWN watermarks to removal using six attacks: (1) double-extraction of a second order surrogate model FA′F^{\prime}_{\mathcal{A}}, (2) fine-tuning (Kornblith et al., 2018), (3) pruning (Blalock et al., 2020), (4) training with noise, (5) adding noise during the inference, and (6) recognizing queries from training data.

Double extraction and fine-tuning. A watermark may be removed by performing an extraction attack against FAF_{\mathcal{A}} to obtain a second order surrogate model FA′F_{\mathcal{A}}^{\prime}. A\mathcal{A} has full control over FAF_{\mathcal{A}}: its prediction API is not protected by DAWN and does not intentionally return incorrect prediction. If A\mathcal{A} uses a disjoint set of queries to extract a surrogate FA′F_{\mathcal{A}}^{\prime} from FAF_{\mathcal{A}}, then FA′F_{\mathcal{A}}^{\prime} may not embed the watermark, preventing the demonstration of its ownership by V\mathcal{V}. Instead of starting the second extraction from scratch, it can use FAF_{\mathcal{A}} as the starting point for FA′F_{\mathcal{A^{\prime}}} and fine-tune it. We call this stealing+fine-tuning.

We observed in Tab. 6 that surrogate models have a lower accuracy than victim models because model extraction incurs a necessary decrease in surrogate model accuracy. We evaluate the extent of the decrease in AcctestAcc_{test} and AccwmAcc_{wm} if A\mathcal{A} launches two successive extraction attacks instead of one or steals+fine-tunes the model to obtain FA′F_{\mathcal{A}}^{\prime}: the first against FVF_{\mathcal{V}} and the second against FAF_{\mathcal{A}}. We evaluate these evasion techniques using the PRADA and KnockOff attacks.

The two successive extraction attacks and stealing+fine-tuning are performed in the same conditions as the first attack. The only difference is that A\mathcal{A} uses half the seed samples for each PRADA attack (5 per class for MNIST-5L and GTSRB-5L, 500 per class for CIFAR10-9L) and runs an additional duplication round to query the same number of inputs. The number of seed samples is a limited adversarial capability in PRADA (Juuti et al., 2019), so we grant A\mathcal{A} with the same capability for single and double extraction attack. For each KnockOff attack, A\mathcal{A} uses a different set of 100,000 inputs from ImageNet. For the second extraction attack (against FAF_{\mathcal{A}}), we query the same number of inputs as for the first one. While this number can be increased, we empirically observed that the test accuracy of FA′F_{\mathcal{A}}^{\prime} reaches its maximum and stagnates before the PRADA and KnockOff attacks finish, i.e., more queries do not improve AcctestAcc_{test} of FA′F_{\mathcal{A}}^{\prime}.

As it can be observed in Tab. 7 and 8, double extraction attack and stealing+fine-tuning effectively remove the watermark from the second order surrogate model FA′F_{\mathcal{A}}^{\prime}. The watermark accuracy is low enough (3-22%) to fail demonstration of ownership for FA′F_{\mathcal{A}}^{\prime}, which empirically confirms that prior DNN watermarking techniques are not resilient to this class of model extraction attacks (Zhang et al., 2018).

While removing the watermark, the extraction of the second order surrogate model using these attacks also increases the degradation in test accuracy by 20% to 80% for FA′F_{\mathcal{A}}^{\prime} compared to FAF_{\mathcal{A}}. Considering a powerful A\mathcal{A} having unlimited access to natural data (e.g., KnockOff adversary model) the extraction of the first order surrogate model incurs little accuracy degradation and so does the extraction of the second order surrogate model. The final model FA′F_{\mathcal{A}}^{\prime} stolen using KnockOff attack has its watermark removed and preserves its utility (from -1pp to -10pp compared to FVF_{\mathcal{V}}). DAWN cannot protect model extraction attacks where A\mathcal{A} has unlimited access to natural data. However, A\mathcal{A}’s access to data is limited in many scenarios, e.g., access to medical imaging that are privacy sensitive, and highly specialized models may not return meaningful predictions to random images. In such a scenario, the KnockOff attack may not be effective and the PRADA attack is more effective. We see that for the PRADA attack the test accuracy of FA′F_{\mathcal{A}}^{\prime} decreases sharply during each attack that removes the watermark (3 top rows in Tab. 7 and 8) . In most cases, the final test accuracy of FA′F_{\mathcal{A}}^{\prime} is less than half of FVF_{\mathcal{V}} (from -35pp to -62pp) and we consider that Acc(FA′)≪Acc(FV)Acc(F_{\mathcal{A}}^{\prime})\ll Acc(F_{\mathcal{V}}) makes FA′F_{\mathcal{A}}^{\prime} too inaccurate to be useful. DAWN can effectively protect a model against extraction attack that uses a limited amount of data - it destroys the utility of the model FA′F_{\mathcal{A}}^{\prime} deprived of the watermark.

Double-extraction and fine-tuning can remove the watermark while preserving test accuracy given that A\mathcal{A} has unlimited access to natural data. However, DAWN defeats both attacks when A\mathcal{A} has limited access to data, which is the case for most model extraction attack scenarios (cf. Sect. 2.1). A\mathcal{A}’s access to data is limited in many scenarios, e.g., medical imaging classifiers, where these removal attacks would be ineffective.

Pruning. Alternatively, A\mathcal{A} may attempt to prune the model by setting random weights of the model to zero. In our experiments we prune weights uniformly randomly. We show that for large values of δ\delta pruning is effective at removing the watermark but it sacrifices model’s utility and renders it useless (c.f. Table 9). Furthermore, we observe that AcctestAcc_{test} and AccwmAcc_{wm} do not fall proportionally i.e. there is no guarantee that sacrificing X% AcctestAcc_{test} results in the same drop in AccwmAcc_{wm}. Also, our experiments show that determining appropriate δ\delta without knowing TAT_{\mathcal{A}} is challenging as there is no consistent drop in accuracy for a particular value of δ\delta across all models.

Training with noise. A\mathcal{A} may attempt to weaken the embedding of the watermark by adding some noise to the samples before they start training. A\mathcal{A}’s goal is not to expose the model to the samples that would be eventually used for the verification. We show that for the values up to ϵ=0.4\epsilon=0.4, AccwmAcc_{wm} is not affected (c.f. Table 10). Beyond that in all but one case, either AccwmAcc_{wm} remains above the effectiveness threshold, or the drop in AcctestAcc_{test} becomes unacceptable. This approach is thus not effective at successfully stealing the model while circumventing DAWN.

Inference with noise. Instead of training with noisy samples, A\mathcal{A} can add noise to all samples during the inference in attempt to avoid verification. However, this will reduce utility for A\mathcal{A}’s clients. In Tab. 11, we show the decrease in AccwmAcc_{wm} and corresponding AcctestAcc_{test} for various amounts of noise ϵ\epsilon. We show that in almost all cases AccwmAcc_{wm} remains high or AcctestAcc_{test} drops below the acceptable utility level (purple, underline). In few cases (red, dashed underline) AcctestAcc_{test} remains high while AccwmAcc_{wm} decreases to <50%<50\%. However, the ϵ\epsilon value that benefits A\mathcal{A} the most is not consistent across the models. Hence, for a particular value of ϵ\epsilon that preserves AcctestAcc_{test}, A\mathcal{A} has no guarantee that watermark verification would fail.

Recognizing queries from training data. Alternatively, A\mathcal{A} having deployed FAF_{\mathcal{A}} and being aware of DAWN can try to prevent watermark verification performed by J\mathcal{J}. A\mathcal{A} can check if queries to FAF_{\mathcal{A}} belong to DAD_{\mathcal{A}} used to steal FVF_{\mathcal{V}}, and return different predictions for them. We evaluated that searching DAD_{\mathcal{A}} for exact matches can be done efficiently using a hash table (28-44ms additional overhead per query for our datasets). Hashing the query is the most time consuming part of this search.

To prevent this evasion by A\mathcal{A}, J\mathcal{J} can slightly perturb samples in TAT_{\mathcal{A}} before submitting them for verification to FAF_{\mathcal{A}}. In Tab. 11, we show the resilience of watermark verification to various amounts of noise (ϵ\epsilon) added to each image in TAT_{\mathcal{A}}. We show that we can maintain Accwm>50%Acc_{wm}>50\% for all models up to ϵ=0.25\epsilon=0.25 or even more for some models. A\mathcal{A} can no longer perform a simple lookup in a hash table to identify queries from J\mathcal{J} if TAT_{\mathcal{A}} is perturbed. A\mathcal{A} must compute the distance to every sample in DAD_{\mathcal{A}} and find the nearest neighbor to the query. We demonstrate that such search incurs a substantial computational overhead (c.f. Tab. 12). Searching in 100,000 ImageNet samples can take over 7s per query (on a server-grade machine with a Xeon CPU and 64 GB of RAM), which is too long to be acceptable for deployment. Also, a small distance to an element in DAD_{\mathcal{A}} does not mean that queried image was in fact part of DAD_{\mathcal{A}}. A\mathcal{A} has to set up a threshold for rejecting the queries that will affect the utility of the model, on top of the described computational overhead.

Discussion

Unremovability W1. We extensively evaluated (Sect 6.1) that manipulation of the training process of FAF_{\mathcal{A}} either does not prevent the embedding of the watermark or if it does, it significantly degrades FAF_{\mathcal{A}}’s test accuracy. Proper use of regularization can effectively mitigate the watermark embedding but it requires A\mathcal{A} to be granted more capabilities e.g. increased access to relevant data. We also showed (Sect. 8) that two successive extraction attacks can remove a watermark from FAF_{\mathcal{A}}. However, it also decreases the test accuracy to an extent that makes FAF_{\mathcal{A}} unusable. Finally, prior work (Adi et al., 2018; Merrer et al., 2017; Zhang et al., 2018) has shown that manipulations after training such as pruning and adversarial fine-tuning are ineffective against backdoor-based DNN watermarks. Our evaluation confirmed that watermark cannot be removed using pruning or fine-tuning (Sect. 8) without sacrificing the utility of the model. We can conclude that DAWN watermarking meets unremovability requirement.

Indistinguishability X2. We defined model-specific watermarking and backdoor functions (Sect. 4.1) that always return the same same output (correct or incorrect) for the same input. We also introduced a solution for mapping inputs with minor differences to similar predictions (Sect 4.1.3).

Reliability W2 and utility X1. The watermark registration and verification protocol that we introduce (Sect. 4.4) ensures that the success of A\mathcal{A} in demonstrating ownership of an arbitrary model is negligible. We showed how to set up DAWN in order to reliably demonstrate ownership of several surrogate models stolen using two state-of-the-art model extraction attacks with high confidence equal to 1−2−641-2^{-64} (Sect. 7). DAWN effectively watermarked every surrogate model FAF_{\mathcal{A}} while causing a negligible decrease of FVF_{\mathcal{V}}’s utility (0.03-0.5%). DAWN allows for reliable ownership demonstration while preserving the utility.

Non-ownership piracy W3 is guaranteed by our watermark registration and verification protocol (Sect. 4.4). In case of contention, the first registered model is deemed the original.

Linkablity W4. DAWN selects watermarked inputs from API client queries and registers one watermark per API client. Different clients make different queries and they will consequently have different watermarks. Given that we meet the reliability requirement W2, a single watermark (TAi,BV(TAi))(T_{\mathcal{A}_{i}},B_{\mathcal{V}}(T_{\mathcal{A}_{i}})) will succeed in proving FAiF_{\mathcal{A}_{i}} is a surrogate of FVF_{\mathcal{V}}. Watermarks are API client-specific which makes a surrogate model linkable to an API client.

Collusion resistance X3. DAWN relies on deterministic functions for watermarking (WVW_{\mathcal{V}}) and backdooring (BVB_{\mathcal{V}}) that are specific to FVF_{\mathcal{V}} but independent of the client sending a query. Thus, the watermark remains indistinguishable despite collusion X2. DAWN is resilient to a Sybil attack (bounded to a certain number of Sybils) and assure successful verification by J\mathcal{J} (Sect. 7.3).

2. Limitations

A\mathcal{A} can attempt to prevent ownership demonstration for FAF_{\mathcal{A}} by ensuring that watermark verification (Eq. 2) fails. This entails reducing the watermark accuracy by training FAF_{\mathcal{A}} using only a subset of the trigger set TAT_{\mathcal{A}}. Since watermarked inputs are indistinguishable (X2) and uniformly distributed in DAD_{\mathcal{A}}, A\mathcal{A} cannot selectively discard them. Nevertheless, A\mathcal{A} can discard x%x\% of the whole DAD_{\mathcal{A}}, which statistically, will result in x%x\% of watermarked input being discarded. This should reduce Accwm(FA)Acc_{wm}(F_{\mathcal{A}}) by 100−x%100-x\%. If xx is high enough, the resulting Accwm(FA)Acc_{wm}(F_{\mathcal{A}}) can be brought down low enough for watermark verification to systematically fail.

While this strategy is effective, it deprives A\mathcal{A} from a large part of DAD_{\mathcal{A}}. This decreases Acctest(FA)Acc_{test}(F_{\mathcal{A}}) and consequently the utility of FAF_{\mathcal{A}} (Sec. 8). Alternatively, A\mathcal{A} must collect a set DAD_{\mathcal{A}} x%x\% larger and make x%x\% more queries to FVF_{\mathcal{V}} to compensate for later discarded training inputs. We already discussed in Sect 2.1 that access to relevant data is the main limitation for A\mathcal{A}. The secondary goal of A\mathcal{A} is to limit the number of queries to FVF_{\mathcal{V}} (cf. Sect. 3.1). This evasion strategy requires more adversarial capabilities (access to data) and it compromises one adversary goal (minimum number of queries). Consequently, even if effective, we do not consider it a realistic evasion strategy.

Another potential limitation of DAWN is circumvention of the mapping function MVM_{\mathcal{V}} (Sec. 4.1.3). If mapping is too aggressive, A\mathcal{A} may probe the input space and try to identify subspaces that are grouped together. However, this is not guaranteed to work because the behavior of the model on synthetic samples is undefined (Goodfellow et al., 2014). MVM_{\mathcal{V}} impacts only the watermarking decision WVW_{\mathcal{V}} and not the returned label - A\mathcal{A} cannot interact directly with the mapping function. On the other hand, if the tolerated modification δ\delta is too small, A\mathcal{A} might identify watermarked queries by submitting several samples with minor modifications and taking the majority vote of the label. We evaluated this attack in Sect. 6.2 showing that mappings from MVM_{\mathcal{V}} are as consistent as predictions from FVF_{\mathcal{V}}. Semantic-preserving modifications to image queries (e.g., translation, rotation, change in color intensity, etc,) could be used to improve this attack. However, by using an embedding from FVF_{\mathcal{V}} to implement MVM_{\mathcal{V}}, both functions should be as resilient to semantic-preserving modifications.

A\mathcal{A} can attempt to weaken the embedding of the watermark by adding a small amount of noise to its training samples before starting the training or to all queries during the inference time (Sect. 8). Although A\mathcal{A} cannot know the optimal ϵ\epsilon value that minimizes accuracy loss while rendering watermark verification ineffective, they can choose a loss budget and incur that loss fully (e.g. 10 pp in our examples) - that implies that in Table 11 A\mathcal{A} would have succeeded in two out of the six cases. How to strengthen DAWN against such an adversary that is ready to incur the maximal allowable accuracy loss is still an open problem.

Related Work

Watermarking DNN models. The first watermarking technique for DNNs (Uchida et al., 2017) explicitly embeds additional information into the weights of a DNN after it is trained. Verifying the watermark requires white-box access to the model in order to analyze the weights. A limitation of this approach is that the watermark can be easily removed by minimally retraining the watermarked model.

Alternative approaches (Merrer et al., 2017; Adi et al., 2018; Zhang et al., 2018; Darvish Rouhani et al., 2019; Jia et al., 2020) that are more robust have been proposed, where the watermark can be verified in a black-box setting. These are based on backdooring and they allow for watermark extraction using only a prediction API, as discussed in Sect. 2.2. These approaches use both a carefully selected trigger set and a specific training process chosen by the model owner. The first proposal for such approach (Merrer et al., 2017) consist in modifying the original model boundary using adversarial retraining (Madry et al., 2017) in order to make the model unique. The watermark is composed of synthetically generated adversarial samples (Goodfellow et al., 2014) that are close to the decision boundary. The impact of selecting a particular distribution for a watermark has been evaluated in (Zhang et al., 2018). It shows that selecting a trigger set from the same distribution as the training data (albeit with minor synthetic modifications) or from a different distribution, does not affect the accuracy of the model for its primary classification task or on its training time, while the watermark gets perfectly embedded. Finally, more formal foundations and theoretical guarantees for backdoor-based DNN watermarking have been provided in (Adi et al., 2018). This work empirically assesses that the removability of a DNN watermark is highly dependent on the training process of the watermarked model (training from scratch vs. re-training).

Defenses against model extraction. It was suggested that the distribution of queries made during an extraction attack is different from benign queries (Juuti et al., 2019). Hence, model extraction can be detected using density estimation methods, namely by assessing the ability for queries to fit a Gaussian distribution or not. However, this technique protects only against attacks using synthetic queries and is not effective against, e.g., the KnockOff attack. Other detection methods analyse subsequent queries close to the classes’ decision boundaries (Quiring et al., 2018; Zheng et al., 2019) or queries exploring abnormally large region of the input space (Kesarwani et al., 2018). Both methods are effective but detect only extraction attacks against decision trees. They are ineffective against complex models like DNNs. Altering predictions returned to API clients can mitigate model extraction attacks. Predictions can be restricted to classes (Tramèr et al., 2016) or adversarially modified to degrade the performance of the surrogate model (Lee et al., 2018; Orekondy et al., 2020). However, some extraction attacks (Juuti et al., 2019) circumvent such defenses because they remain effective using just prediction classes.

Prior defenses to model extraction are designed to protect only simple models (Quiring et al., 2018; Kesarwani et al., 2018) or to prevent only specific extraction attacks (Lee et al., 2018; Zheng et al., 2019). It is arguable if a generic defense would ever be effective at detecting/preventing model extraction. Consequently, with DAWN we take a different approach where we assume a surrogate model can be extracted. Then we propose a generic defense to identify surrogate DNN models that have been extracted from any victim model using any extraction attack.

References

Appendix A Datasets and Models

Table 13 presents the characteristics of the datasets we used in our experiments. These are divided into a training and a testing set. Images were resized to fit the corresponding model architectures used in prior work. Table 14 presents model architectures used for conducting experiments with low capacity models - the perfect-knowledge attacker in Sect. 6.1 and reproduction of the PRADA (Juuti et al., 2019) attack in Sect. 7.2.

Appendix B Detecting watermarked inputs

We assess if watermarked inputs can be identified such that A\mathcal{A} could remove them from the DAD_{\mathcal{A}} before training the surrogate model.

This defense consists in first training a DNN model with the whole training dataset. Then, training data is predicted using the trained model and we record the activations of the last hidden layer of the DNN model. These activations are projected to three dimensions using Independent Component Analysis (ICA) and clustered into two clusters using k-means. These clusters are expected to group benign training data and poisoned data (watermarked inputs) respectively. The intuition for this approach is that incorrectly labeled inputs (watermark) trigger different activations than correctly labeled inputs in the trained DNN model. The size and silhouette score (Rousseeuw, 1987) of the two clusters are analyzed to conclude (1) if there is backdoor in the model and (2) which training inputs compose the backdoor. According to authors, a low silhouette score (0.1/0.15) and a high difference in relative cluster size is expected if the model embeds a watermark. The smallest cluster should contain the watermarked inputs.

To evaluate this defense against DAWN, we trained two sets of DNN models, plain models using a correctly labeled training set only, and watermarked models each embedding a watermark of size ∣TA∣=250|T_{\mathcal{A}}|=250. We applied the watermark detection process discussed above on these models and report results in Tab. 15. Clustering is not able to isolate watermarked inputs into a single cluster; the main part of watermarked inputs belongs to large clusters. The recommendation (Chen et al., 2019) to discard small clusters from training would deprive DAD_{\mathcal{A}} from a large number of correctly labeled samples while a large part of the watermark would be preserved. Using this approach, the defense wrongly discards 26.4% of clean data from DAD_{\mathcal{A}} on average while detecting only 67 out of 250 watermarked samples (26.8% of TAT_{\mathcal{A}}). The silhouette score is not useful for detecting the watermark either since watermarked and plain models have close scores that are all above the recommended detection threshold (0.1/0.15) (Chen et al., 2019). Our plain models are detected as embedding a watermark using this defense. We conclude that this defense is ineffective at detecting watermarks generated by DAWN.

We assume the reason for this ineffectiveness is due to the nature of our watermark, selected from the same distribution as the training set. In contrast to prior DNN watermarking solutions (Merrer et al., 2017; Adi et al., 2018; Zhang et al., 2018), our watermarked inputs do not come from a single manifold distant from the training data manifold. Consequently the model does not learn a “single” activation that generalizes to the whole watermark but rather learns individual exceptions for each watermarked input. The activations of watermarked inputs are thus different from each other and they are scattered among the activations of the remaining training data (correctly labeled). This can be observed in Fig. 3 - left, where we see that watermarked inputs are scattered among correctly labeled inputs. This explains why generated clusters cannot isolate watermarked inputs from correctly labeled inputs (Fig. 3 - right).