SSLGuard: A Watermarking Scheme for Self-supervised Learning Pre-trained Encoders
Tianshuo Cong, Xinlei He, Yang Zhang
Introduction
Deep learning, in particular supervised learning (SL), has gained tremendous success during the past decade, and the development of SL relies on a large amount of high-quality labeled data. However, high-quality data is often difficult to collect and the cost of labeling is expensive. Self-supervised learning (SSL) is proposed to resolve such restrictions by generating “labels” from the unlabeled dataset (called pre-training dataset) and uses the derived “labels” to pre-train an encoder which can output informative embeddings. SSL encoders have shown great promise in various downstream tasks. For instance, on the ImageNet dataset , Chen et al. show that, by using SimCLR pre-trained with ImageNet (unlabeled), the downstream classifier can achieve top-5 accuracy with only labels, which outperforms a supervised AlexNet but uses 100 fewer labels. He et al. show that SSL can surpass SL under 7 downstream tasks including segmentation and detection. Therefore, compared to the SL-based classifier which only suits a specific classification task, the SSL pre-trained encoder can achieve remarkable performance on different downstream tasks.
However, the data collection and training process of SSL encoders are also expensive as they benefit from larger datasets and more powerful computing devices. For example, the performance of MoCo pre-trained with the Instagram-1B dataset ( billion images) outperforms that of the encoder pre-trained with the ImageNet-1M dataset (1.28 million images), and SimCLR requires 32 TPU v3 cores to train a ResNet-50 due to the large batch size setting (i.e., 4096) . Therefore, the cost to train a powerful encoder by SSL is prohibitive for individuals, and the high-performance encoders are usually pre-trained by leading AI companies with sufficient computing resources and shared via cloud platforms for commercial usage, i.e., Encoder-as-a-Service (EaaS) . For instance, Clarifai provides image encoders for different downstream services. OpenAI provides access to GPT-3 which can be considered as a powerful encoder for a variety of natural language processing (NLP) downstream tasks, such as code generation, style transfer, etc.
Once deployed on the cloud platform, the encoders are not only accessible to legitimate users but also threatened by potential adversaries. As illustrated in Figure 1, for the legitimate user, the encoder is used to train a downstream classifier. On the other hand, an adversary may perform model stealing attacks which aim to learn a surrogate encoder that has similar functionality. Such attacks may not only compromise the intellectual property of the service provider but also serve as a stepping stone for further attacks such as membership inference attacks (MIA) (i.e., mount MIA offline by using surrogate encoders), backdoor attacks (i.e., publish another backdoored encoder), and adversarial attacks . The security and privacy of SSL encoders are threatened by these attacks, which call for effective defenses.
As one major technique to protect the machine learning model’s copyright, model watermarking inserts a secret pattern into the model. Then, the ownership can be claimed if a similar or the same pattern is successfully extracted from the model. Recent studies on model watermarking mainly focus on the classifier that targeted a specific task . However, watermarking SSL encoders may face several intrinsic challenges. First, model watermarking against the classifier usually needs to specify a target class, while the SSL encoder does not have such information. Second, downstream tasks for SSL encoders are flexible, which challenges the traditional model watermarking scheme that is only suitable for one specific downstream task. Therefore, a new watermarking scheme should be designed to overcome those challenges to protect the copyright of SSL encoders. To the best of our knowledge, this has been left largely unstudied.
Our Work. In this paper, we first quantify the copyright breaching threat against SSL encoders through the lens of model stealing attacks. Then, we introduce SSLGuard, the first watermarking scheme for the SSL encoders to protect their copyrights. Note that in this work, we consider image encoders only.
For model stealing attacks, we first assume that the adversary only has black-box access to the victim encoder. We then characterize the adversary’s background knowledge into two dimensions, i.e., the surrogate dataset and the surrogate encoder’s architecture. Regarding the surrogate dataset which is used to train the surrogate encoder, we consider the adversary may or may not know the victim encoder’s pre-training dataset. Regarding the surrogate encoder’s architecture, we first assume that it shares the same architecture as the victim encoder. Then, we relax this assumption and find that the effectiveness of model stealing attacks can even increase by leveraging a larger model architecture. We empirically show that the model stealing attacks achieve remarkable performance. For instance, given a ResNet-50 encoder pre-trained on ImageNet by SimCLR, the ResNet-101 surrogate encoder can achieve 0.944 accuracy on STL-10 while the accuracy for the victim encoder is 0.948. We also show that the cost of stealing an encoder is much smaller than pre-training it from scratch, e.g., pre-training a BYOL ResNet-50 encoder costs 72.49 (see Table 3 for the detailed comparison). Such observation emphasizes the underlying threat of jeopardizing the model owner’s intellectual property and the emergence of copyright protection.
To protect the copyright of SSL encoders, we propose a robust black-box watermarking scheme named SSLGuard. Concretely, the goal of SSLGuard is to inject a watermark based on a given secret vector into a clean SSL encoder. The output of SSLGuard contains a watermarked encoder and a key-tuple. To be specific, the key-tuple consists of the secret vector, a verification dataset, and a decoder. SSLGuard fine-tunes a clean encoder to a watermarked encoder which can keep the utility and map samples in the verification dataset to secret embeddings. We further introduce a decoder to transform these secret embeddings into the secret vector. For other encoders, the decoder only transforms the embeddings generated from the verification dataset into random vectors. Recent research has shown that if a watermarked model is stolen, its corresponding watermark usually vanishes . To remedy this situation, SSLGuard adopts a shadow dataset and a shadow encoder to locally simulate model stealing attacks. Meanwhile, SSLGuard optimizes a trigger that can be recognized by both the watermarked encoder and the shadow encoder. We later show in Section 5 that such a design can strongly preserve the watermark even in the surrogate encoder stolen by the adversary.
Empirical evaluations over 7 datasets (i.e., ImageNet, CIFAR-10, CIFAR-100, STL-10, GTSRB, MNIST, and FashionMNIST) and 3 encoder pre-training algorithms (i.e., SimCLR, MoCo v2, and BYOL) show that SSLGuard can successfully inject/extract the watermark to/from the SSL encoder without sacrificing its performance and is robust to model stealing attacks. Moreover, we consider various types of watermark removal attacks including input preprocessing (noising), output perturbing (noising and truncation), and model modification (overwriting, pruning, and fine-tuning). We empirically show that SSLGuard is still effective in such a scenario.
In summary, we make the following contributions:
We unveil that the SSL pre-trained encoders are highly vulnerable to model stealing attacks.
We propose SSLGuard, the first watermarking scheme against SSL pre-trained encoders, which can protect the intellectual property of published encoders.
Extensive evaluations show that SSLGuard is effective in injecting and extracting watermarks, and it is robust against model stealing and other watermark removal attacks such as input noising, output perturbing, overwriting, model pruning, and fine-tuning.
Background
Self-supervised learning is a rising AI paradigm that aims to train an encoder by a large scale of unlabeled data. A high-performance pre-trained encoder can be shared into the public platform as an upstream service. In downstream tasks, customers can use the embeddings output from the pre-trained encoder to train their classifiers with limited labeled data or even no data . One of the most remarkable self-supervised learning paradigms is contrastive learning . In general, encoders are pre-trained through contrastive losses which calculate the similarities of embeddings in a latent space. In this paper, we consider three representative contrastive learning algorithms, i.e., SimCLR , MoCo v2 , and BYOL .
SimCLR . SimCLR is a simple framework for contrastive learning. It consists of 4 components, including Data augmentation, Base encoder , Projection head and Contrastive loss function.
where denotes the cosine similarity function and denotes a temperature parameter. SimCLR jointly trains the base encoder and projection head by minimizing the final loss function:
where and are the indexes for each positive pair. Once the model is trained, SimCLR discards the projection head and keeps the base encoder only, which serves as the pre-trained encoder.
MoCo v2 . Momentum Contrast (MoCo) is a famous contrastive learning algorithm, and MoCo v2 is the modified version (using a projection head and more data augmentations).
MoCo points out that contrastive learning can be regarded as a dictionary lookup task. The “keys” in the dictionary are the embeddings output from the encoder. A “query” matches a key if they are encoded from the same image. MoCo aims to train an encoder that outputs similar embeddings for a query and its matching key, and dissimilar embeddings for others. The dictionary is desirable to be large and consistent, which contains rich negative images and helps to learn good embeddings. MoCo aims to build such a dictionary with a queue and momentum encoder.
MoCo contains two parts: query encoder and key encoder . Given a query sample , MoCo gets an encoded query . For other samples , MoCo builds a dictionary whose keys are , . The dictionary is a dynamic queue that keeps the current mini-batch encoded embeddings and discards the ones in the oldest mini-batch. The benefit of using a queue is decoupling the dictionary size from the mini-batch size, so the dictionary size can be set as a hyper-parameter. Assume is the key that matches, the loss function will be defined as:
Here is a temperature hyper-parameter. MoCo trains by minimizing contrastive loss and updates by gradient descent. However, it is difficult to update by back-propagation because of the queue, so is updated by moving-averaged as:
where denotes a momentum coefficient. Finally, we keep the as the final pre-trained encoder.
BYOL . Bootstrap Your Own Latent (BYOL) is a novel self-supervised learning algorithm. Different from previous methods, BYOL does not rely on negative pairs, and it has a more robust selection of image augmentations.
BYOL’s architecture consists of two neural networks: online networks and target networks. The online networks, with parameters , consist of an encoder , a projector and a predictor . The target networks are made up of an encoder and a projector . The two networks bootstrap the embeddings and learn from each other.
Given an input sample , BYOL produces two augmented views and by using image augmentations and , respectively. The online networks output a projection and target networks output a target projection . The online networks’ goal is to make the prediction similar to . Formally, the similarity can be defined as the following:
Conversely, BYOL feeds to the online networks and to the target networks separately and gets . The final loss function can be formulated as:
BYOL updates the weights of the online and target networks by:
where is the learning rate of the online networks. The target networks’ weight is updated in a weighted average way, and denotes the decay rate of the target encoder. Once the model is trained, we treat the online networks’ encoder as the pre-trained encoder.
2 Model Stealing Attacks
Model stealing attacks aim to steal the parameters or the functionality of the victim model. To achieve this goal, given a victim model , the adversary can issue a bunch of queries to the victim model and obtain the corresponding responses. Then the queries and responses serve as the inputs and “labels” to train the surrogate model, denoted as . Formally, given a query dataset , the adversary can train by
where is a similarity function.
Note that if the victim model is a classifier, the response can be the prediction probability of each class. If the victim model is an encoder, the response can be the embeddings. A successful model stealing attack may not only breach the intellectual property of the victim model but also serve as a springboard for further attacks such as MIA , backdoor attacks and adversarial attacks . Previous work has demonstrated that neural networks are vulnerable to model stealing attacks. In this paper, we concentrate on model stealing attacks on SSL encoders, which have not been studied yet.
3 DNNs Watermarking
Considering the cost of training deep neural networks (DNNs), DNNs watermarking algorithms have received wide attention as it is an effective method to protect the copyright of the DNNs. Watermarking is a traditional concept for media such as audio and video, and it has been extended to protect the intellectual property of machine learning models recently . Concretely, the watermarking procedure can be divided into two steps, i.e., injection and verification. In the injection step, the model owner injects a watermark and a pre-defined behavior into the model in the training process. The watermark is usually secret, such as a trigger that is only known to the model owner . In the verification step, the ownership of a suspect model can be claimed if the watermarked encoder has the pre-defined behavior when the input samples contain the trigger.
So far, the watermarking algorithms mainly focus on the classifiers in a specific task. However, how to design a watermarking algorithm for SSL pre-trained encoders that can fit various downstream tasks remains largely unexplored.
Threat Model
In this paper, we consider two parties: the defender and the adversary. The defender is the owner of the victim encoder, whose goal is to protect the copyright of the victim encoder when publishing it as an online service. The adversary, on the contrary, aims to steal the victim encoder, i.e., by model stealing attacks or directly obtaining the model (insider threat), and bypass the copyright protection method for the victim encoder.
Adversary’s Motivation. Adversary’s motivation lies in two areas: Firstly, EaaS is being popular and high-performance SSL encoders are often pre-trained by top AI companies . Pre-training an encoder requires collecting a huge amount of data, expert knowledge for designing architectures/algorithms, and many failure trials, which are expensive. This makes the model architectures or training algorithms regarded as trade secrets and will not be publicly available, which makes it less possible for the adversary to directly train a comparable performance SSL encoder from scratch. Secondly, the cost of stealing an SSL encoder is quite less than training an SSL encoder from scratch. For instance, pre-training a ResNet-50 by BYOL needs 72.49 (see Table 3 for more details). Once the adversary steals the victim encoder successfully, they can resell it or deploy it on the cloud platform to be a commercial competitor.
Adversary’s Background Knowledge. For the adversary, we first assume that they only have black-box access to the victim encoder, which is the most challenging setting for the model stealing attacks . In this setting, the adversary can only query the victim encoder with data samples and obtain their corresponding responses, i.e., the embeddings, to train the surrogate encoders. We categorize the adversary’s background knowledge into two dimensions, i.e., the pre-training dataset and the victim encoder’s architecture. Concretely, we assume that the adversary has a query dataset to perform the attack. Note that the query dataset does not need to be in the same distribution as the victim encoder’s pre-training dataset. Regarding the victim encoder’s architecture, we first assume that the adversary can obtain it since such information is usually publicly accessible. Then we empirically show that this assumption can be relaxed, and the attack is even more effective when the adversary leverages a deeper model architecture.
Adaptive Adversary. We then consider an adaptive adversary who knows that the victim encoder has already been watermarked. This means they can leverage watermark removal techniques including input preprocessing (noising), output perturbing (noising and truncation), and model modification (overwriting, pruning, and fine-tuning) on the encoder to bypass the watermark verification.
Design of Watermarking Scheme
In this section, we present SSLGuard, a watermarking scheme to preserve the copyright of the SSL pre-trained encoders. SSLGuard should have the following properties:
Fidelity: To minimize the impact of SSLGuard on the legitimate users, the influence of SSLGuard on the clean pre-trained encoders should be negligible, which means SSLGuard should keep the utility of downstream tasks.
Effectiveness: SSLGuard should judge whether a suspect model is a watermarked (or a clean) model with high precision. In other words, SSLGuard should extract watermarks from watermarked encoders effectively.
Undetectability: The watermark cannot be extracted by a no-matching secret key-tuple. Undetectability ensures that ownership of the SSL pre-trained encoder could not be misrepresented.
Efficiency: SSLGuard should inject and extract watermark efficiently. For instance, the time cost for the watermark injection and extraction process should be less than pre-training an SSL model.
Robustness: SSLGuard should be robust against model stealing attacks and other watermark removal attacks such as input noising, output perturbing, overwriting, model pruning, and fine-tuning.
In the following subsections, we will introduce the design methods for SSLGuard. Table 1 summarizes the notations used in this paper.
The workflow of SSLGuard is shown in Figure 2. Concretely, given a clean encoder which is pre-trained by a certain SSL algorithm, SSLGuard will output a watermarked encoder and a secret key-tuple as:
The secret key-tuple consists of three items: a verification dataset , a decoder , and a secret vector . is an MLP that maps the embeddings generated from the encoder to a new latent space (same dimension as ) to calculate the cosine similarity with . Concretely, given an input image , the decoded vector can be defined as:
SSLGuard contains two processes, i.e., watermark injection and extraction. For the injection process, SSLGuard uses a secret key-tuple to inject the watermark into a clean encoder and outputs watermarked encoder as: The defender can release to the cloud platform and keep secret. For the extraction process, given a suspect encoder , the defender can use to extract decoded vectors from by: where is a set of decoded vectors. Then, the defender can measure the cosine similarity between and , and judge if a suspect encoder is a copy by:
here we adopt watermark rate (WR) as the metric to denote the ratio of the verified samples whose outputs are close to . Concretely, WR is defined as:
In summary, we need two thresholds here: and . is used to calculate WR, and is a threshold to verify the copyright. We set and by default. Note that the can be set to a smaller value as we show in Section 5 that the WR is 0 for the clean encoders. The overview of SSLGuard is depicted in Figure 3. Concretely, we first train a watermarked encoder that contains the information of the verification dataset and the secret vector. The clean encoder serves as a query-based API to guide the training process. The shadow encoder is used to simulate the model stealing process to better preserve the watermark under model stealing attacks. The watermarked encoder should keep the utility of the clean encoder while preserving the watermark injected in it.
2 Preparation
To watermark a pre-trained encoder, the defender should prepare a private dataset , a mask , and a random trigger . The mask is a binary matrix that contains the position information of trigger , which means and have the same size as the private samples . Following , we inject the trigger into by:
where denotes the element-wise product. Therefore, given the trigger , we can generate the verification dataset as:
Here we define three loss functions, i.e., correlated loss , uncorrelated loss , and embedding match loss to achieve three goals. Our first goal is to let the decoded vectors transformed from the verification dataset to be similar to the secret vector , and we define correlated loss function as:
where is a similarity function. If not otherwise specified, we use cosine similarity as the similarity function. The goal of is to train an encoder and an decoder together to transform into , where is correlated with . The more similar and are, the smaller will be.
Secondly, given a clean dataset and an encoder , the decoder transforms embeddings to the orthogonal direction of for uncorrelated samples . Therefore, we could get another loss function, uncorrelated loss function, as:
Finally, we here define an embedding match loss function to match the embeddings generated from two encoders and :
SSLGuard leverages to maintain the utility of the watermarked encoder and simulate the model stealing attacks.
3 Watermark Injection
As shown in Figure 3, SSLGuard adopts three encoders: a clean encoder , a watermarked encoder and a shadow encoder . Meanwhile, SSLGuard also uses three datasets: a target dataset , a shadow dataset , and a verification dataset . In the following part, we will introduce our loss functions for each module.
Shadow Encoder. For the shadow encoder, its task is to mimic the model stealing attacks. Here we use to simulate the query process. The loss function of the shadow encoder is:
Trigger and Decoder. Given a verification dataset, we aim to optimize a trigger and a decoder to extract from both the watermarked encoder and the shadow encoder, but not the clean encoder. The corresponding loss can be defined as:
Besides, for the clean encoder , watermarked encoder , and the shadow encoder , the decoder should not map the decoded keys closely to from the target dataset, the loss to achieve this goal can be defined as:
Given the above losses, the final loss function for trigger and decoder can be defined as:
Watermarked Encoder. For the watermarked encoder, we want it to keep the utility of the clean encoder. Therefore, for the samples from , we force the embeddings from and to become similar through . The loss can be defined as:
Meanwhile, the decoder should successfully extract from the verification dataset instead of the target dataset . The corresponding loss to achieve this goal is defined as:
The final loss function for the watermarked encoder is:
Optimization Problem. After designing all loss functions, we formulate SSLGuard as an optimization problem. Concretely, we update the parameters as follows:
where , , and are learning rates of shadow encoder, watermarked encoder, trigger, and decoder, respectively. We note that we update , , , and sequentially in one iteration, and we stop the optimization until the iteration reaches the max iteration number.
Evaluation
Datasets. We use the following 7 datasets to conduct our experiments.
ImageNet . The ImageNet dataset contains 1.2 million training images distributed in 1,000 classes. Each image has size .
CIFAR-10 The CIFAR-10 dataset has images in classes. Among them, there are images for training and images for testing. The size of each image is .
CIFAR-100 . Similar to CIFAR-10, The CIFAR-100 dataset contains images with size in classes, and there are 500 training images and 100 testing images in each class.
STL-10 . The STL-10 dataset consists of training images and testing images in 10 classes. Besides, it also contains unlabeled images. Note that the images on STL-10 are acquired from labeled images on ImageNet.https://cs.stanford.edu/~acoates/stl10/ The size of each image is .
GTSRB . German Traffic Sign Recognition Benchmark (GTSRB) contains training images and testing images. It contains -category traffic signs.
MNIST . MNIST is a handwritten digits dataset that contains 60,000 training images and 10,000 testing images in 10 classes. Each image has size .
FashionMNIST . FashionMNIST (F-MNIST) is a Zalando’s article image dataset that has 10 classes. It has 60,000 training images and 10,000 testing images. Each sample is a grayscale image with size .
We resize images of all datasets to in our experiments. We use ImageNet as the pre-training dataset; STL-10, CIFAR-10, F-MNIST, and MNIST as the downstream datasets; and STL-10, CIFAR-10, CIFAR-100, and GTSRB as the query dataset (to launch model stealing attacks). Note that for the STL-10 dataset, we randomly split the unlabeled samples (100,000) of it into two parts (each containing 50,000 samples). We consider the first part as the unlabeled STL-10 dataset and the second part as the same distribution unlabeled STL-10 dataset which is denoted as STL-10 (s).
Pre-trained Encoder. In our experiments, we adopt real-world contrastive learning pre-trained encoders as the victim encoders. Concretely, we download the checkpoints of the encoders from the official website (i.e., SimCLRhttps://github.com/google-research/simclr and MoCo v2https://github.com/facebookresearch/moco) or the public platform (i.e., BYOLhttps://github.com/yaox12/BYOL-PyTorch). All the encoders are ResNet-50 pre-trained on ImageNet.
Downstream Classifier. We use a -layer MLP as the downstream classifier with and neurons in its hidden layer. For each downstream task, we freeze the parameters of the pre-trained encoders and train the downstream classifier for epochs using Adam optimizer with learning rate.
SSLGuard. We reload the clean encoder and fine-tune it to be the watermarked encoder. Note that we freeze the weights in batch normalization layers following the settings by Jia et al. . We consider the unlabeled STL-10 dataset (with only 50,000 images as mentioned above) as both and , and adopt a ResNet-50 as the shadow encoder’s architecture. We sample 100 images from 5 random classes on ImageNet as our . Note that each class contains 20 images and the for watermarking SimCLR, MoCo v2, and BYOL are non-overlapping. For each sample in , space will be patterned by the trigger. We leverage the SGD optimizer with learning rate to train both the watermarked encoder and shadow encoder for 50 epochs. The batch size in our experiment is . The dimension of is . For the trigger, we randomly generate a tensor from a uniform distribution in $GG0.005$ learning rate to update both the decoder and the trigger.
2 Clean Downstream Accuracy
Given three clean SSL pre-trained encoders (i.e., pre-trained by SimCLR, MoCo v2, and BYOL on ImageNet), we first measure their downstream accuracy, denoted as clean downstream accuracy (CDA), for different tasks. We consider downstream classification tasks, i.e., STL-10, CIFAR-10, MNIST, and F-MNIST. The CDA are shown in Table 2. We observe that the SSL pre-trained encoders can achieve remarkable performance on different downstream tasks, which means the SSL pre-trained encoders can learn high-level semantic information from one task (i.e., ImageNet), and the informative embeddings can generalize to other tasks (i.e., STL-10 and CIFAR-10). Meanwhile, the cost of pre-training SSL encoders is expensive (see Table 3), such observation further demonstrates the necessity of protecting the copyright of the SSL pre-trained encoders. Note that we adopt CDA as our baseline accuracy. Later we measure an encoder’s performance by comparing its DA with CDA.
3 Model Stealing Attacks
Since the SSL pre-trained encoders (clean encoders) are powerful, we then evaluate whether they are vulnerable to model stealing attacks. To build a surrogate encoder, we consider three key information, i.e., the surrogate encoder’s architecture, the distribution of the query dataset, and the similarity function used to “copy” the victim encoder.
Surrogate Encoder’s Architecture. We first investigate the impact of the surrogate encoder’s architecture. Note that here we adopt the unlabeled STL-10 dataset (with 50,000 unlabeled samples) as the query dataset and cosine similarity as the similarity function to measure the difference between the victim and surrogate encoders’ embeddings. Since the architecture of the victim encoder can be non-public, the adversaries may try different surrogate encoder architectures to perform the model stealing attacks. Concretely, we assume the adversaries may leverage ResNet-18, ResNet-34, ResNet-50, or ResNet-101 as the surrogate encoder’s architecture. If the output dimension is different from ResNet-50 (the architecture of the victim encoder), e.g., ResNet-18/ResNet-34 outputs 512-dimensional embeddings, we leverage an extra linear layer to transform them into 2048-dimension. The DA of surrogate encoders is summarized in Figure 4. A general trend is that the deeper the surrogate encoder’s architecture, the better performance it can achieve on the downstream tasks. For instance, for SimCLR (4(a)), the DA on STL-10 and CIFAR-10 are 0.728 and 0.657 when the surrogate encoder’s architecture is ResNet-18, while the DA increases to 0.759 and 0.697 when the surrogate encoder’s architecture is changed to ResNet-50. This may be because a deeper model architecture can provide a wider parameter space and greater representation ability. Therefore, in general, deeper surrogate encoder’s architectures can better “copy” the functionality from victim encoders. Note that in the following experiments, the adversary uses ResNet-50 as the surrogate encoder’s architecture by default as it has comparable performance to ResNet-101 while requiring fewer resources.
Distribution of the Query Dataset. Secondly, we evaluate the impact of the query dataset’s distribution. In the real-world scenario, the adversary may or may not have the query dataset that is from the same distribution as the victim encoder’s pre-training dataset. Here the adversary leverages ResNet-50 as the surrogate model’s architecture and cosine similarity as the similarity function. Regarding the query dataset, the adversary may leverage the training dataset of CIFAR-10, CIFAR-100, and GTSRB and the unlabeled dataset of STL-10 to perform the attacks. The results are shown in Figure 5. First, we observe that the model stealing attack is more effective with querying by the same distribution dataset as the pre-training dataset. For instance, given the victim model trained by SimCLR (5(a)), when the downstream task is STL-10 classification, the DA for the surrogate encoders are 0.759, 0.646, 0.651, and 0.538 when the query dataset is STL-10, CIFAR-10, CIFAR-100, and GTSRB, respectively. This demonstrates that the same distribution query dataset can better steal the functionality of the victim encoder.
Another observation is that the distribution of the surrogate dataset may also influence DA on different tasks. For instance, given the victim model trained by BYOL (5(c)), when the downstream task is CIFAR-10 classification, the DA is 0.814 with CIFAR-10 as the query dataset, while only 0.769 with STL-10 as the query dataset. However, when the downstream task is STL-10 classification, the DA is 0.799 with CIFAR-10 as the query dataset but increases to 0.946 with STL-10 as the query dataset. Therefore, if the adversary is aware of the downstream task, they can construct a query dataset that is close to the downstream task to improve the stealing performance.
Similarity Function. Finally, we investigate the effect of similarity functions used in model stealing attacks. Besides cosine similarity, the adversary can also use mean absolute error (MAE) and mean square error (MSE) to match the victim encoder’s embeddings. Here we assume that the adversary leverages ResNet-50 as the surrogate model’s architecture and STL-10 as the query dataset. The results are shown in Figure 6. We can see that cosine similarity outperforms MAE and MSE in most settings. For instance, given the victim model trained by MoCo v2 (6(b)), the DA are all below 0.5 when using MAE and MSE. This can be credited to the normalization effect of cosine similarity, which helps to better learn the embeddings . This indicates that cosine similarity may better facilitate the stealing process.
Monetary Cost. We compare the monetary costs of pre-training an SSL encoder from scratch and stealing an SSL encoder. We first measure the training cost of the encoders. To pre-train a ResNet-50 encoder, SimCLR needs 60 hours with 32 TPU v3s, MoCo v2 uses 212 hours with 8 NVIDIA V100 GPUs, and BYOL takes 72 hours with 32 NVIDIA V100 GPUs (the training information is from the official or open-source implementation as mentioned in Section 5.1). The cost of model stealing contains two parts: querying the victim encoders and training the surrogate encoders locally. We use the GPU price from Google cloudhttps://cloud.google.com/compute/gpus-pricing to calculate the price for pre-training (i.e., We run our experiments on one NVIDIA A100 GPU whose price is 1 per 1,000 queries, from AWS.https://aws.amazon.com/rekognition/pricing We adopt the unlabeled STL-10 dataset (50,000 samples), cosine similarity, and different architectures to launch model stealing attacks. The monetary costs are shown in Table 3. We observe that the cost of stealing the pre-trained encoder is much smaller than pre-training it from scratch. For instance, pre-trains a BYOL ResNet-50 encoder takes \5,713.92\. This indicates that an adversary can “copy” the victim encoder with much less cost.
4 SSLGuard
In this section, we adopt SSLGuard to inject the watermarks into the clean encoders pre-trained by SimCLR, MoCo v2, and BYOL. We aim to validate four properties of SSLGuard, i.e., effectiveness, utility, undetectability, and efficiency. We will discuss the robustness of SSLGuard separately in Section 5.5.
Effectiveness. We first evaluate the effectiveness of SSLGuard. Concretely, we check whether the model owner can extract the watermark from the watermarked encoders. Ideally, the watermark should be successfully extracted from the watermarked encoder and shadow encoder , but not the clean encoder . We use the generated key-tuple to measure the watermark rate (WR) for , , and on three SSL algorithms. As shown in Table 4, the WR of and are all 1.00, which means encoder and both contain the information of and . Meanwhile, the WR of is 0.00. This means SSLGuard is generic and does not judge a clean encoder to be a watermarked encoder.
Fidelity. One of the initial intentions of SSLGuard is to maintain the utility of the original downstream task. To verify its fidelity, we first take BYOL as an example and visualize embeddings output from (the clean encoder pre-trained by BYOL) and using t-Distributed Neighbor Embedding (t-SNE) , which is depicted in Figure 7. We observe that the t-SNE results of and are almost identical and the embeddings are successfully separated by both encoders. This demonstrates that watermarked encoder trained by SSLGuard can faithfully reproduce the embeddings generated from the clean encoder. Also, we train downstream classifiers by using three watermarked encoders , and on STL-10, CIFAR-10, F-MNIST, and MNIST. Table 5 shows the DA in different scenarios. We observe that the DA of the watermarked encoders are almost the same as that of the clean encoders. For instance, compared to , the DA for only drops up to 0.009 from CDA. The evaluation shows that SSLGuard does not sacrifice the utility of the clean encoders.
Undetectability. We then check if the watermark can be extracted by a no-matching key-tuple. Through SSLGuard, we generate three key-tuples: , and . We use one of the key-tuples to verify other watermarked encoders, such as using to judge . As shown in Table 6, we see that the WR are all 0.00 in no-match pairs, which means we cannot use a non-matching to verify a watermarked encoder.
Efficiency. SSLGuard injects watermark into SimCLR, MoCo v2, and BYOL using 17.5hrs, 17.36hrs, and 10.70hrs, respectively, which are only 29.17%, 8.19%, and 14.86% of the time cost to pre-train SSL encoders, and the watermark extraction time is only 1.51s, 2.08s, and 1.82s, respectively. Note also that we use only a single GPU (A100) in the watermark injection process, which is much less than the requirement for pre-training the SSL encoders. This demonstrates that SSLGuard can inject and extract watermarks efficiently.
5 Robustness
We now quantify the robustness of SSLGuard. Concretely, we evaluate SSLGuard against model stealing and the following watermark removal attacks: Input preprocessing, output perturbing, and model modification. For instance, the adversary can add noise to the input samples or output embeddings. Also, the adversary can modify the parameters of the encoder by overwriting, pruning, and fine-tuning. Since watermark removal attacks may affect the performance of the encoders, and the adversary aims to "clean" the encoder but keep its functionality, we measure DA and WR simultaneously of these surrogate encoders. We note that the victim encoders are the watermarked encoders, and we leverage SimCLR, MoCo, and BYOL to denote , , and in this subsection. Regarding the downstream accuracy, we only show the results on BYOL (SimCLR and MoCo have similar trends).
Here we consider that the adversary may add i.i.d.Gaussian noise to each input image by . We evaluate DA on four downstream tasks and WR when we use different . The results of WR are shown in 8(a) and DA are shown in 9(a). We first observe that DA drops as increases. For instance, the DA on CIFAR-10 drops from 0.932 to 0.865 when increases from 0.05 to 0.15. On the other hand, the WR are all 1.00 for different on SimCLR, MoCo, and BYOL, respectively. This may be because when we inject the trigger into , the distribution of is too special, so our watermarked encoder can remember these special samples, which is robust to the input noising attacks.
5.2 Output Perturbing.
The adversary can also add some perturbations to the embeddings before returning them as the outputs. Here we consider two kinds of perturbations, i.e., random noising and truncation.
Output Noising. The adversary may return the perturbed embeddings by adding i.i.d.Gaussian noise as where is the original embedding, is the perturbed embedding, and is a hyper-parameter to control the noise level. Then, we evaluate DA and WR on different . From 9(b), we observe that DA decreases when increases. For instance, when increases from 0.05 to 0.15, DA on STL-10 drops from 0.940 to 0.905. However, the WR remains above 0.50 for all watermarked encoders (see 8(b)), which means when we feed the embeddings with noise into the decoders, the secret vector can still be successfully extracted. Therefore, the adversaries cannot remove the watermark even if they add random noise to the embeddings at the expense of decreasing the model’s performance.
Truncation. The adversary may decrease the precision of the embeddings by leveraging truncation. For instance, the adversary retains decimal places for each value in the embeddings, e.g., if , the adversary modifies the value to , and changes to when . 8(c) and 9(c) shows WR and DA under different . We observe that DA has a sharp drop when decreases from 1 to 0. Meanwhile, WR are all above 0.5 instead of MoCo, i.e., WR of MoCo drops to 0.00 when , but the DA on STL-10 is only 0.10. Therefore, adversaries cannot remove the watermark from the encoder while remaining its functionality.
5.3 Model Modification.
When adversaries have white-box access to the encoder, they can try to remove the watermark by modifying the encoder’s parameters. In this section, we consider three methods of model modification: watermark overwriting, model pruning, and fine-tuning.
Overwriting. The adversary can also leverage SSLGuard to inject a new watermark into an SSL encoder whether or not they know that the encoder has already been injected with a watermark. The adversary aims to generate a new watermarked encoder from with a different key-tuple. We want to confirm if our original watermark can remain in as well. For each , we measure the DA on different downstream tasks and the WR of the original key-tuple. The results are shown in Table 7. We observe that although we overwrite the watermarked encoder with a new key-tuple to generate a new encoder, the original watermark is still preserved, i.e., the WR of the original watermark in the new encoder is 1.00. This indicates that the original watermark can still be preserved even if the adversary overwrites a new watermark into the model.
Pruning. Pruning is an effective technology for model compression . It is also considered a watermark removal attack since many neurons may be disabled which reduces the effectiveness of the watermark . In this part, we leverage global and local unstructured pruning methods to the watermarked encoders. In the global pruning setting, we set fraction of weights in the convolutional layers which have the smallest absolute values in all layers to 0. Compared to global pruning, i.e., putting together all the connections across different layers and comparing them, local pruning aims to prune a proportion of connections with the smallest absolute values in the same layer. We show the WR and DA in the first two sub-figures of Figure 10 and Figure 11, respectively. We observe that DA and WR drop a little as the ratio increases in global pruning. However, for local pruning, there is a larger downward trend in DA. For instance, DA is 0.954 when and 0.871 when , this is because local pruning cannot preserve the global information in the model properly. In general, most of the WR are 1.0, which means SSLGuard is robust to different pruning settings. We also notice a special case here, i.e., on BYOL, when , the WR is 0.50. This is the worst case in our experiment, which demonstrates that we use watermark verification threshold in SSLGuard is reasonable. Also, note that for all clean encoders we evaluate in this paper, the WR is 0. This means the can be set to a smaller value to better verify the watermarked encoder as we discussed in Section 4.1.
Fine-tuning. After pruning, the adversary can fine-tune the surrogate encoders under the victim encoder’s supervision, which is following the setting in . This process is also called fine-pruning . The goal of fine-tuning is to regain DA’s drop. We fine-tune all the weights of the pruned encoders (global and local) by the MSE loss function. We note that we freeze the BatchNorm layers of the pruned encoders due to reducing inaccurate batch statistics estimation caused by a small batchsize . The WR are shown in 10(c) and 10(d), and the DA are shown in 11(c) and 11(d). We observe that fine-tuning can recover lost information from the victim encoder. For instance, when in the local pruned model, DA on STL-10 is 0.917. After fine-tuning the pruned model, DA comes to 0.954. Meanwhile, WR increases as DA recovers. This means SSLGuard is robust to fine-tuning.
5.4 Model Stealing.
We then quantify the robustness of SSLGuard through the lens of model stealing attacks. Note that we only consider the most powerful surrogate encoder’s architectures and most effective query datasets. Concretely, based on the evaluation in Section 5.2, we consider ResNet-50 and ResNet-101 as the surrogate encoder’s architectures and STL-10 as the query dataset. We name the three attacks Steal-1, Steal-2, and Steal-3. The details of each attack are shown in Table 8.
The WR and DA for different attacks are shown in Table 9. We observe that although the model stealing attack is effective against the watermarked encoder, we can still verify the ownership of the surrogate model as the WR is also high. For instance, for Steal-2 against the watermarked encoder pre-trained by BYOL, the DA is 0.937 and 0.815 on STL-10 and CIFAR-10, while the WR is 1.00, which indicates that the watermark injected by SSLGuard can still preserve in the surrogate encoder stolen by the adversary. We also have similar observations on Steal-1 and Steal-3, which demonstrate the robustness of SSLGuard under model stealing attacks.
Discussion
The Necessity of the Shadow Encoder. The reason why SSLGuard can extract watermarks from the surrogate encoder is that it locally simulates a model stealing process by using a shadow dataset and shadow encoder. In this part, we aim to demonstrate the need for such a design. We discard the shadow encoder and inject the watermark into a clean pre-trained encoder on SimCLR, MoCo v2, and BYOL. Then we get the corresponding key-tuples. The key-tuples can extract watermarks successfully. However, when We mount Steal-1 to the watermarked encoders to generate three surrogate encoders (i.e., , , and ), the WR are all 0.00, which means the watermark may not be verified. Meanwhile, DA for are 0.945, 0.735, 0.843, and 0.926 on STL-10, CIFAR-10, F-MNIST, and MNIST, respectively. This indicates that the adversary can successfully steal the victim encoder as the DA for the surrogate encoder are close to the target encoder. In conclusion, SSLGuard cannot work well without the shadow encoder as the adversary can steal a surrogate encoder with high utility while bypassing the watermark verification process. Therefore, the shadow encoder is crucial for defending against model stealing attacks.
The Choice of Mask. In our experiments, we set the covering space of the mask as . We also leverage different masks , i.e., and to inject watermark into BYOL, then we mount Steal-1 to the watermarked encoders, the WR are 0.99 and 1.00. The results show that the WR is similar when we leverage different covering spaces of the masks, which indicates that SSLGuard is effective under different masks.
Extension to Other Types of Datasets. In this paper, we only focus on encoders pre-trained on image datasets. To extend SSLGuard into encoders pre-trained on other types of datasets such as texts or graphs , the main challenge is to define a suitable trigger pattern in the language or graph domain. Then we can apply a similar method to watermark those models. We leave it as our future work to further explore the effectiveness of SSLGuard on other domains such as texts or graphs.
Related Work
Privacy and Security for SSL. There have been more and more studies on the privacy and security of self-supervised learning. Jia et al. sum up 10 security and privacy problems for SSL. Among them, only a small part has been studied. Liu et al. study MIA against contrastive learning-based pre-train encoder. Concretely, Liu et al. leverage data augmentations over the original samples to generate multiple augmented views. Then, the authors measure the similarities among the embeddings of the augmented samples. The intuition is that, if the sample is a member, then the similarities should be higher than a non-member. He and Zhang perform the first privacy analysis of contrastive learning. Concretely, the authors observe that the contrastive models are less vulnerable to membership inference attacks, while more vulnerable to attribute inference attacks. The reason is that contrastive models are more generalized with less overfitting level, which leads to fewer membership inference risks, but the representations learned by contrastive learning are more informative, thus leaking more attribute information. Jia et al. propose the first backdoor attack against SSL pre-trained encoders. By injecting the trigger pattern in the pre-training process of an encoder that correlated to a specific downstream task, the backdoored encoder can behave abnormally for this downstream task. The author further shows that triggers for multiple tasks can be simultaneously injected into the encoder.
DNNs Copyright Protection. In recent years, several techniques for DNNs copyright protection have been proposed. Among them, DNNs watermarking is one of the most representative algorithms. Jia et al. propose an entangled watermarking algorithm that encourages the classifiers to represent training data and watermarks similarly. The goal of the entanglement is to force the adversary to learn the knowledge of the watermarks when he steals the model. DNN fingerprinting is another protection method. Unlike watermarking, the goal of fingerprinting is to extract a specific property from the model. Cao et al. introduce a fingerprinting extraction algorithm, namely IPGuard. IPGuard regards the data points near the classification boundary as the model’s fingerprint. If a suspect classifier predicts the same labels for these points, then it will be judged as a surrogate classifier. Chen et al. propose a testing framework for supervised learning models. They propose six metrics to measure whether a suspect model is a copy of the victim model. Among these metrics, four of them need white-box access, and black-box access is enough for the rest.
Conclusion
In this paper, we first quantify the copyright breaching threats of SSL pre-trained encoders through the lens of model stealing attacks. We empirically show that the SSL pre-trained encoders are highly vulnerable to model stealing attacks. This is because the rich information in the embeddings can be leveraged to better capture the behavior of the victim encoder. To protect the copyright of the SSL pre-trained encoder, we propose SSLGuard, a robust black-box watermarking scheme for the SSL pre-trained encoders. Concretely, given a secret vector, SSLGuard injects a watermark into a clean pre-trained encoder and outputs a watermarked version. The shadow training technique is also applied to preserve the watermark under potential model stealing attacks. Extensive evaluations show that SSLGuard is effective in embedding and extracting watermarks and robust against model stealing and different types of watermark removal attacks such as input noising, output perturbing, overwriting, model pruning, and fine-tuning.
Acknowledgement
We thank all anonymous reviewers for their constructive comments. This work is partially funded by the Helmholtz Association within the project “Trustworthy Federated Data Analytics” (TFDA) (funding number ZT-I-OO1 4), by the National Key Research and Development Program of China (2018YFA0704701, 2020YFA0309705), by the Major Program of Guangdong Basic and Applied Research (2019B030302008), and by the Major Scientific and Technological Innovation Project of Shandong Province (2019JZZY010133).