Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models

Zijin Yang, Kai Zeng, Kejiang Chen, Han Fang, Weiming Zhang, Nenghai Yu

Introduction

Diffusion models signify a noteworthy leap forward in image generation. These well-trained diffusion models, especially commercial diffusion models like Stable Diffusion (SD) , Glide , and Muse AI , enable individuals with diverse backgrounds to create high-quality images effortlessly. However, this raises concerns about intellectual property and whether diffusion models will be stolen or resold twice.

On the other hand, the ease of generating realistic images raises concerns about potentially misleading content generation. For example, on May 23, 2023, a Twitter-verified user named Bloomberg Feed posted a tweet titled “Large explosion near the Pentagon complex in Washington DC-initial report,” along with a synthetic image. This tweet led to multiple authoritative media accounts sharing it, even causing a brief impact on the stock marketFake image of Pentagon explosion on Twitter. On October 30, 2023, White House issued an executive order on AI security, emphasizing the need to protect Americans from AI-enabled fraud and deception by establishing standards and best practices for detecting AI-generated content and authenticating official contentFACT SHEET: President Biden Issues Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence. The urgency of labeling generated content for copyright authentication and prevention of misuse is evident.

Watermarking is highlighted as a fundamental method for labeling generated content, as it embeds watermark information within the generated image, allowing for subsequent copyright authentication and the tracking of false content. Existing watermarking methods for the diffusion model can be divided into three categories, as shown in Fig. 1. Post-processing-based watermarks adjust robust image features to embed watermarks, thereby directly altering the image and degrading its quality. To mitigate this concern, recent research endeavors propose fine-tuning-based methods , which amalgamate the watermark embedding process with the image generation process. Intuitively, these methods need to modify model parameters, introducing supplementary computational overhead. Recently, Wen et al. proposed the latent-representation-based Tree-Ring watermark, which conveys information by adapting the latent representations to match specific patterns. However, it restricts the randomness of sampling, which impacts generative performance.

Through the above analysis, we can find that these methods compromise model performance to embed watermarks. In practical applications, model performance is paramount for both business interests and user experience. Substantial resource investment is often necessary to pursue enhanced model performance. This leads to a fundamental question: Can watermarks be embedded without compromising model performance?

We affirmatively address the question presented above. Succinctly, the generation process can be delineated into two key phases: latent representation sampling and decoding. Our goal is to align the distribution of the latent representation in watermarked images with that of the latent representation in normally generated images. By keeping the model unaltered, the distribution of watermarked images is naturally consistent with that of normally generated images, enabling the seamless embedding of watermarks without compromising model performance.

Building upon this insight, we propose a watermarking method named Gaussian Shading, designed to ensure no deterioration in model performance. The embedding process encompasses three primary elements: watermark diffuse, randomization, and distribution-preserving sampling. Watermark diffusion spreads the watermark information throughout the latent representation. During the generation process, the watermark information will be diffused to the whole semantics of the image, thus achieving excellent robustness. Watermark randomization and distribution-preserving sampling guarantee the congruity of the latent representation distribution with that of watermark-free latent representations, thereby achieving performance-lossless. In the extraction phase, the latent representations are acquired through Denoising Diffusion Implicit Model (DDIM) inversion , allowing for the retrieval of watermark information. Harnessing the extensive scope of the SD latent space, we can achieve a high-capacity watermark of 256 bits, surpassing prior methods.

To the best of our knowledge, ours is the first technique that tackles this challenging problem of performance-lossless watermarking for diffusion models, and we provide theoretical proof. Moreover, this technique leaves the architecture and parameters of SD unaltered, necessitating no supplementary training. It can seamlessly integrate as a plug-and-play module within the generation process. Model providers can easily replace watermarked models with non-watermarked ones without affecting usability experiences.

We conducted thorough experiments on SD. Under strong noise perturbation, the average true positive rate and bit accuracy can exceed 0.99 and 0.97, respectively, validating the superiority of Gaussian Shading in both detection and traceability tasks compared to prior methods. Additionally, experiments on visual quality and image-text similarity serve as indicators of performance preservation in our approach. Lastly, we deliberated on various watermark erasure attacks, affirming the steadfast performance of our watermark in the face of such adversities.

Related Work

2 Image Watermarking

Digital watermarking is an effective means to address copyright protection and content authentication by embedding copyright or traceable identification information within carrier data. Typically, the functionality of a watermark depends on its capacity. For example, a single-bit watermark can determine whether an image was generated by a particular diffusion model, i.e., copyright protection; a multi-bit watermark can further determine which user of the diffusion model generated the image, i.e., traceability.

Image watermarking is a method that employs images as carriers for watermarking. Initially, watermark embedding methods primarily focused on the spatial domain , but later, to enhance robustness, transform domain watermarking techniques were developed. In recent years, with the advancement of deep learning, researchers have turned their attention to neural networks , harnessing their powerful learning capabilities to develop watermarking techniques .

3 Image Watermarking for Diffusion Models

Existing Image watermarking methods for the diffusion model can be divided into three categories, as shown in Fig. 1. The image watermarking methods described in the previous section can be applied directly to the images generated by the diffusion model, which is called post-processing-based watermarks . These methods directly modify the image, thus degrading image quality. Recent research endeavors have amalgamated the watermark embedding process with the image generation process to mitigate this issue. Stable Signature fine-tunes the LDM decoder using a pre-trained watermark extractor, facilitating watermark extraction from images produced by the fine-tuned model. Zhao et al. and Liu et al. suggest fine-tuning the diffusion model to implant a backdoor as a watermark, enabling watermark extraction by triggering. These fine-tuning-based approaches enhance the quality of watermarked images but introduce supplementary computational overhead and modify model parameters. Furthermore, Wen et al. introduced the Tree-Ring Watermark, which conveys copyright information by adapting the frequency domain of latent representations to match specific patterns. This method achieves an imperceptible watermark. However, it directly disrupts the Gaussian distribution of noise, limiting the randomness of sampling and resulting in affecting model performance.

Methods

In this section, we provide an overview of the application scenarios and functionalities in Fig. 2. We then proceed to detail the embedding and extraction processes shown in Fig. 3. Finally, we present a mathematical proof of the performance-lossless characteristic of the watermark.

Scenarios. See Fig. 2, the scenario involves the operator Alice, the thief Carol, and two types of users Bob and Trudy.

Alice is responsible for training the model, deploying it on the platform, and providing the corresponding API for users, but she does not open-source the code or model weights. Carol does not use Alice’s services but steals images generated by her model, claiming ownership of the copyrights. Bob and Trudy, as community users, can utilize the API to generate and disseminate images. While Bob faithfully adheres to the community guidelines, Trudy aims to generate deep fake, and infringing content. To evade detection and traceability, Trudy can employ various data augmentation to modify illicit images.

Detection. This scenario satisfies the detection (copyright protection) requirement. Alice embeds a single-bit watermark into each generated image. The successful extraction of the watermark from an image serves as evidence of Alice’s rightful ownership of the copyright, while also indicating that the image is artificially generated (as opposed to natural images).

Traceability. This scenario fulfills the traceability requirement. Alice allocates a watermark to each user. By extracting the watermark from the illicit content, it enables tracing Trudy, through comparison with the watermark database. Traceability is a higher pursuit than detection and can also achieve copyright protection for different users.

Details of the statistical tests in both scenarios are shown in Supplementary Material.

2 Watermark Embedding

Watermark diffusion. The dimensions of the latent representations are given by c×h×w{c\times h\times w}, where each dimension can represent ll bits of the watermark. Therefore, the watermark capacity becomes l×c×h×w{l\times c\times h\times w} bits. To enhance the robustness of the watermark, we represent the watermark using 1fhw\frac{1}{f_{hw}} of the height and width, and 1fc\frac{1}{f_{c}} of the channel, and replicate the watermark fc⋅fhw2f_{c}\cdot f^{2}_{hw} times. Thus, the watermark ss with dimensions l×cfc×hfhw×wfhwl\times\frac{c}{f_{c}}\times\frac{h}{f_{hw}}\times\frac{w}{f_{hw}} is expanded into a diffused watermark sds^{d} with dimensions l×c×h×wl\times c\times h\times w. The actual watermark capacity is k=l×c×h×wfc⋅fhw2k=\frac{{l\times c\times h\times w}}{f_{c}\cdot f^{2}_{hw}} bits.

Watermark randomization. If we know the distribution of the diffused watermark sds^{d}, we can directly utilize distribution-preserving sampling to obtain the corresponding latent representations zTsz^{s}_{T}. However, in practical scenarios, its distribution is always unknown. Hence, we introduce a stream key KK to transform sds^{d} into a distribution-known randomized watermark mm through encryption. Considering the use of computationally secure stream cipher, such as ChaCha20 , mm follows a uniform distribution, i.e., mm is a random binary bit stream.

Distribution-preserving sampling driven by randomized watermark. When each dimension represents ll-bit randomized watermark mm, this ll bits can be regarded as an integer y∈[0,2l−1]y\in[0,2^{l}-1]. Since mm is a ciphertext, yy follows a discrete uniform distribution, i.e., p(y)=12lp(y)=\frac{1}{2^{l}} for y=0,1,2,…,2l−1y=0,1,2,\dots,2^{l}-1. Let f(x)f(x) denote the probability density function of the Gaussian distribution N(0,I)\mathcal{N}(0,I), and ppfppf denotes the quantile function. We divide f(x)f(x) into 2l2^{l} equal cumulative probability portions. When y=iy=i, the watermarked latent representation zTsz^{s}_{T} falls into the ii-th interval, which means zTsz^{s}_{T} should follow the conditional distribution:

The probability distribution of zTsz^{s}_{T} is given by:

Eq. 2 indicates that zTsz^{s}_{T} follows the same distribution as the randomly sampled latent representation zT∼N(0,I)z_{T}\sim\mathcal{N}(0,I). Next, we elaborate on how this sampling is implemented.

Let the cumulative distribution function of f(x)f(x) be denoted as cdfcdf. We can obtain the cumulative distribution function of Eq. 1 as follows,

Given y=iy=i, we aim to perform random sampling of zTsz^{s}_{T} within the interval [ppf(i2l),ppf(i+12l)][ppf(\frac{i}{2^{l}}),ppf(\frac{i+1}{2^{l}})]. The commonly used method is rejection sampling , which can be time-consuming as it requires repeated sampling until zTsz^{s}_{T} falls into the correct interval. Instead, we can utilize the cumulative probability density. When randomly sampling F(zTs∣y=i)F(z^{s}_{T}|y=i), the corresponding zTsz^{s}_{T} is naturally obtained through random sampling. Since F(zTs∣y=i)F(z^{s}_{T}|y=i) takes values in $,samplingfromitisequivalenttosamplingfromastandarduniformdistribution,denotedas, sampling from it is equivalent to sampling from a standard uniform distribution, denoted asu=F(z^{s}_{T}|y=i)\sim\mathcal{U}(0,1).ShiftthetermsofEq.3,andtakeintoaccountthat. Shift the terms of Eq. 3, and take into account thatcdfandandppf$ are inverse functions, we have

Eq. 4 represents the process of sampling the watermarked latent representation zTsz^{s}_{T} driven by the randomized watermark mm. To extract the watermark, its inverse map is

Image generation. After the sampling process, the watermark is embedded in the latent representation zTsz^{s}_{T}, and the subsequent generation process is no different from the regular generation process of SD. Here, we employ the DPMSolver for iterative denoising of zTsz^{s}_{T}. In addition to DPMSolver , other continuous-time samplers based on ordinary differential equation (ODE) solvers , such as DDIM , DEIS , PNDM , and UniPC , can be used too. After obtaining denoised z0sz^{s}_{0}, the watermarked image XsX^{s} is generated using the decoder D\mathcal{D}: Xs=D(z0s)X^{s}=\mathcal{D}(z^{s}_{0}).

3 Watermark Extraction

DDIM Inversion. Using the SD encoder E\mathcal{E}, we first restore X′sX^{\prime s} to the latent space z0′s=E(X′s)z^{\prime s}_{0}=\mathcal{E}(X^{\prime s}). Then, we introduce the DDIM inversion to estimate the additive noise. It can be considered that zT′s≈zTsz^{\prime s}_{T}\approx z^{s}_{T}. We also observe that although DDIM inversion is derived from DDIM, it can apply to other continuous-time samplers based on ODE solvers.

Watermark reduction from latent representations. After obtaining zT′sz^{\prime s}_{T}, according to the inverse transformation defined in Eq. 5, the tensor can be converted into a bit stream m′m^{\prime}. Subsequently, m′m^{\prime} is decrypted using KK to obtain s′ds^{\prime d}. Inverse diffusion of the watermark results in fc⋅fhw2f_{c}\cdot f^{2}_{hw} copies of the watermark. Similar to voting, if the bit is set to 1 in more than half of the copies, the corresponding watermark bit is set to 11; otherwise, it is set to . This process restores the true binary watermark sequence s′s^{\prime}.

4 Proof of Lossless Performance

In prior works, the incorporation of watermark embedding modules inevitably results in a decline in model performance, as typically evaluated using metrics such as Peak Signal-to-Noise Ratio (PSNR) and Fréchet Inception Distance (FID) , which are more suitable for assessing post-processing methods. To assess methods that integrate the watermark embedding and generation processes, we propose a definition for the impact of watermark embedding on model performance, drawing on the complexity-theoretic definition of steganographic security . This definition is based on a probabilistic game between a watermarked image XsX^{s} and a normally generated image XX. The tester A\mathcal{A} can use any watermark to drive the sampling process and generate XsX^{s}, similar to the chosen hidden text attacks , which we refer to as chosen watermark tests. The watermarking method is performance-lossless under chosen watermark tests, if for any polynomial-time tester A\mathcal{A} and key K←KeyGenG(1ρ)K\leftarrow\textsf{KeyGen}_{\mathcal{G}(1^{\rho})}, it holds that

Here, ρ\rho represents the length of the security parameter, such as the key KK, and negl(ρ)\textsf{negl}(\rho) is a negligible term relative to ρ\rho.

We prove the statement using a proof by contradiction. First, assume that the watermarked image XsX^{s} and the normally generated image XX are distinguishable, meaning

where δ\delta is non-negligible with respect to the key KK. Let the iterative denoising process be denoted as Q(⋅)Q(\cdot), and substitute the LDM encoder E\mathcal{E} into Eq. 7, we have

where the randomized watermark mm is obtained by encrypting the diffused watermark sds^{d} using the encryption algorithm EE with key KK. Note that Eq. 2 contains the fact that distribution-preserving sampling driven by randomized watermark and random sampling are equivalent. Therefore, we denote sequence-driven sampling as S(⋅)S(\cdot). zTsz_{T}^{s} can naturally be obtained by sampling driven by mm, i.e., zTs=S(m)z^{s}_{T}=S(m). On the other hand, zTz_{T} can be considered as obtained by sampling driven by a truly random sequence rr of the same length as mm, i.e., zT=S(r)z_{T}=S(r). Eq. 8 can be written as

Sampling S(⋅)S(\cdot), denoising Q(⋅)Q(\cdot), and encoder E\mathcal{E} can be considered as subroutines that the tester AE,Q,S\mathcal{A}_{\mathcal{E},Q,S} can use. Thus, Eq. 9 can be simplified,

Note that S(⋅)S(\cdot), Q(⋅)Q(\cdot), and E\mathcal{E} are all polynomial-time programs, so the time taken by the tester AE,Q,S\mathcal{A}_{\mathcal{E},Q,S} to make the distinction is also polynomial. Eq. 10 essentially states that it is possible to distinguish between mm and rr in polynomial time. However, we have used the computationally secure stream cipher ChaCha20 in watermark randomization, which means that mm as a pseudorandom sequence cannot be distinguished from a truly random sequence in polynomial time. Eq. 10 contradicts the computational security property of ChaCha20 . Therefore, Eq. 10 is not valid, leading us back to our initial assumption that Eq. 7 is also not valid. This implies that the watermarked image XsX^{s} and the normally generated image XX are indistinguishable in polynomial time. Hence, Gaussian Shading is performance-lossless under chosen watermark tests.

Experiments

This section focuses on experimental analysis, including details of the experimental setup, performance evaluation of Gaussian Shading, comparison with baseline methods, ablation experiments, and potential attacks.

SD models. In this paper, we focus on text-to-image LDM, hence we select SD provided by huggingface. We evaluate Gaussian Shading as well as baseline methods, using three versions of SD: V1.4, V2.0, and V2.1. The size of the generated images is 512×512512\times 512, and the latent space dimension is 4×64×644\times 64\times 64. During inference, we employ the prompt from Stable-Diffusion-Prompt Stable-Diffusion-Prompts, with a guidance scale of 7.5. We sample 50 steps using DPMSolver . Considering that users tend to propagate the generated images without retaining the corresponding prompts, we use an empty prompt for inversion, with a scale of 1. We perform 50 steps of inversion using DDIM inversion .

Watermarking methods. In the main experiments, the settings for Gaussian Shading are fc=1,fhw=8,l=1f_{c}=1,f_{hw}=8,l=1, resulting in an actual capacity of 256 bits. We select five baseline methods: three officially used by SD, namely DwtDct , DwtDctSvd , and RivaGAN , a multi-bit watermarking called Stable Signature , and a train-free invisible watermarking called Tree-Ring .

Robustness evaluation To evaluate the robustness, we select nine representative types of noise shown in Fig. 4. We conduct experiments following the noise strength in Fig. 4.

Evaluation metrics. In the detection scenario, we calculate the true positive rate (TPR) corresponding to a fixed false positive rate (FPR). In the traceability scenario, we calculate the bit accuracy. To measure the bias in model performance, we compute the FID and CLIP-Score for 10 batches of watermarked images and perform a tt-test on the mean FID and CLIP-Score compared to that of watermark-free images.

All experiments are conducted using the PyTorch 1.13.0 framework, running on a single RTX 3090 GPU.

2 Performance of Gaussian Shading

Detection. In the detection scenario, we consider Gaussian Shading as a single-bit watermark, with a fixed watermark ss. We approximate the FPR to be controlled at 100,10−1,…,10−1310^{0},10^{-1},\dots,10^{-13}, calculate the corresponding threshold τ\tau , and test the TPR on 1,0001,000 watermarked images To mitigate the effects of randomness, we perform 5 trials with different ss and compute the average TPR. See Fig. 5(a), when the FPR is controlled at 10−1310^{-13}, the TPR remains at least 0.990.99 for eight out of the nine cases. Although the TPR for Brightness is only 0.9530.953, it is still a promising result.

Traceability. In this scenario, Gaussian Shading serves as a multi-bit watermark. Assuming Alice provides services to NN users, Alice needs to allocate one watermark for each user. In our experiments, we assume that N′=1,000N^{\prime}=1,000 users generate images, with each user generating 1010 images, resulting in a dataset of 10,00010,000 watermarked images.

During testing, we calculate the threshold τ\tau to control the FPR at 10−610^{-6}. Note that when computing traceability accuracy, we need to consider two types of errors: false positives, where watermarked images are not detected, and traceability errors, where watermarked images are detected but attributed to the wrong user. Therefore, we first determine whether the image contains a watermark. If it does, we calculate the number of matching bits AccAcc with all NN users on the platform. The user with the highest AccAcc is considered the one who generated the image. Finally, we verify whether the correct user has been traced. When N>N′N>N^{\prime}, it can be assumed that some users have been assigned a watermark but have not generated any images.

See Fig. 5(b), when N=106N=10^{6}, Gaussian Shading exhibits almost perfect traceability in seven cases. Although the traceability accuracy for Brightness is only 95.47%95.47\%, if a user generates two images, the probability of successfully tracing him is still no less than 99%99\%.

3 Comparison to Baselines

In this section, we compare the performance of Gaussian Shading with baselines on SD V1.4, V2.0, and V2.1. We use our implementations for each method, see details in Supplementary Material.

We conduct tests on 1,0001,000 generated images for each method respectively. See Tab. 1. Gaussian Shading exhibits strong robustness and significantly outperforms baselines in both scenarios. In terms of bit accuracy, it surpasses the best-performing baseline by approximately 7%7\%. This can be attributed to the extensive diffusion of the watermark throughout the entire latent space, establishing a profound binding between the watermark and the image semantics.

To measure the performance bias introduced by the watermark embedding, we apply a tt-test to evaluate. The hypotheses are H0:μs=μ0,H1:μs≠μ0H_{0}:\mu_{s}=\mu_{0},H_{1}:\mu_{s}\neq\mu_{0}, where μs\mu_{s} and μ0\mu_{0} represent the average FID or CLIP-Score of multiple sets of watermarked and watermark-free images, respectively. A lower tt-value indicates a higher probability that H0H_{0} holds. If the tt-value is larger than a threshold, H0H_{0} is rejected, and model performance is considered to have been affected. See Tab. 1, Gaussian Shading achieves the smallest tt-value, which indirectly reflects its performance-lossless characteristic. For a detailed analysis of the tt-test, please refer to the Supplementary Material.

4 Ablation Studies

In this section, we conduct comprehensive ablation experiments on SD V2.1 to demonstrate hyperparameter selection. Unless specified, we generate 1,0001,000 images and test the TPR and the bit accuracy with a theoretical FPR of 10−610^{-6}.

Watermark capacity. The watermark capacity is determined by three parameters: channel diffusion factor fcf_{c}, height-width diffusion factor fhwf_{hw}, and embedding rate ll. See Tab. 2, to balance the capacity and robustness of Gaussian Shading, we chose fc=1f_{c}=1 and fhw=8f_{hw}=8. After fixing fcf_{c} and fhwf_{hw}, we vary ll to examine if it could enhance the capacity, and additional results can be found in Supplementary Material. Considering all factors, we determine that the optimal solution is fc=1f_{c}=1, fhw=8f_{hw}=8, and l=1l=1, resulting in a watermark capacity of 256256 bits.

Sampling methods. To validate the generalization, we select five commonly used sampling methods, all continuous-time samplers based on ODE solvers . See Tab. 3, all of them exhibit excellent performance with a bit accuracy of approximately 97%97\% against noises.

Impact of the inversion step. In practice, the inference step is often unknown, which introduces a mismatch with the inversion step. See Tab. 4, such mismatch introduces minimal loss in accuracy. Considering the high efficiency of existing samplers, the inference step generally does not exceed 5050. Therefore, we set the inversion step to 5050.

Guidance scales. Given diverse user preferences for image-prompt alignment, larger guidance scales ensure faithful adherence to prompts, while smaller scales grant the model greater creative freedom. In SD, the guidance scale is typically selected from the range of $.Hence,experimentscovertherangeof. Hence, experiments cover the range of2toto18.Fortheinversion,anemptypromptisusedforguidance,andtheguidancescaleisfixedat1,assumingunknowninformationduringextraction.InFig.6(a),thebitaccuracyofGaussianShadingsurpasses. For the inversion, an empty prompt is used for guidance, and the guidance scale is fixed at 1, assuming unknown information during extraction. In Fig. 6(a), the bit accuracy of Gaussian Shading surpasses99.9\%$, showing its reliability in real-world-like scenes.

Noise intensities. To further test the robustness, we conduct experiments using different intensities of noises. See Figs. 6(b), 6(c), 6(d), 6(f), 6(g), 6(h), 6(i), 6(e) and 6(j), for Random Crop and Gaussian Noise, performance declines significantly with higher intensities. However, for the other seven types of noise, even at high intensities, the bit accuracy remains approximately 80%80\%.

5 Attacks against Gaussian Shading

We consider two malicious attacks: compression attack, where the attacker employs a neural network to compress watermarked images, and inversion attack, assuming the attacker is aware of the watermark embedding method, enabling them to modify the image’s latent representations.

Compression attack. We utilize popular auto-encoders to compare Stable Signature (SS) with Gaussian Shading across various compression rates. Additionally, we assess the compression quality through the PSNR between the compressed and watermarked images. See Figs. 7(a) and 7(b), Gaussian Shading significantly outperforms Stable Signature. This is because Gaussian Shading diffuses the watermark across the entire semantic space of images, while Stable Signature relies solely on the image texture.

Inversion attack. Assuming the attacker is aware of the embedding method, a more effective approach to erasing is through inversion to obtain latent representations and subsequently modify them. We validate the robustness against such attacks. Importantly, our experiments assume the strongest attacker capability of using the same model as Alice for precise inversion. In real-world scenarios, where the watermark embedding is not publicly available, the attacker’s capabilities would be weaker.

Specifically, we perform inversion to obtain latent representations and randomly flip a certain rate of them. Using the flipped latent representations, we regenerate the images and extract the watermark. See Fig. 7(c). the watermark can still be reliably extracted when the flipping rate (FR) is less than 0.4. At high FRs, significant changes in images are observed. Although the watermark cannot be accurately extracted, we consider the image transformed into a different one, resulting in the content not intended to be protected.

From another perspective, the attacker can launch a forgery attack by performing inversion on an innocuous image from Bob and generating harmful content using a different prompt. See Fig. 7(c), when the FR is 0, Alice can accurately trace Bob based on the forgeries, enabling the attacker to successfully frame Bob. Therefore, protecting the model from leakage is crucial for operators.

Limitations

Despite extensive experimental validation of Gaussian Shading’s superior performance, our work still has certain limitations. Firstly, the usage scenarios are restricted due to the reliance on DDIM inversion , which necessitates the utilization of continuous-time samplers based on ODE solvers like DPMSolver . Secondly, Gaussian Shading employs stream ciphers, necessitating proper key usage and management on the deployment platform. Additionally, we assume that the model is not publicly accessible, and only operators can verify the watermark, providing a certain level of protection against white-box attacks and ensuring security. However, if a legitimate third party requires watermark verification, cooperation from the operators becomes necessary. Lastly, Gaussian Shading is vulnerable to forgery attacks, emphasizing the importance for operators to safeguard the model parameters.

Conclusion and Future Work

We propose Gaussian Shading, a provably performance-lossless watermarking applied to diffusion models. Compared to baseline methods, Gaussian Shading offers simplicity and effectiveness by making a simple modification in the sampling process of the initial latent representation. Extensive experiments validate the superior performance in both detection and traceability scenarios. To our knowledge, we are the first to propose and implement a performance-lossless approach in image watermarking.

Regarding future work, we will introduce more efficient inversion methods and include a wider range of sampling methods. Additionally, careful consideration should be given to counteracting forgery attacks.

Acknowledgement. This work was supported in part by the Natural Science Foundation of China under Grant U2336206, 62102386, 62072421, 62372423, and 62121002.

References

Details of Gaussian Shading

Detection. Alice embeds a single-bit watermark, represented by kk-bit binary watermark s∈{0,1}ks\in\{0,1\}^{k}, into each generated image using Gaussian Shading. This watermark serves as an identifier for her model. Assuming the watermark s′s^{\prime} is extracted from image XX, the detection test for the watermark can be represented by the number of matching bits between two watermark sequences, Acc(s,s′)Acc(s,s^{\prime}). When the threshold τ∈{0,…,k}\tau\in\{0,\dots,k\} is determined, if

it is deemed that XX contains the watermark.

In previous works , it is commonly assumed that the extracted watermark bits s1′,…,sk′s^{\prime}_{1},\dots,s^{\prime}_{k} from the vanilla images are independently and identically distributed, with si′s^{\prime}_{i} following a Bernoulli distribution with parameter 0.50.5. Thus, Acc(s,s′)Acc(s,s^{\prime}) follows a binomial distribution with parameters (k,0.5)(k,0.5). It is worth noting that if we extract from a vanilla image and decrypt it using a computationally secure stream key , the resulting diffused watermark s′ds^{\prime d} should be a pseudorandom bit stream, and the corresponding watermark s′s^{\prime} would also be pseudorandom. In other words, the bits s1′,…,sk′s^{\prime}_{1},\dots,s^{\prime}_{k} are independently and identically distributed, and each si′s^{\prime}_{i} follows a Bernoulli distribution with a parameter of 0.50.5. This aligns perfectly with the above assumption.

Once the distribution of Acc(s,s′)Acc(s,s^{\prime}) is determined, the false positive rate (FPR⁡\operatorname{FPR}) is defined as the probability that Acc(s,s′)Acc(s,s^{\prime}) of a vanilla image exceeds the threshold τ\tau. This probability can be further expressed using the regularized incomplete beta function Bx(a;b)B_{x}(a;b) ,

Traceability. To enable traceability, Alice needs to assign a watermark si∈{0,1}ks^{i}\in\{0,1\}^{k} to each user, where i=1,…,Ni=1,\dots,N and NN represents the number of users. During the traceability test, the bit matching count Acc(s1,s′),…,Acc(sN,s′)Acc(s^{1},s^{\prime}),\dots,Acc(s^{N},s^{\prime}) needs to be computed for all NN watermarks. If none of the NN tests exceed the threshold τ\tau, the image is considered not generated by Alice’s model. However, if at least one test passes, the image is deemed to be generated by Alice’s model, and the index with the maximum matching count is traced back to the corresponding user, i.e., argmax⁡i=1,…,NAcc(si,s′)\operatorname*{argmax}_{i=1,\dots,N}Acc(s^{i},s^{\prime}). When a threshold τ\tau is given, the FPR⁡\operatorname{FPR} can be expressed as follows ,

2 Details of Denoising and Inversion

Markov chains of diffusion models. DDPM proposed that the diffusion model consists of two Markov chains used for adding and removing noise. The forward chain is pre-designed to transform the data distribution q0(x0)q_{0}(x_{0}) into a simple Gaussian distribution qT(xT)≈N(xT∣0,σ2I)q_{T}(x_{T})\approx\mathcal{N}(x_{T}|0,\sigma^{2}I) over a time interval of TT. Here, σ>0\sigma>0, and the transition probability q(xt∣xt−1)q(x_{t}|x_{t-1}) is defined as N(xt;αtx0,(1−αt)I)\mathcal{N}(x_{t};\sqrt{\alpha_{t}}x_{0},(1-\alpha_{t})I), where αt\alpha_{t} is a predetermined hyperparameter. By virtue of the Markov property, we have

with βt=α‾t\beta_{t}=\sqrt{\overline{\alpha}_{t}}, σt2=1−α‾t\sigma^{2}_{t}=1-\overline{\alpha}_{t}, and α‾t=∏i=0tαi{\overline{\alpha}_{t}}=\prod_{i=0}^{t}\alpha_{i}.

The transition kernel of the reverse chain is learned by a neural network θ\theta and aims to generate data from a Gaussian distribution with the transition probability distribution defined as

For LDM , since the diffusion process occurs in the latent space Z\mathcal{Z}, Eq. 14 and Eq. 15 should be rewritten for the latent representations zz of LDM as follows:

Denoising method for Gaussian Shading. DPMSolver is a higher-order ODE solver , and in this paper, we employ its second-order version during image generation, whose denoising process is as follows,

where λt=λ(t)=log⁡(βtσt)\lambda_{t}=\lambda(t)=\log\left(\frac{\beta_{t}}{\sigma_{t}}\right), tλ(⋅)t_{\lambda}(\cdot) represents the inverse function of λt\lambda_{t}, ht−1=λt−1−λth_{t-1}=\lambda_{t-1}-\lambda_{t}, t=1,2,…,Tt=1,2,\dots,T, and cc indicates the prompt used for text-to-image generation.

Inversion method for Gaussian Shading. We note that in DDIM , Song et al. proposed an inversion method where they used the Euler method to solve the ODE and obtained an approximate solution for the inverse process:

According to Eq. 19 it is possible to estimate the noise to be added, which enables latent representation restoration.

Experimental Details and Additional Experiments

To test the actual FPR of Gaussian Shading, and to validate the accuracy of Eq. 12 and Eq. 13, we performed watermark extraction on 50,00050,000 vanilla images from the ImageNet2014 validation set. See Fig. 8, the theoretical and actual measured curves are very close, indicating that the theoretical thresholds derived from Eq. 12 and Eq. 13 can effectively guarantee the actual FPR.

2 Details of Comparison Experiments

Watermarking methods settings. To ensure a fair comparison, we set the watermark capacity to 256 bits for DwtDct and DwtDctSvd . As RivaGAN has a maximum capacity of only 32 bits, we retain this setting. The capacity and robustness of Stable Signature are determined by Hidden trained in the first stage. However, in our experiments, we find that Hidden with a capacity of 256 bits did not converge during training. Additionally, if there are too many types of noise in the noise layer, Hidden does not converge either. As an alternative, we use the open-source model of Stable Signature with a capacity of 48 bitsThe GitHub Repository for Stable Signature. During fine-tuning, we utilize 400 images from the ImageNet2014 validation set, with a batch size of 4 and 100 training steps. Tree-Ring is a single-bit watermark, and we only compare it in the detection scenario. Since its Rand mode is more closely aligned with the concept of performance-lossless, we adopt this setting.

The specific experimental results in both scenarios are shown in Tab. 5 and Tab. 6, respectively. In the detection scenario, the average TPR of Gaussian Shading remains above 0.995 in the presence of noise, surpassing the subpar performance of Tree-Ring by approximately 0.1. In the traceability scenario, the average bit accuracy of Gaussian Shading exceeds 97%97\% against noises, outperforming the second-best method, RivaGAN, by around 7%7\%. In both scenarios, Gaussian Shading exhibits superior performance compared to baseline methods.

The tt-test for model performance. To measure the performance bias introduced by the watermark embedding, we apply a tt-test to evaluate.

We first generate 50,000/10,000 images using SD V2.1 for each watermarking method, divided into 10 groups of 5,000/1,000 images each. We then calculate the FID /CLIP-Score for each group and compute the average value μs\mu_{s}. Similarly, we generate 50,000/10,000 watermark-free images using SD V2.1, test the FID/CLIP-Score for 10 groups, and calculate the average value μ0\mu_{0}. For the FID, we randomly select 5000 images from MS-COCO-2017 validation set and calculate the scores using the aforementioned groups. For the CLIP-Score, we utilize OpenCLIP-ViT-G to compute the image-text relevance.

If the model performance is maintained, then μs\mu_{s} and μ0\mu_{0} should be statistically close to each other. Therefore, the hypotheses are

The statistic t-vt\text{-}v is calculated as follows:

nsn_{s} and n0n_{0} represent the number of testing times, which are both set to 10 in the experiments, and SsS_{s} and S0S_{0} represent the standard deviations of the FID/CLIP-Score for watermarked and watermark-free images, respectively.

A lower tt-value indicates a higher probability that H0H_{0} holds. If the tt-value is larger than a threshold, H0H_{0} is rejected, and model performance is considered to have been affected. The significance level for the test is set to t-v0.05(ns+n0−2)=t-v0.05(18)≈2.101t\text{-}v_{0.05}(n_{s}+n_{0}-2)=t\text{-}v_{0.05}(18)\approx 2.101. In terms of the FID, the tt-values of the baseline methods, as depicted in Tab. 7, are all greater than the critical value t-v0.05(18)≈2.101t\text{-}v_{0.05}(18)\approx 2.101, except for Gaussian Shading. Regarding the CLIP-Score, Tree-Ring, Stable Signature, and Gaussian Shading all exhibit competitive results. Note that the CLIP-Score tends to measure the alignment between generated images and prompts, while the FID is solely used to assess image quality. In summary, these baseline methods demonstrate a noticeable impact on the model’s performance in a statistically significant manner. On the other hand, Gaussian Shading achieved the smallest tt-value, which indirectly confirms its performance-lossless characteristic.

3 Details of Ablation Studies

Watermark capacity. The watermark capacity is determined by three parameters: channel diffusion factor fcf_{c}, height-width diffusion factor fhwf_{hw}, and embedding rate ll. To investigate the impact of these hyperparameters on watermark performance, we first fix ll to find an optimal value for fcf_{c} and fhwf_{hw}. Experimental results are shown in Tab. 8. Subsequently, we fix fcf_{c} and fhwf_{hw} to search for the highest possible ll, and the corresponding experimental results are presented in Tab. 9.

Considering all factors, we determine that the optimal solution is fc=1f_{c}=1, fhw=8f_{hw}=8, and l=1l=1, resulting in a watermark capacity of 256256 bits.

Sampling methods. Experimental results about sampling methods under different noises are shown in Tab. 10, and all of them exhibit excellent performance with an average bit accuracy of approximately 97%97\% against noises.

4 Additional Visual Results

See Fig. 9 and Fig. 10, we present the visual results of different watermarking methods on prompts from the MS-COCO-2017 validation set. From the residual images in Fig. 9, it can be observed that DwtDct , DwtDctSvd , RivaGAN , and Stable Signature introduce noticeable watermark artifacts, leading to a degradation in model performance. As shown in Fig. 10, although Tree-Ring watermark is imperceptible, its embedding may directly impair the image quality. Additionally, it may also introduce changes in the object count and spatial relationships, causing inconsistency with the prompt. In the case of Gaussian Shading, as long as the latent representations where the watermark is mapped remain consistent with that of the original image, no changes occur in the generated image.

To further showcase the visual performance of Gaussian Shading, we present the visual results at multiple embedding rates ranging from 1 to 5 on prompts from Stable-Diffusion-Prompt Stable-Diffusion-Prompts. See Fig. 11, with the increase in watermark length, the model maintains a good generation quality. Moreover, the diversity and randomness of watermarked images indirectly reflect the performance-lossless characteristic of Gaussian Shading.