An Undetectable Watermark for Generative Image Models

Sam Gunn, Xuandong Zhao, Dawn Song

Introduction

As AI-generated content grows increasingly realistic, so does the threat of AI-generated disinformation. AI-generated images have already appeared online in attempts to influence politics (Ryan-Mosley, 2023). Watermarking has the potential to mitigate this issue: If AI providers embed watermarks in generated content, then that content can be flagged using the watermarking key. Recognizing this, governments have begun putting pressure on companies to implement watermarks (Biden, 2023; California State Legislature, 2024; European Union, 2024). However, despite an abundance of available watermarking schemes in the literature, adoption has remained limited (Seetharaman & Barnum, 2023). There are at least a few potential explanations for this.

First, some clients are willing to pay a premium for un-watermarked content. For instance, a student using generative AI for class might find a watermark problematic. Any company implementing a watermark could therefore put itself at a competitive disadvantage.

Second, existing watermarking schemes noticeably degrade the quality of generated content. Some schemes guarantee that the distribution of a single response is unchanged, but introduce correlations across generations.One scheme with this guarantee is to fix the randomness of the sampling algorithm, so that the model can only generate one unique response for each prompt. While this might be acceptable for small models, it is questionable whether anyone would be willing to use such a watermark for a model that cost over $100 million just to train (Knight, 2023). In other words, given the vast effort put into optimizing models, any observable change in the model’s behavior is probably unacceptable.

Undetectability, originally defined in the context of watermarking by Christ et al. (2024), addresses both of these issues. For an undetectable watermark, it is computationally infeasible for anyone who doesn’t hold the detection key to distinguish generations with the watermark from generations without — even if one is allowed to make many adaptive queries. Crucially, an undetectable watermark provably preserves quality under any efficiently-computable metric, including quality metrics that are measured across many generations (such as FID (Heusel et al., 2017), CLIP (Radford et al., 2021) and Inception Score (Salimans et al., 2016) for images). Therefore one can confidently use an undetectable watermark without any concern that the quality might degrade. And if the detection key is kept sufficiently private, then the competitive disadvantage to using the watermark can be minimized: One can give the detection key only to mass distributors of information (like Meta and X), so that broad dissemination of AI-generated disinformation can be filtered without interfering in users’ personal affairs. Since only the mass information distributors would be able to detect the watermark, it would not harm the value of the content except to bad actors.

In this paper, we introduce the first undetectableSee Section C.3 for a discussion of the extent to which our scheme is cryptographically undetectable for various choices of parameters. watermarking scheme for image generation models. Our scheme works for latent diffusion models (Rombach et al., 2022), with which we generate watermarked images by progressively de-noising initial watermarked noise within the latent space. The key component in our scheme is a pseudorandom error-correcting code, or pseudorandom code (PRC), a cryptographic object introduced by Christ & Gunn (2024). We therefore refer to our scheme as the PRC watermark in this work.

At a high level, a PRC allows us to embed a cryptographically pseudorandom pattern that is robustly distributed across the entire latent space, ensuring that the watermark operates at a semantic level. The fact that our watermark is embedded at the semantic level, combined with the PRC’s error-correcting properties, makes our watermark highly robust — especially to pixel-level watermark removal attacks (Zhao et al., 2023).

Additionally, since the PRC from Christ & Gunn (2024) can be used to encode and decode messages, we can robustly embed large messages within the PRC watermark. While the decoder is somewhat less robust than the detector, the detector can still be effectively used in cases where the decoder fails.

Finally, the PRC watermark is highly flexible, requiring no additional model training or fine-tuning, and can be seamlessly incorporated into existing diffusion model APIs. It allows the user to independently set the message length and a desired upper bound on the false positive rate (FPR) at the time of watermark key generation. The false positive rate is rigorous, rather than empirical: If the user sets the desired upper bound on the false positive rate to FF, then we prove in Theorem 2 that the false positive rate will not exceed FF.

Experiments on quality and detectability are presented in Section 4.2. We emphasize that undetectability theoretically ensures quality preservation, and our scheme is undetectable by the results of Christ & Gunn (2024). Therefore we perform experiments on quality and detectability only to ensure that our scheme is secure enough with our finite choice of parameters.

We demonstrate the undetectability of our scheme in three key ways:

We show in Table 1 that the quality, as measured by the FID, CLIP, and Inception Score, are all preserved by the PRC watermark. This is in contrast to every other scheme we tested.

We show in Table 2 that the perceptual variability of responses, as measured by the LPIPS score (Zhang et al., 2019), is preserved under the PRC watermark. This is in contrast to every other comparable scheme we tested.

We show in Figure 2 that an image classifier fails to learn to detect the PRC watermark. The same image classifier quickly learns to detect every other scheme we tested.

We demonstrate the robustness of our scheme in Section 4.3. We find that watermark removal attacks fail to remove the PRC watermark without significantly degrading the quality of the image. We test the robustness under ten different types of watermark removal attacks with varying strengths and compare PRC watermark to eight different state-of-the-art watermarking schemes. Among the three watermarking schemes with the lowest impact on image quality,These are the DwtDct, DwtDctSvd, and PRC watermarks. the PRC watermark is the most robust to all attacks.

Finally, we show in Section 4.4 that the PRC watermark can be used to encode and decode long messages in generated images. The encoding algorithm is exactly the same, except that the user passes it a message, and the scheme remains heuristically undetectable. These messages could be used to encode, for instance, timestamps, user IDs, digital signatures, or model specifications. We find in Figure 11 that the robustness of the decoder for 512-bit messages is comparable to, although slightly less than, the robustness of the detector. For non-attacked images, we show in Figure 12 that we can increase the message capacity to at least 2500 bits.

Related work

There is a rich history of digital watermarking techniques, ranging from conventional steganography to modern methods based on generative models. Following the taxonomy in An et al. (2024), watermarking methods are categorized into two types: post-processing and in-processing schemes.

Post-processing schemes embed the watermark after image generation and have been used for decades due to their broad applicability.

In-processing schemes modify the generative model or sampling process to embed the watermark directly in the generated content.

Our PRC watermark falls under the in-processing category. Note that post-processing watermarks cannot be made undetectable without introducing extra modeling assumptions: One can always distinguish between a fixed image and any modification of it. We refer the reader to surveys (Cox et al., 2008; Wan et al., 2022; An et al., 2024) for more on post-processing methods. Below, we focus on two popular in-processing techniques: Tree-Ring and Gaussian Shading watermarks.

Wen et al. (2023) introduced Tree-Ring watermark, the first in-processing watermark that modifies the latent sampling distribution and employs an inverse diffusion process for detection. Our PRC watermark builds on this framework but adopts a different latent distribution. The Tree-Ring watermark works by fixing concentric rings in the Fourier domain of the latent space to be 0. To detect the watermark, one uses DDIM inversion (Song et al., 2021) to estimate the initial latent, and the watermark is considered present if the latent estimate has unusually small values in the watermarked rings. Follow-up works have extended this approach by refining the heuristic latent pattern in the watermarking process (Zhang et al., 2024; Ci et al., 2024). However, under Tree-Ring’s strategy, the initial latent significantly deviates from the Gaussian latent distribution, leading to reduced image quality and variability, as shown in Tables 1 and 2. Furthermore, the Tree-Ring watermark is a zero-bit scheme and cannot encode messages. While the Tree-Ring watermark is robust to several attacks, it is highly susceptible to the adversarial surrogate attack since the latent pattern is easy to learn with a neural network. In Figure 5, we find that this attack removes the Tree-Ring watermark with minimal effect on image quality.

The basic Gaussian Shading watermark (Yang et al., 2024) works by choosing a fixed quadrant of latent space as the watermarking key, and only generating images from latents in that quadrant. Detection involves recovering the latent and determining if it lies unusually close to the watermarked quadrant. In their paper, Yang et al. (2024) include a proof that Gaussian Shading has “lossless performance.” However, this proof only shows that the distribution of a single watermarked image is the same as that of a single un-watermarked image. Crucially, even standard quality metrics such as the FID (Heusel et al., 2017), CLIP Score (Radford et al., 2021), and Inception Score (Salimans et al., 2016) account for correlations between generated images, so their proof of lossless performance does not guarantee perfect quality under these metrics. Indeed, we find in Table 1 that the Gaussian Shading watermark significantly degrades the FID and Inception Score.Table 1 of Yang et al. (2024) appears to show that the FID against the COCO dataset is preserved under Gaussian Shading. However, from their code repository it appears that this table is generated by re-sampling the watermarking key for every generation. To be consistent with the intended use case, in this work we use the same random watermarking key to generate many images and compute the quality score. We expand on this further by measuring the “variability” of watermarked images. Since images under the Gaussian Shading watermark all come from the same quadrant in latent space, we expect that the variability should be reduced. We use the LPIPS perceptual similarity score (Zhang et al., 2018) to measure the diversity among different watermarked images for a fixed prompt. As shown in Table 2, the perceptual similarity between images is significantly higher with the Gaussian Shading watermark, confirming the diminished variability. The diminished variability turns out to be easily observable to the human eye — we show an example of this in Figure 10.

Undetectable watermarks were initially defined by Christ et al. (2024) in the context of language models. Subsequent to Christ & Gunn (2024), alternative constructions of PRCs have been given by Golowich & Moitra (2024) and Ghentiyala & Guruswami (2024). It would be interesting to see if these PRCs yield improved image watermarks, but we did not investigate this.

Method

We consider a setting where users make queries to a provider, and the provider responds to these queries with images produced by some image generation model. In watermarking, the provider holds a watermarking key that is used to sample from a modified, watermarked distribution over images. Anyone holding the watermarking key can, with high probability, distinguish between samples from the un-watermarked distribution and the watermarked distribution.

Since the watermark may be undesirable to some users, some of them may attempt to remove the watermark. We are therefore interested in robust watermarks, for which watermark detection still functions even when the image is subjected to a watermark removal attack. We assume that the adversary performing such a removal attack is restricted in two ways. First, the adversary should have weaker capabilities than the provider. If the adversary can generate their own images of equal quality to the provider, then they don’t need to engage in watermark removal attacks. Second, we are only interested in adversaries that produce high-quality images after the removal attack. If removal attacks require significantly degrading the quality of the image, then there is incentive to leave the watermark.

We are also interested in spoofing attacks, whereby an adversary who doesn’t know the watermarking key attempts to add a watermark to an un-watermarked image. We only perform limited experiments on spoofing attacks, so we do not discuss the adversarial capabilities here. However, we note that our techniques, together with the ideas on unforgeable public attribution from Christ & Gunn (2024), immediately yield a scheme that is provably resilient to spoofing attacks.

2 Overview of the PRC watermark

Some of the most popular generative image models today are latent diffusion models (Rombach et al., 2022), which consist of a diffusion model specified by a de-noising neural network ϵ\epsilon, a (possibly-randomized) function fϵf_{\epsilon} depending on ϵ\epsilon, a number of diffusion iterations TT, and an autoencoder (E⁡,D⁡)(\operatorname{\mathcal{E}},\operatorname{\mathcal{D}}). For a latent diffusion model, Generate\mathsf{Generate} works as follows.

In words, Generate\mathsf{Generate} works by starting with a normally distributed latent and iteratively de-noising it. The de-noised latent is then decoded by the autoencoder. In order to produce an image for the prompt π{\bm{\pi}} using Generate\mathsf{Generate}, we use Sample\mathsf{Sample} defined as follows.

There has been increasing interest in tracing the diffusion model generative process back (Recover\mathsf{Recover}). Diffusion inversion has been important for various applications such as image editing (Hertz et al., 2022) and style transfer (Zhang et al., 2023). A commonly used method for reversing the diffusion process is Denoising Diffusion Implicit Models (DDIM) (Song et al., 2021) inversion, which leverages the formulation of the denoising process in diffusion models as an ordinary differential equation (ODE). However, the result of DDIM inversion, z(T){\bm{z}}^{(T)}, is an approximation even when the input text is known. For our implementation of Generate\mathsf{Generate}, we employ Stable Diffusion with DPM-solvers (Lu et al., 2022) for sampling. In our implementation of Recover\mathsf{Recover}, we adopt the exact inversion method proposed in Hong et al. (2023) for more accurate inversion.

Our PRC consists of four algorithms, given in Appendix B:

PRC.KeyGen(n,F,t)\mathsf{PRC}.\mathsf{KeyGen}(n,F,t) samples a PRC key \key\key, which will also serve as the watermarking key. The parameter nn is the block length, which in our case is the dimension of the latent space; FF is the desired false positive rate; and tt is a parameter which may be increased for improved undetectability at the cost of robustness.

PRC.Encode\key\mathsf{PRC}.\mathsf{Encode}_{\key} samples a PRC codeword.

PRC.Detect\key(c)\mathsf{PRC}.\mathsf{Detect}_{\key}({\bm{c}}) tests whether the given string c{\bm{c}} came from the PRC.

PRC.Decode\key(c)\mathsf{PRC}.\mathsf{Decode}_{\key}({\bm{c}}) decodes the message from the given string c{\bm{c}}, if it exists. The decoder is slower and less robust than the detector.

As our PRC, we use the LDPC construction from Christ & Gunn (2024), modified to handle soft decisions. Essentially, this PRC works by sampling random tt-sparse parity checks and using noisy solutions to the parity checks as PRC codewords. For appropriate choices of parameters, Christ & Gunn (2024) prove that this distribution is cryptographically pseudorandom. We describe how the PRC works in detail in Appendix B, and we describe our watermarking algorithms in detail in Appendix C.

For a full description of the algorithm, see Algorithm 6.

Let PRC\mathsf{PRC} be any PRC, and let PRCWat.Sample\mathsf{PRCWat}.\mathsf{Sample} be as defined above. Then for any efficient algorithm A⁡\operatorname{\mathcal{A}} and any c>0c>0,

The notation \advO⁡(\secparam)\adv^{\operatorname{\mathcal{O}}}(\secparam) means that \adv\adv is allowed to run in any time that is polynomial in \secpar\secpar, making an arbitrary number of queries to O⁡\operatorname{\mathcal{O}}. For our experiments, we do not strictly adhere to the parameter bounds required for the pseudorandomness proof of Christ & Gunn (2024) to hold; as a result of this and the fact that we use small finite choices of parameters, our scheme should not be used for undetectability-critical applications. See Section C.3 for a discussion on this point.

To detect the watermark with the watermarking key, we use (roughly) the following algorithm. As long as Recover\mathsf{Recover} reproduces a good enough approximation to the latent that was originally used to generate an image, PRCWat.Detect\mathsf{PRCWat}.\mathsf{Detect} will recognize the watermark.

For our actual detector, we use a slightly more complicated algorithm that accounts for the fact that coordinates of z(T){\bm{z}}^{(T)} with larger magnitude are more reliable. The complete algorithm is given in Algorithm 7.

It turns out that, for low error rates, the PRC from Christ & Gunn (2024) can be used to encode and decode long messages using an algorithm called belief propagation. We can therefore include long messages in our watermark. Our algorithm for decoding the message from an image is PRCWat.Decode\mathsf{PRCWat}.\mathsf{Decode}, described in Algorithm 8. PRCWat.Decode\mathsf{PRCWat}.\mathsf{Decode} is slower and less robust than PRCWat.Detect\mathsf{PRCWat}.\mathsf{Detect}, but we find that it still achieves an interesting level of robustness.

Finally, our PRC watermark allows the user to set a desired false positive rate, FF. We prove Theorem 2, which says that our PRC watermark detector has false positive rate at most FF, in Section C.2.

In words, Theorem 2 says that any image generated independently of the watermarking key has at most a probability of FF of being identified as “watermarked” by our watermark detector or decoder.

Experiments

In our primary experiments, we focus on text-to-image latent diffusion models, utilizing the widely adopted Stable Diffusion framework (Rombach et al., 2022). Specifically, we evaluate the performance of various watermarking schemes using the Stable Diffusion-v2.1https://huggingface.co/stabilityai/stable-diffusion-2-1-base model, a state-of-the-art generative model for high-fidelity image generation. Additionally, we explore applying PRC watermarking to other generative models, as demonstrated with VAE (Kingma & Welling, 2013) models in Appendix D. All images are generated at a resolution of 512×\times512 with a latent space of 4×\times64×\times64. During inference, we apply a classifier-free guidance scale of 3.0 and sample over 50 steps using DPMSolver (Lu et al., 2022). As described in Section 3, we perform diffusion inversion using the exact inversion method from Hong et al. (2023) to obtain the latent variable z(T){\bm{z}}^{(T)}. In particular, we use 50 inversion steps and an inverse order of 0 to expedite detection, balancing accuracy and computational efficiency. All experiments are conducted on NVIDIA H100 GPUs.

We conduct comparative evaluations against various watermarking schemes, including in-, and post-processing techniques, as defined in Section 2. For post-processing methods, we compare with DwtDct (Al-Haj, 2007), DwtDctSvd (Navas et al., 2008), RivaGAN (Zhang et al., 2019), StegaStamp (Tancik et al., 2020), and SSL Watermark (Fernandez et al., 2022). For in-processing methods, we include a comparison with Stable Signature (Fernandez et al., 2023), Tree-Ring (Wen et al., 2023) and Gaussian Shading (Yang et al., 2024). Most baseline methods are designed to embed multi-bit strings within an image. Specifically, we set 32 bits for DwtDctSvd, RivaGAN, and SSL Watermark; 96 bits for StegaStamp; and 48 bits for Stable Signature. We employ publicly available code for each method, using the default inference and fine-tuning parameters specified in original respective papers for post- and in-processing methods. For Tree-Ring and Gaussian Shading watermarks, we use the same diffusion model and inference parameter settings as those used in PRC. We encode 512 random bits in the PRC watermark. If the decoder is successful, then with high probability, the bits are recovered correctly. Figure 1 illustrates examples of different watermarking schemes applied to a specific text prompt, highlighting the visual impact of each approach.

We evaluate watermarking methods on two datasets: MS-COCO (Lin et al., 2014) and the Stable Diffusion Prompt (SDP) dataset.https://huggingface.co/datasets/Gustavosta/Stable-Diffusion-Prompts We generate 500 un-watermarked images using MS-COCO captions or SDP prompts, and apply post-processing watermark methods to generate watermarked images. In-processing methods directly generate watermarked images from prompts. To assess the performance of the different watermarking schemes, we primarily examine four aspects: effectiveness, image quality, robustness, and detectability. For effectiveness, which involves performing binary classification between watermarked and un-watermarked images, we calculate the true positive rate (TPR) at a fixed false positive rate (FPR). Specifically, we report TPR@FPR=0.01. Without any attacks, the PRC watermark achieves TPR=1.0@FPR=0.01. Note that for PRC watermarking, the FPR is set at 1%, though it can be easily made smaller depending on the use case (see long message experiments in Section 4.4).

2 Quality and detectability

To evaluate the image quality of watermarked images, we compute the Frechet Inception Distance (FID) (Heusel et al., 2017), CLIP Score (Radford et al., 2021), and Inception Score (Salimans et al., 2016) to measure the distance between generated watermarked and un-watermarked images, and between watermarked and real images. For our comparison to real images, we use the MS-COCO-2017 training set; for the comparison to un-watermarked images, we use 8,000 images generated by the un-watermarked diffusion model using prompts from the SDP dataset. We calculate FID and CLIP Scores over five-fold cross-validation and report the mean and standard error. To assess perceptual variability (diversity), we select 10 diverse prompts from the PromptHero websitehttps://prompthero.com/ and use different in-processing watermark methods to generate 100 images for each prompt. We calculate perceptual similarity for all image pairs using the LPIPS (Zhang et al., 2019) score, averaging the results over the 10 prompts and reporting the standard error. Higher LPIPS scores indicate better variability for a given prompt. This evaluation is essential since, for image generation tasks, users typically generate multiple images from a single prompt and then select the best one (e.g., Midjourney).

Table 1 presents the empirical results for FID, CLIP, and Inception Scores for different watermarking schemes on both the COCO and SDP datasets. The table compares original image quality, post-processing watermark quality, and in-processing watermark quality (separated by dashed lines). We observe that StegaStamp results in the most significant quality degradation among post-processing watermark schemes, and the PRC watermark is the only method that consistently preserves image quality across all three metrics on both datasets. Table 2 presents the results of the variability analysis. Since post-processing methods are expected to have minimal impact on image variability, they are excluded from this table. The PRC watermark demonstrates variability comparable to un-watermarked images, outperforming the other in-processing schemes in this regard.

To evaluate detectability, we use ResNet18 (He et al., 2016) as the backbone model and train it on 7,500 un-watermarked images and 7,500 watermarked images (or 7,500 images watermarked with key 1 and 7,500 with key 2) to perform binary classification. Each experiment tests different watermarking schemes, with results shown in Figure 2. For the PRC watermark, the neural network slowly converges to perfect detection on the training set but achieves only 50% accuracy (random guess) on the validation set, indicating that the network is memorizing the training samples rather than learning the watermark pattern. In contrast, for all other schemes, the network performs perfectly on the validation set, demonstrating that the watermark is learnable.

3 Robustness of the Detector

To comprehensively evaluate the robustness of the PRC watermark and compare it to baseline watermarking methods, we tested nine distinct watermarking techniques against ten different types of attacks. Detailed descriptions of the attack configurations can be found in Section A.1.

The robustness of the various watermarking methods under these attacks is shown in Figure 5. We evaluated the quality of the attacked images using PSNR, SSIM, and FID metrics, comparing them to the original watermarked images. Notably, the PRC watermark demonstrates high resilience to most attacks. Even under sophisticated attacks, no method successfully reduced the true positive rate (TPR) below 0.99 while keeping the FID score under 70. This demonstrates that current watermark removal techniques struggle to erase our watermark without significantly degrading image quality. For instance, as shown in Figure 5, a JPEG compression attack with a quality factor of 20 only reduced the TPR from 1.0 to 0.94, but the resulting images displayed noticeable blurriness and a loss of detail (see Figure 4). Finally, in Figure 7 we demonstrate increased robustness for t=2t=2.For our other experiments we set t=3t=3. See Appendix B for details on the meaning of the parameter tt. However, for t=2t=2 there exist fast attacks on the undetectability of the PRC watermark, so we do not explore this choice further.

4 Encoding long messages in the watermark

The use of a PRC allows us to embed long messages in our watermarks, as described in Appendix C. We find in Figure 11 that the decoder is highly robust for 512-bit messages, although the detector is slightly more robust in this case. We find in Figure 12 that the decoder can reliably recover up to 2500 bits of information if the images are not subjected to removal attacks.

5 Security of the PRC watermark under spoofing attacks

To test the spoofing robustness of different watermarks, we followed the approach in Saberi et al. (2023), aiming to classify non-watermarked images as watermarked (increasing the false positive rate). Spoofing attacks can damage the reputation of generative model developers by falsely attributing watermarks to images. We used a PGD-based (Madry et al., 2018) method similar to that of the surrogate model adversarial attacks, flipping the surrogate model’s prediction from un-watermarked to watermarked. Just as with the adversarial surrogate attack, this attack cannot work against any undetectable watermark such as PRC watermark.

6 Possibility of extension

The PRC watermark can also be applied to other generative models, particularly those sampling from Gaussian distributions. We have set up a demo experiment working for traditional VAE models, as detailed in Appendix D. We would also be interested to see the PRC watermark applied to emerging generative models such as Flow matching (Lipman et al., 2022); whether or not this is possible hinges only on the existence of a suitable Recover\mathsf{Recover} algorithm.

Conclusion

We give a new approach to watermarking for generative image models that incurs no observable shift in the generated image distribution and encodes long messages. We show that these strong guarantees do not preclude strong robustness: Our watermarks achieve robustness that is competitive with state-of-the-art schemes that incur large, observable shifts in the generated image distribution.

SG is supported by a Google PhD Fellowship. This work is also partially supported by the National Science Foundation under grant no. 2229876. Any opinions, findings and conclusions or recommendations expressed in this material are those of the authors and do not reflect the views of the supporting entities.

References

Appendix A Additional experiment results and details on robustness

The figures included in this section are:

Figure 5, a comprehensive evaluation of watermarking schemes under the attacks described in Section A.2.

Figure 6, the performance of the embedding attack on in-processing watermarks.

Figure 7, a brief evaluation of the robustness of our PRC watermark with t=2t=2.

Figure 8, the performance of the spoofing attack against in-processing watermarks.

Figure 9, example images under the embedding attack.

Figure 11, a brief evaluation of the robustness of our PRC watermark decoder for 512 bits.

Figure 12, the length of messages which can be reliably encoded and decoded with out PRC watermark when there is no watermark removal attack.

A.2 Details on robustness

We applied a range of attacks, categorized into photometric distortions, degradation distortions, regeneration attacks, adversarial attacks, and spoofing attacks. Each type is described in detail below.

We applied two photometric distortion attacks: brightness and contrast adjustments. For brightness, we tested enhancement factors of ,whereafactorof0.0resultsinacompletelyblackimage,and1.0retainstheoriginalimage.Similarly,forcontrast,weusedenhancementfactorsof, where a factor of 0.0 results in a completely black image, and 1.0 retains the original image. Similarly, for contrast, we used enhancement factors of, where a factor of 0.0 produces a solid gray image, and 1.0 preserves the original image.

Three types of degradation distortions were applied: Gaussian blur, Gaussian noise, and JPEG compression. Specifically:

Gaussian Blur: We varied the radius from $$.

Gaussian Noise: Noise was introduced with a mean of 0 and standard deviations of $$.

JPEG Compression: Compression quality was set at $$, with lower quality levels leading to higher degradation.

Regeneration attacks (Zhao et al., 2023) alter an image’s latent representation by first introducing noise and then applying a denoising process. We implemented two forms of regeneration attacks: diffusion model-based and VAE-based approaches.

Diffusion Model Regeneration: We employed the Stable-Diffusion-2-1base model as the backbone and conducted $$ diffusion steps to attack the image. As the number of steps increased, the image diverged further from the original, often causing a performance drop. Interestingly, for the FID metric, we observed that more diffusion steps sometimes improved the FID score, as the diffusion model’s inherent purification process preserved a natural appearance while altering textures and styles.

VAE-Based Regeneration: We used two pre-trained image compression models from the CompressAI libraryhttps://github.com/InterDigitalInc/CompressAI: Bmshj2018 Ballé et al. (2018) and Cheng2020 Cheng et al. (2020), referred to as Regen-VAE-B and Regen-VAE-C, respectively. Compression factors were set to $$, where lower compression factors resulted in more heavily degraded images.

We also explored adversarial attacks, focusing on surrogate detector-based and embedding-based adversarial methods.

Surrogate Detector Attacks: Following (Saberi et al., 2023), we trained a ResNet18 model (He et al., 2016) on watermarked and non-watermarked images to act as a surrogate classifier. Specifically, we train the model for 10 epochs with a batch size of 128 and a learning rate of 1e-4. Using this model, we applied Projected Gradient Descent (PGD) adversarial attacks (Madry et al., 2018) on test images, simulating an adversary who knows either un-watermarked images and watermarked images (Adversarial-Cls), or watermarked images with two different keys (Adversarial-Cls-Diff-Key). The goal was to perturb the images with one key such that the detector misclassifies them as being associated with the other key. The attack was tested on four watermarking methods: Tree-Ring, Gaussian Shading, PRC, and StegaStamp watermark, with epsilon values of $$. Since the PRC watermark is undetectable, we find in Figure 2 that the surrogate classifier cannot even be trained!

Embedding-Based Adversarial Attacks: Adversarial perturbations were also applied to the image embedding space. Given an encoder f:X→Zf:\mathcal{X}\rightarrow\mathcal{Z} that maps images to latent features, we crafted adversarial images xadvx_{\text{adv}} to diverge from the original watermarked image xx, constrained within an l∞l_{\infty} perturbation limit. This was solved using the PGD algorithm (Madry et al., 2018). The VAE model for the original diffusion model stabilityai/sd-vae-ft-mse was assumed to be known for this attack.

Appendix B The pseudorandom code

We use the construction of a PRC from Christ & Gunn (2024), which is secure under the certain-subexponential hardness of LPN. The proof of pseudorandomness, assuming the 2ω(λ)2^{\omega(\sqrt{\lambda})} hardness of LPN, from the technical overview of Christ & Gunn (2024) applies identically here. The PRC works by essentially embedding random parity checks in codewords. The key generation and encoding algorithms are given in Algorithms 1 and 2.

Since the work of Christ & Gunn (2024), at least two new constructions of PRCs have been introduced using different assumptions (Golowich & Moitra (2024); Ghentiyala & Guruswami (2024)). It would be interesting to see if any of these new constructions yield image watermarks with improved robustness.

The main difference between the PRC used here and the one from the technical overview of Christ & Gunn (2024) is that ours is optimized for our setting by allowing soft decisions on the recovered bits. That is, PRC.Detect\mathsf{PRC}.\mathsf{Detect} takes in not a bit-string but a vector s{\bm{s}} of values in the interval $.IfthePRCcodewordis. If the PRC codeword is{\bm{c}},then, thens_{i}shouldbetheexpectedvalueshould be the expected value(-1)^{c_{i}}conditionedontheuser’sobservation.Wepresentconditioned on the user’s observation. We present\mathsf{PRC}.\mathsf{Detect}$ in Algorithm 3 and explain how we designed it in Section C.1.

Christ & Gunn (2024) show that any zero-bit PRC (i.e., a PRC with a Detect\mathsf{Detect} algorithm but no Decode\mathsf{Decode}) can be generically converted to one that encodes information at a linear rate. However, that construction requires increasing the block-length of the PRC, which could harm the practical performance of our watermark. Instead, we use belief propagation with ordered statistics decoding to directly decode the message. Note that belief propagation cannot handle a constant rate of errors if the sparsity is greater than a constant; therefore, this only works when Recover\mathsf{Recover} produces an accurate approximation to the initial latent. Still, since our robustness experiments use a small sparsity of t=3t=3, we find that our decoder functions even when the image is subjected to significant perturbation.

The only parameters that need to be set in PRC.KeyGen\mathsf{PRC}.\mathsf{KeyGen} are:

nn, the block length, which is the dimension of the image latents in the PRC watermark. Holding the other parameters constant, larger nn will yield a more robust PRC.

FF, the desired false positive rate. We prove in Theorem 2 that the scheme will always have a false positive rate of at most FF, as long as the string being tested does not depend on the PRC key.

tt, the sparsity of parity checks. Larger tt yields undetectability against more-powerful adversaries, but decreased robustness.

For watermark detection and decoding, we allow the user to set an estimated error σ\sigma. This should be the standard deviation of the error z′−z{\bm{z}}^{\prime}-{\bm{z}} that the user expects. In cases where the watermark does not need to be robust to perturbations of the image, one can set σ=0\sigma=0. If σ\sigma is not set by the user, we use a default of σ=3/2\sigma=\sqrt{3/2} which we found to be effective for robust watermarking.

Appendix C Details on the PRC watermark

Watermark key generation, Algorithm 5, is exactly the same as PRC key generation.

Watermarked image generation works by sampling the initial latents to have signs chosen according to a PRC codeword. If a message is to be encoded in the watermark, the message is simply encoded into the PRC.

Our detection algorithm PRC.Detect\mathsf{PRC}.\mathsf{Detect} is given in Algorithm 3. In Section C.1 we will explain how we designed the detector, and in Section C.2 we will prove Theorem 2 which says that PRC.Detect\mathsf{PRC}.\mathsf{Detect} and PRC.Decode\mathsf{PRC}.\mathsf{Decode} have false positive rates of at most FF. Note that PRC.Decode\mathsf{PRC}.\mathsf{Decode} is guaranteed to have a false positive rate of at most FF simply because of testbits\mathsf{testbits}.

Let z{\bm{z}} be the initial latent and z′{\bm{z}}^{\prime} be the recovered latent. We will compute the probability that a given parity check ww is satisfied by sign⁡(z)\operatorname{sign}({\bm{z}}) (after accounting for the noise and one-time pad), conditioned on the observation of z′{\bm{z}}^{\prime}. In order for this to be possible, we need to model the distributions of z{\bm{z}} and z′{\bm{z}}^{\prime}: We use z∼N⁡(0,In){\bm{z}}\sim\operatorname{\mathcal{N}}({\bm{0}},{\bm{I}}_{n}) and z′∼N⁡(z,σ2In){\bm{z}}^{\prime}\sim\operatorname{\mathcal{N}}({\bm{z}},\sigma^{2}{\bm{I}}_{n}) for some σ>0\sigma>0.

Crucially, when we bound the false positive rate in Section C.2, we will do it in a way that does not depend on the distribution of z′{\bm{z}}^{\prime}; we only use the facts that z∼N⁡(0,In){\bm{z}}\sim\operatorname{\mathcal{N}}({\bm{0}},{\bm{I}}_{n}) and z′∼N⁡(z,σ2In){\bm{z}}^{\prime}\sim\operatorname{\mathcal{N}}({\bm{z}},\sigma^{2}{\bm{I}}_{n}) to inform the design of our detector. In other words, Theorem 2 holds unconditionally, even though our detector is designed to have the highest true positive rate for a particular distribution of z′{\bm{z}}^{\prime}.

Our first step is to compute the posterior distribution on sign⁡(z)\operatorname{sign}({\bm{z}}), conditioned on the observation z′{\bm{z}}^{\prime}.

If z∼N⁡(0,1)z\sim\operatorname{\mathcal{N}}(0,1) and z′∼N⁡(z,σ2)z^{\prime}\sim\operatorname{\mathcal{N}}(z,\sigma^{2}) then

The joint distribution of (z,z′)(z,z^{\prime}) is

Using the formula for the conditional multivariate normal distribution,See, for instance, (Holt & Nguyen, 2023, Theorem 3). the distribution of zz conditioned on z′z^{\prime} is

where Φ\Phi is the cumulative distribution function of the standard normal distribution, so

where we have used the fact that Φ(x)=(1+\erf(x/2))/2\Phi(x)=(1+\erf(x/\sqrt{2}))/2. ∎

for each i∈[n]i\in[n]. Let aw=∏j∈w(−1)otpja_{\bm{w}}=\prod_{j\in{\bm{w}}}(-1)^{\mathsf{otp}_{j}} and s^w=∏j∈wsj\hat{s}_{\bm{w}}=\prod_{j\in{\bm{w}}}s_{j} for each w∈P{\bm{w}}\in{\bm{P}}. Then (1+aws^w)/2(1+a_{\bm{w}}\hat{s}_{\bm{w}})/2 is the probability that (−1)otp⊕e⋅sign⁡(z)(-1)^{\mathsf{otp}\oplus{\bm{e}}}\cdot\operatorname{sign}({\bm{z}}) satisfies w{\bm{w}}.

is greater than some threshold. We set the threshold by computing a bound on the false positive rate.

C.2 Bounding the false positive rate: Proof of Theorem 2

To compute a bound on the false positive rate of the detector, we use a “bounded from above” version of Hoeffding’s inequality due to Fan et al. (2015):

We are now ready to prove Theorem 2. We state the theorem for the PRC detector and decoder; note that this immediately implies the same result for the PRC watermark detector and decoder.

Observe that the use of testbits\mathsf{testbits} immediately implies that PRC.Detect\mathsf{PRC}.\mathsf{Detect} has a false positive rate of at most FF, i.e.,

We therefore turn to analyzing the false positive rate of PRC.Detect\mathsf{PRC}.\mathsf{Detect}. We adopt the notation from Section C.1, with

By construction of the parity check matrix, the parity checks w∈P{\bm{w}}\in{\bm{P}} are linearly independent. Since otp\mathsf{otp} is uniformly random, it follows that the values awa_{\bm{w}} are independent and uniformly random from {−1,1}\{-1,1\}. Therefore by 2 it suffices to show that

where rr is the number of parity checks in P{\bm{P}} and C=12∑w∈Plog⁡2(1+s^w1−s^w)C=\frac{1}{2}\sum_{{\bm{w}}\in{\bm{P}}}\log^{2}\left(\frac{1+\hat{s}_{\bm{w}}}{1-\hat{s}_{\bm{w}}}\right).

for each w∈P{\bm{w}}\in{\bm{P}}. Since a{\bm{a}} is random, each fw(aw)f_{\bm{w}}(a_{\bm{w}}) is uniformly random from (1±s^w)/2(1\pm\hat{s}_{\bm{w}})/2.

where C=12∑w∈Plog⁡2(1+s^w1−s^w)C=\frac{1}{2}\sum_{{\bm{w}}\in{\bm{P}}}\log^{2}\left(\frac{1+\hat{s}_{\bm{w}}}{1-\hat{s}_{\bm{w}}}\right). The claim follows by setting τ=Clog⁡(1/F)\tau=\sqrt{C\log(1/F)}. ∎

C.3 Practical undetectability

We have not yet discussed the extent to which our scheme is undetectable for practical image sizes. As observed by Christ et al. (2024), the undetectability of any watermarking scheme can be broken with enough samples and computational resources: Undetectability just means that the resources required to detect the watermark without the key scale super-polynomially with the resources required to detect the watermark with the key. And under the same assumptions as in Christ & Gunn (2024), our scheme is asymptotically undetectable for the right scaling of parameters. We refer to our scheme as “undetectable” because of this, and because our experiments on quality and detectability demonstrate that it is undetectable enough for the main practical applications. However, for the specific, concrete choices of parameters used in our experiments, undetectability is not guaranteed against motivated adversaries.

For the PRC watermark, there exists a brute-force attack on undetectability that runs in time O(nt−1)O(n^{t-1}), counting queries to the generative model as O(1)O(1), where nn is the dimension of the image latents and tt is the sparsity of parity checks which can be set by the user (larger tt decreases the robustness). This attack works by simply iterating over tt-sparse parity checks until one used by the watermark is found. We did not attempt to optimize the attack, so it is possible that faster attacks could be found.

In our experiments we have n=214n=2^{14} dimensional image latents, and we set t=3t=3 for most of our experiments demonstrating robustness. To ensure cryptographic undetectability, a better choice would be t=log⁡2(n)/2=7t=\log_{2}(n)/2=7. The watermark detector still works with t=7t=7 for non-perturbed images, but we choose t=3t=3 for most experiments because of the improved robustness. Note that O(n2)O(n^{2}) is far greater than the O(1)O(1) time required to detect prior watermarks without the key, but a motivated adversary can still break the undetectability of our scheme. We therefore stress that our scheme, in its current form, should not be used for undetectability-critical applications such as steganography.

The reason there exists a relatively fast brute-force distinguishing attack against our scheme is that there exist quasi-polynomial time attacks against the PRC of Christ & Gunn (2024). The alternative constructions of PRCs due to Golowich & Moitra (2024) and Ghentiyala & Guruswami (2024) also suffer from quasi-polynomial time attacks. It is an interesting open question to construct PRCs that do not have quasi-polynomial time attacks; using our transformation, any such PRC would generically yield a watermarking scheme with improved undetectability. We hope that generative image model watermarks with improved undetectability can be built in the future.

Appendix D Demo: PRC watermark for VAEs

The PRC watermark can be applied to VAEs (Kingma & Welling, 2013) as well. Using the same gradient descent technique as Hong et al. (2023), we optimize the latent to obtain the decoder inversion result for watermark detection. We test the PRC watermark on a VAE with a 256-dimensional latent space, trained on the CelebA dataset (Liu et al., 2018). By setting t=2t=2 and FPR as 0.05, we achieve over 90% TPR when embedding a zero-bit PRC watermark in the images. We show example generated images in Figure 13. We did not investigate the robustness or quality of the PRC watermark for VAEs in-depth, so this section is only to demonstrate the generality of our technique.