Hidden in the Noise: Two-Stage Robust Watermarking for Images

Kasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde, Niv Cohen

Introduction

Generative AI is capable of synthesizing high-quality images indistinguishable from real ones. This capability can be used to deliberately deceive. These fake image generations, called deepfakes, have the potential to cause severe societal harms through the spread of confusion and misinformation (Peebles & Xie, 2022; Esser et al., 2024; Chen et al., 2024; Ramesh et al., 2021). In addition, owners of different models and images may want to control the spread of their derivatives for copyright reasons and safeguard their intellectual property. One way to mitigate these harms is model watermarking. The study of watermarking has a rich history and has recently been adopted for AI-generated content (Pun et al., 1997; Langelaar et al., 2000; Craver et al., 1998). For an extended discussion of recent work in this area, we direct the reader to Appendix B. Unfortunately, most current image watermarking methods are not robust to watermark removal attacks utilizing image diffusion generative models (Zhao et al., 2023a).

Recently, new watermarking methods utilize the inversion property of DDIM to achieve more robust watermarking (Wen et al., 2023; Ci et al., 2024; Yang et al., 2024b). These methods embed patterns in a diffusion model’s initial noise and then detect them in the noise pattern reconstructed from the generated image. This technique provides strong robustness against various attacks, making it effective at resisting watermark removal attempts. Yet, these prior methods are themselves vulnerable to new types of attacks. Tree-Ring Wen et al. (2023) add a pattern to the initial noise, making it distinct from a random Gaussian initial noise in a way that an attacker can detect (Yang et al., 2024a). This may enable forgery attacks, aimed at applying the watermark without the owner’s permission. Such attacks are often even more concerning than removal attacks, as they can cause severe damage to model owners if their watermark is associated with illegal content.

Therefore, there is a need for image watermarking methods that generate images that are not distinct from non-watermarked images (to anyone but the model owner). As suggested by previous works, since the model already takes random noise as initialization, we may initialize it with a pseudo-random noise pattern that we can detect later (Yang et al., 2024b; Kuditipudi et al., 2023). Namely, reconstructing an approximation of the initial noise used in the diffusion process from a given image allows the detection of the noise pattern used by the model. Although this reconstructed noise is not completely identical to the used initial noise, it is much more similar to the initial noise than it is to other randomly distributed noise patterns. Thus, it can serve as a watermark that can be identified in the generated images (see also Appendix C for similar ideas used in previous works).

While using a pseudo-random initial noise does not distort the distribution of single-generated images, it may carry information about the watermark when groups of images are examined together. Specifically, works such as Yang et al. (2024b) embeds the watermark in an initial noise such that the resulting generated image comes from the same distribution as non-watermarked images. Yet, when many images generated from the same noise pattern are examined together, the correlation between them may expose that they are not distortion-free as a set. E.g., the average of many similarly-watermarked images may differ from the average of non-watermarked ones (Yang et al., 2024a). A natural solution to this distortion of sets is to use more than one initial noise for each watermark we deploy.

Yet, given a sufficiently small set of initial noises (denoted as NN) and an enormous number of images generated by a model, an attacker could potentially still collect many images sharing the same initial noise in order to perform removal and forgery attacks as was applied to previous methods (Yang et al., 2024a). Using many initial noises (a large value of NN) will make such attacks much more difficult, if not infeasible. Surprisingly, we find that a very large number of random initial noises remain distinguishable from one another, even after reconstructing the noise from a generated image. However, a large value of NN might incur a negative effect on the runtime and accuracy of the approach. In order to lower the effective quantity of noises we need to scan at detection while retaining strong robustness, we propose a two-stage efficient watermarking framework. We supplement our NN initial noise samples with MM Fourier patterns as a group identifier - a unique identifier of a subset of initial noises we might have used for generating a given image (Figure 1). During detection, we may first recover the group identifier (stage 1) and use it to find an exact match (stage 2). Thus, we reduce our search space to the number of initial noises per group (N/MN/M).

We demonstrate that the initial noise used in the diffusion process is itself a distortion-free watermarking method for images (Section 3).

We present WIND, our two-stage method for effectively using the initial noise as a watermark (Section 4).

We demonstrate that WIND achieves state-of-the-art results for its robustness to removal and forgery attempts (Section 5).

Preliminaries

In a watermarking scheme we usually consider the owner, trying to mark images as an output of their model; and an attacker, trying to remove or forge the watermark on unrelated images.

The Owner releases a private model (diffusion model in our case) that clients can access through an API, allowing them to generate images that contain a watermark. The watermark is designed to have a negligible impact on the quality of the generated images. There are a few settings regarding the watermark detection, including public infomation and private information watermarking (Cox et al., 2007; Wong & Memon, 2001). We focus on the setting where the watermark is detectable only by the owner, enabling them to verify whether a given image was generated by their model using private information.

The Attacker uses the API to generate an image and subsequently attempts to launch a malicious attack aimed at either removing or forging the embedded watermark, with the intention of using the image or watermark for unauthorized purposes.

2 Diffusion Models Inversion

3 Tree-Ring and RingID Watermarks

In order to watermark images in a human-imperceptible and robust way, previous works have encoded specific patterns in the Fourier space of the initial noise. Tree-Ring (Wen et al., 2023) first transforms the initial noise into the Fourier space. A key pattern is then embedded into the center of the transformed noise. The noise is subsequently transformed back into the spatial domain. During the detection phase, the diffusion process is inverted, and the Fourier domain is examined to verify the presence of the imprinted pattern. RingID (Ci et al., 2024) shows that Tree-Ring struggles to distinguish between different keys. Therefore, the number of unique keys (distinguishable from one another) that can be embedded with Tree-Ring is low. They increase the possible number of unique keys that can be encoded using Fourier patterns.

Systematic Distribution Shifts in Generated Images Enable Attacks. Systematic distribution shifts in the generated content make it easier to verify the existence of a watermark. However, in the case of Tree-Ring and other watermarking techniques, it also opens up an avenue of attack (Wen et al., 2023; Yang et al., 2024b; Xian et al., 2024; Bui et al., 2023). Emblematic is the method of Yang et al. (2024a), whose attack approximates the difference between watermarked and non-watermarked images. Increasing the number of images with the watermark can improve the accuracy of the approximation. The impact of distribution shifts is significant, as the attack remains effective even when the watermarked and non-watermarked images are not paired (Figure 3).

Initial Noise is a Distortion Free Watermark

Watermarks which systematically perturb the distribution of image generations are more vulnerable to removal and forgery attacks. A distortion-free watermarking method, by contrast, is more robust (Kuditipudi et al., 2023). Our first finding is that the initial noise already in standard use in diffusion models can be such a watermark.

Let NN be the number of initial noises we can generate. We will secure our watermarking process with a long, secret salt ss. We begin by sampling a random (and reproducible) initial noise. Let i∗∼Unif([N])i^{*}\sim\textrm{Unif}([N]) be the index of the initial noise. We will use a hash function to get a seed hash(i∗,s)\textrm{hash}(i^{*},s). Plugging the seed into a pseudorandom generator, we generate a reproducible initial noise vector zi∗∼N(0,I)\mathbf{z}_{i^{*}}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) drawn from a centered Gaussian distribution. When we generate fewer than NN images, we can use each initial noise at most once and the noise appears distortion-free. We discuss the case when the number of images exceeds NN in Appendix F.

Runtime considerations. Our method requires searching over all NN watermarks, leading to a naive runtime complexity of O(N)\mathcal{O}(N). However, more efficient algorithms for similarity-based search, such as HNSW (Malkov & Yashunin, 2018), can reduce this complexity to O(log⁡N)\mathcal{O}(\log N), at the expense of additional memory usage. We provide empirical runtime analysis of our method in Appendix H. For large enough values of NN, this cost may eventually become undesirable. Together with our aim to maintain high robustness with an increasing number of keys, it motivates a more efficient method, which is presented in the next section.

Method

While always using a single initial noise for our model might imply good robustness properties, to make forgery and removal more difficult, it is generally preferable to maintain a large set of NN initial noises to be used by the model. More importantly, using a large number of different noises NN may serve as different keys, encoding some metadata about each image. This metadata might include information about the specific model that generated it, as well as additional information about the generation for further validation of the image source, once detected.

In order to make the search over a large number of noises more efficient, we introduce a two-stage efficient watermarking approach we name WIND (Watermarking with Indistinguishable and Robust Noise for Diffusion Models). First, we initialize MM groups of initial noise, each group associated with its own Fourier-pattern key. In contrast to prior work, we employ these Fourier patterns not as a watermark, but as a group identifier to reduce the search space.

2 Resilience to Forgery

In addition to empirical evaluations of specific attacks as in Figures 2 and 3; we discuss below the attacker’s ability to infer knowledge about the used noise pattern across different watermarked images. Even if the attacker is able to obtain information about a specific initial noise zi\mathbf{z}_{i} for an index ii (which is an extreme case), the other noise vectors for j≠ij\neq i are still safeWe note that obtaining a single noise pattern might not be enough to effectively forge the watermark, as the model owner may encode this pattern with additional metadata as described in Section 4.1. This is because we use a cryptographic hash function and a secret salt. Formally, Theorem 4.1 shows that, as long as the cryptographic hash function remains unbroken and the secret salt is kept private, the watermarking algorithm maintains its security properties against even very powerful adversaries.

Generate valid reconstructed noise zj\mathbf{z}_{j} for any other initial noise index j≠ij\neq i

3 Watermarking Non-Synthetic Images.

Until now, we have addressed watermarking only for AI-generated synthetic images. Yet, protecting copyrights, or preventing the spread of misinformation, may also apply to modified natural images. Most previous approaches to watermark diffusion models overlook attempting to expand their method to non-generated images. To allow using our framework for non-generated images, we expand our framework. By using diffusion inpainting, our watermark can be applied to a natural image. Later, by inverting the inpainted image we can verify the presence of the watermark.

As demonstrated in Figure 5, our inpainting method injects a watermark with minimal visual impact, preserving the original image’s integrity. Please see Appendix D for additional results.

Experiments

Setting. For a fair comparison with previous methods (Ci et al., 2024; Wen et al., 2023), we employed Stable Diffusion-v2 (Rombach et al., 2022), with 50 inference steps for both generation and inversion. Other implementation-details can be found in Appendix I.

Image Transformation Attacks. Following previous methods (Wen et al., 2023; Ci et al., 2024) we applied these image transformations to the generated images: 75∘75^{\circ} rotation, 25%25\% JPEG compression, 75%75\% random cropping and scaling (C & S), Gaussian blur with an 8×88\times 8 filter size, Gaussian noise with σ=0.1\sigma=0.1, and color jitter with a brightness factor uniformly sampled between 0 and 6. In Table 1 we compare our methods to both Tree-Ring and RingID. As the results demonstrate, using multiple keys with RingID (Ci et al., 2024) is possible. Yet, it remains vulnerable to cropping and scaling attacks. In contrast, WIND effectively addresses this challenge. It enables accurate watermark detection under all image transformation attacks. We note that the incorporation of the keys in the RingID method not only allows us to embed keys but also increases the robustness of the full method to certain attacks.

Steganalysis Attack. We assess the robustness of our method against the attack proposed by Yang et al. (2024a), which is capable of forging and removing the Tree-Ring and RingID keys. As discussed in Figure 2, this attack attempts to approximate the watermark by subtracting watermarked images from non-watermarked images. The results, presented in Figure 3, indicate that while the attack could be able to forge or remove our group identifier, it is unable to forge or remove our watermark (initial noises). Even when the Fourier pattern type key is removed through an exhaustive search, our method remains robust in identifying the correct initial noise.

Regeneration Attacks. Recently, Zhao et al. (2023a) introduced a two-stage regeneration attack: (i) adding noise to the representation of a watermarked image, and (ii) reconstructing the image from this noisy representation. To assess the resilience of our approach to regeneration attacks, we applied the attack from Zhao et al. (2023a) to watermarked images generated by our model. As shown in Table 3, the attack has a minimal impact on the distribution of the cosine similarities between the initial noise and the inverted noise. The attacked noise similarity still maintains a significant gap compared to random noise.

To examine the performance of our inpainting method, we report the Fréchet Inception Distance (FID) (Heusel et al., 2018) on the MS-COCO-2017 (Lin et al., 2015) training dataset in Table 3. Notably, our method achieves the lowest FID among the compared methods, indicating a closer alignment with real images. Additionally we include some images generated by our framework in Figure 4.

Discussion and Limitations

Editing a Given Image vs. Forging. While forging our watermark by obtaining the initial noise is hard (Section 3), an easier path to obtaining harmful watermarked images might be to apply a slight edit to an already watermarked image. An harmful image in this context might include a copy-right infringing image, NSFW image, or any other content the model owner wish to avoid being associated with. Naturally, there is a trade-off between the severity of the applied edit, and the edit ability to preserve the initial watermark. We present one solution to mitigating this issue in the next discussion point.

Storing a Database of Generations. Model owners wishing to protect themselves from an attacker modifying a watermark image may keep a database of the past generations by their model. For these extreme cases, the model owner might only save the used prompts and initial noiseseeds, and use the reconstructed noise to retrieve the entire set of prompts used with that specific seed (Huang & Wan, 2024). While this process may be resource-intensive, it is only required in the rare event that an attacker intentionally modifies a benign image into a harmful one while preserving the watermark.

Private Model. Our watermark robustness is based to a large extent on the inability of an attacker to invert a model, which is empirically validated but not mathematically proven. Yet, as discussed in Section 2.2, the ability to successfully invert our model may be nearly equivalent to the ability to steal the forward diffusion process, effectively stealing the model (in which case, any watermarking attempt might be deemed quite useless anyhow). Still, a better framing of the mathematical assumptions behind this claim is a limitation of this work, as well as of previous works on watermarking using inversion of the diffusion generative process.

Attacker’s Advantage. There exists a large set of diverse attacks aimed at watermark removal (Zhang et al., 2023; Yang et al., 2024a; Zhao et al., 2023a), along with image transformations such as rotation and crops that also achieve some limited success against our watermark. As in many security applications, we suspect that an attacker capable enough will still be able to remove the watermark using new techniques we might not expect. However, a more robust watermark may nevertheless help to decrease the spread of false information.

Additional discussion and limitations can be found in Appendix C.

Conclusion

In this work, we present a robust and distortion-free watermarking method that leverages the initial noises employed in diffusion models for image generation. By integrating existing techniques, we enhanced the approach to achieve improved efficiency and robustness against various types of attacks. Furthermore, we outlined a strategy for applying our method to non-generated images through inpainting.

References

Appendix A Notation

Appendix B Related Works

Memorization in Diffusion Models. Diffusion models (Ho et al., 2020; Sohl-Dickstein et al., 2015) have demonstrated a capacity not only to generalize but also to memorize training data. This can lead to the reproduction of specific patterns or, in some cases, exact content from the training set, including sensitive or proprietary information. This memorization poses significant risks of unintended intellectual property leakage, particularly in large-scale generative models. Several studies have shown that information from training data can be extracted from diffusion models (Carlini et al., 2023b; Somepalli et al., 2023b; Carlini et al., 2023a; Gu et al., 2023; Somepalli et al., 2023a).

Image Watermarking. Image watermarking is essential for protecting intellectual property, verifying content authenticity, and maintaining the integrity of digital media. The field has ranged from traditional signal processing techniques to recent deep learning methods (Potdar et al., 2005; Singh & Singh, 2023).

Among Early watermarking strategies, one of the simplest methods was Least Significant Bit (LSB) embedding, which modifies the least significant bits of image pixels to imperceptibly embed watermarks (Wolfgang & Delp, 1996). Another classical approach utilized frequency-domain transformations and Singular Value Decomposition (SVD) to hide watermarks within image coefficients. (Chang et al., 2005; Al-Haj, 2007).

Recent developments leverage deep learning for watermarking. For instance, HiDDeN (Zhu et al., 2018) introduced an end-to-end trainable framework for data hiding. RivaGAN (Zhang et al., 2019) utilizes adversarial training to embed watermarks, while Lukas & Kerschbaum (2023) proposed an embedding technique that optimizes efficiency by avoiding full generator retraining.

Watermarking for Diffusion Models. Existing watermark methods for diffusion models can be divided into three categories:

(i) Post-processing methods which adjust image features to embed watermarks (Zhao et al., 2023c; Fernandez et al., 2023b). This approach alters the image and its distribution, which can result in significant changes to the generated image. However, recent work by Zhao et al. (2023b) shows that pixel-level perturbations are removable by regeneration attacks makes. To date, this approach is not robust.

(ii) Fine-tuning-based approaches combine the watermark within the generation process (Zhao et al., 2023c; Xiong et al., 2023; Liu et al., 2023; Fernandez et al., 2023a; Cui et al., 2023). To date, these methods have robustness issues as well (Zhao et al., 2023a).

(iii) Tree-Ring introduced an approach to proposing a method to imprint a tree-ring pattern into the initial noise of a diffusion model (Wen et al., 2023). Each pattern is used as a key, which is added in the Fourier space of the noise. The verification of the presence of the key involves recovering the initial noise from the generated image and checking if the key is still detectable in Fourier space. This approach makes Tree-Ring and its follow-up works the most robust approach against attacks (Zhao et al., 2023a; An et al., 2024).

Recently, Yang et al. (2024a) took advantage of the distribution shift present in Tree-Ring that occurs with impainting keys and arranged the first successful black box attack against it, as we detailed in Figure 2.

Appendix C Additional Discussion and Limitations

The seminal work by Wen et al. (2023) innovated the use of initial noise in DDIM for watermarking. Most related to our work, Yang et al. (2024b) also embeds a watermark in the initial noise already used by a DDIM diffusion model. Yet, while Yang et al. (2024b) proposes a watermark that is distortion-free for a single image, it is not distortion-free when examining sets of images; therefore it is vulnerable to attacks such as Yang et al. (2024a). We aim to be robust to attacks even when many images are examined together.

There are additional technical differences between our approach and Yang et al. (2024b). Most notably: (i) Our work also studies applying our watermark to non-synthetic (natural) images, or images coming from other generative models. (ii) While Yang et al. (2024b) design a function to embed specific bits into the initial noise, we take another approach. Namely, we view the entire initial noise (with generation and inversion) as a noisy channel. Inspired by Shannon (1949), we use a random encoding of the watermark identities into the channel.

Computational Requirements. As discussed in Section 3 our similarity search can be accelerated given well-known methods. Yet, the computational requirements of our method might be limiting when trying to use our method on edge devices. However, similarly to Tree-Ring and Ring-ID (Wen et al., 2023; Ci et al., 2024) our method assumes a private model, which is usually not deployed on edge devices anyhow.

Trade-Offs Between the Watermarking Overhead, and Detection Accuracy.

We suggest the following variants of our method for different possible requirements of runtime scaling, detection robustness, and ease of adaptation.

A. Detection of the group identifier alone: This operation takes a search of O(M)\mathcal{O}(M), but is vulnerable to both removal and forgery attempts, as we use a weaker watermark for group identifiers.

B. Detection of the Fourier pattern, followed by a validation of the exact initial noise (WINDfast{}_{\text{fast}}): within the group. This operation takes O(N/M)\mathcal{O}(N/M) search. It is vulnerable to removal attempts, but more resilient to forgery attempts (see Table 1).

C. An exhaustive search of the initial noise, also outside the identified group (WINDfull{}_{\text{full}}): This operation takes O(N)\mathcal{O}(N) search. It is more resilient to both removal and forgery attempts (see Figure 3, and Table 1).

This method, while slower is also easier to adapt. A user that wishes to use a fast version of this variant may apply a similar algorithm to the one described above using only a few possible random noises. This would replace the distinguishability of many different watermarks with the ability to rapidly and simply detect the watermarked images.

Practically, an NN search can be accelerated using many methods and can be scaled to tens of millions without significantly affecting the detection time (Wei et al., 2024; Douze et al., 2024; Malkov & Yashunin, 2018; Andoni et al., 2018).

Image Quality Considerations.

Our method relies on using an initial random noise, drawn from the same distribution of initial noises already used by the model. Therefore, the core of our method (the initial noise stage) is not compromising the visual quality of the generated images at all.

The only effect on visual quality comes from the group identifier, where we use existing off-the-shelf watermarking images. In our implementation, we used the RingID (Ci et al., 2024) method that adds the Fourier pattern to the initial noise.

When a model owner wishes to preserve image quality even better, they may use any other existing watermarking method for the group identifier stage. This will still not compromise the security provided by the initial noise stage.

Inversion Attack.

As discussed in Section 2.2 in our paper, accurately inverting the model is as difficult as copying the forward process of the model (image generation). While hard, an attacker able to do so is effectively also capable of generating novel images using the same diffusion process. Therefore, At this stage, the model itself is effectively compromised (and not only the watermark signature). We believe that being as hard to forge as the model itself, is a reasonable level of security for almost all use cases.

Yet, approximately inverting the model might also be a threat. While even approximately inverting a model is also very hard, it might be easier than stealing the model. Still, we would like to emphasize that our method is more secure than other diffusion-process-based watermarking techniques, where image distortion themselves may allow easier forging (Yang et al., 2024a).

Appendix D Additional Results

We expect our watermark to be effective directly for any model for which some inversion to the original noise is possible. Namely, as the correlation between random noises in a very high dimension is very much concentrated around 0, even a very slight success in the inversion process is enough to be distinguishable. In higher generation resolutions the dimensionality of the noise is even higher, and therefore the separation would be even better (El Karoui, 2009).

Empirically, to validate the generality of our method, we also report results for the SD 1.4 model (Rombach et al., 2022). Using N=10000N=10000 noises and M=2048M=2048 group identifiers, our method achieved a detection accuracy of 97% to identify the correct watermark (initial noise).

In any case, our method of the reported SD 2.1 model can also be used to watermark images collected from other sources (please see Section 4.3, Section D.2).

D.2 Non-Synthetic Images Watermark Detection

Our inpainting method allows us to watermark both images generated by any model and non-generated images. To evaluate the robustness of the inpainting watermarking approach, we present results in Table 1 for this method, utilizing N=100N=100 noises. Results are shown in Table 6.

D.3 Further Exploration of the Regeneration Attack Perturbation Strength

In Section 5.1, we discussed the robustness of WIND against regeneration attacks. However, using it iteratively might still be a stronger adversary. We applied the regeneration attack proposed by Zhao et al. (2023a), up to 50 times. We see that iterative regeneration indeed decreases the similarity between the original noise and the reconstructed one. This happens as the image becomes less and less correlated to the original generation Figure 6.

Yet, the detection rate of our algorithm remains very high Table 14. We attribute this to the fact that even a slight remaining correlation between the attacked image and the initial noise remains significant with respect to the correlation expected from non-watermarked images. This happens because of the very low correlation between random (non-watermarked) noises (Figure 2).

D.4 Quantitative Analysis of the Effect on Image Quality

We reported the FID of our model on Table 3. To further assess the effect of WIND watermark on image quality we report the CLIP score Hessel et al. (2021) before and after watermarking on Table 10. Results indicate that adding the watermark has a negligible effect on the CLIP score for generated images.

To further quantify the distortion introduced by each model, we report pixel-base matrices, SSIM and PSNR in the two settings we study:

Images Generated by the Diffusion Model. WIND’s distortion arises from using group identifiers, enabling faster detection. To disentangle this effect, we also evaluate WINDw/o{}_{\text{w/o}}, which omits group identifiers. As can be seen in Table 11, the image quality generated using our full method is comparable to that of previous techniques. Users who wish to generate distortion-free images, without affecting image quality, can do so by omitting the group identifier (at the cost of a slower detection phase for very large values of NN).

Watermarking Non-Synthetic Images. Additionally, we present results for WINDinpainting{}_{\text{inpainting}}, our inpainting-based approach capable of watermarking both non-synthetic images and outputs from other generative models (Table 15). Although other watermarking methods may preserve image quality better, our image quality remains high. Importantly, to the best of our knowledge, our approach is the only one capable of watermarking non-synthetic images while remaining robust against the regeneration attack (Zhao et al., 2023a). Therefore, it is preferable when an adversary may try to remove the watermark.

In addition, the inpainting technique can be applied selectively to specific parts of the image if the copyright owner wishes to perfectly preserve fine details in certain areas.

D.5 Robustness Comparison to Different Number of Inference Steps

We evaluate the impact of inference steps on detection accuracy, as shown in Table 7. The results indicate that using 100 steps yields better detection accuracy compared to other step counts, including the 50 steps used in our main experiments.

D.6 True Positive and AUC

Expanding on the detection assessment settings discussed in Section 5, we reported WIND’s error bars. AUC and True Positive (TPR@1%FPR) results are available on Table 8. Demonstrate strong performance, emphasizing WIND’s robustness and reliability.

D.7 Evaluation Against Additional Attacks

We evaluate WIND against a diverse set of attacks, including transfer-based, query-based, and white-box methods. Specifically, we employ the WeVade white-box attack (Jiang et al., 2023), the transfer attack described in Hu et al. (2024), a black-box attack utilizing NES queries (Ilyas et al., 2018), and a random search approach discussed in Andriushchenko et al. (2024), adopted to attempt watermark removal. The success rates of these attacks are detailed in Table 9. Notably, none of these methods succeed against WIND, as the correct watermark remains detectable in over 97% of cases even after applying these attacks.

Appendix E Proof of Resilience to Forgery

The WIND method is an approach for generating multiple watermarked images. Theorem 4.1 tells us that compromising one or more watermarked images does not give away any information about any other watermarked images. E.g., the adversary cannot “generate valid reconstructed noise for any other initial noise index j≠ij\neq i”. That said, Theorem 4.1 does leave open the possibility that an adversary can take a watermarked image, reconstruct the initial noise only for that image, and use it to attack the method, which we evaluate empirically.

We will prove each part of the theorem separately:

1. The adversary cannot recover the secret salt ss: Given the output seed=\textrm{seed}= hash(i∗,s)(i^{*},s) and partial input i∗i^{*} the adversary aims to find ss. This is equivalent to finding a pre-image of given partial information about the input. By the pre-image resistance property of cryptographic hash functions, this task is computationally infeasible. Even if the adversary knows all possible values of ii, the space of possible secret salts ss is too large to search exhaustively (as ss is a sufficiently long random string). Therefore, the adversary cannot recover ss.

2. The adversary cannot generate valid reconstructed noise for any other initial noise index j≠ij\neq i. This security guarantee is ensured by two properties of hash: a) Second pre-image resistance: Given (i∗,s)(i^{*},s), it’s computationally infeasible to find (i′,s)(i^{\prime},s) where i′≠i∗i^{\prime}\neq i^{*} such that hash(i∗,s)=(i^{*},s)= hash(i′,s)(i^{\prime},s). b) Collision resistance: It’s computationally infeasible to find any two distinct inputs that hash to the same output. These properties ensure that the adversary cannot find alternative inputs that produce the same hash output, and thus cannot generate valid reconstructed noise for different index numbers jj. ∎

Theorem 4.1 leaves open the possibility that an adversary can recover the noise from a watermarked image and use that noise to forge a new watermarked image. However, empirically we show that this attack fails without access to the weights of the private diffusion model.

Appendix F Further Discussion on Distortion

Using the same initial noise for multiple generations is not distortion-free when examining groups of images. For example, all images with the same prompt pp and the same initial noise zz will be identical, distorted away from the distribution of groups of images generated with i.i.d noises. Luckily, the huge gap between the similarities distribution of (i) reconstructed vs. used noise and (ii) reconstructed vs. another noise, allows us to use as many different noise patterns, while still keeping the noise we used distinguishable more similar to the reconstructed noises. Therefore, limiting the level of distortion in practice.

Appendix G Number of Groups

In our framework, we divide the initial noises into NN groups and associate a Tree-Ring-type key with each group. The use of Fourier Pattern keys enables robustness against rotation, and grouping reduces the search space for inverted noise.

To investigate the impact of the number of groups, we performed an experiment with 10,00010,000 noises and varied the number of groups from 32 to 20482048. As expected, Figure 8 demonstrates that increasing the number of groups leads to better accuracy in detecting the correct initial noise. This is because a larger number of groups results in fewer noises per group, which facilitates more accurate detection. Detailed results for each number of groups under transformation attacks are reported in Table 13.

Appendix H Empirical Runtime Analysis

However, the runtime is highly sensitive to the available computational resources. To provide a practical estimate, we measured the detection time using a single NVIDIA GeForce RTX 3090. Specifically, we divided 100,000 initial noise samples into 32 groups and reported the detection. Under these conditions, the detection phase for 100,000 noise samples takes approximately 22 seconds per detection. We include a comparison with other methods in Table 12.

Appendix I Implementation Details

For all evaluations we used the set of prompts taken from Gustavosta (2024).

Threshold for Detection.

General Retrieval Details.

We included simple rotation (using intervals of 22 degrees) and sliding window (window size of 32, stride of 8) searches as part of the retrieval process. These searches do not involve directly optimizing for the specific degrees of rotation or cropping encountered, ensuring that robustness remains intrinsic to the method.

Appendix J Additional Qualitative Results