RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification
Hai Ci, Pei Yang, Yiren Song, Mike Zheng Shou
Introduction
With the advancement and popularity of diffusion models, a vast number of high-quality diffusion generated images are circulating on the internet. It has become increasingly important to identify AI-generated images to detect fabrication and to track and protect copyrights. Luckily, watermarking offers a reliable solution.
Research on watermarking has a long history in computer vision . A series of works imprint imperceptible watermarks on images through small pixel perturbations. Recently, demonstrates that pixel-level perturbations are provably removable by regeneration attacks. In contrast, Tree-Ring Watermarking has recently introduced a novel approach, proposing to imprint a specially designed tree-ring pattern into the initial noise of a diffusion model. Different patterns are also called different "keys". The verification of the watermark’s existence involves recovering the initial diffusion noise from the image and comparing the extracted key with the injected one. Tree-Ring is considered a semantic watermark since the watermark changes the layout of the generated image and semantically hidden in it. It has demonstrated strong robustness against regeneration attacks and various other image transformations , credited to the meticulously designed tree-ring pattern. However, previous research only studies its performance in the verification task, which aims at distinguishing between watermarked and unwatermarked images. It’s still unclear if Tree-Ring can be used to distribute multiple distinct keys and distinguish between them, a.k.a. watermark identification, which is crucial for source tracing and user attribution. Will Tree-Ring still be strong and robust as in the verification task? In this work, we will explore and answer this question.
We start by studying the source of Tree-Ring’s robustness under verification setting. For the first time, we uncover that the robustness of Tree-Ring stems not only from the carefully designed tree-ring pattern but also from the unintentionally introduced distribution shift during watermark imprinting. Such shift makes a big difference, especially in its robustness to Rotation and Crop/Scale image transformations, which cannot be handled by tree-ring’s pattern design. More importantly, this overlooked power does not help in the identification scenario, because different watermarks experience the same shift.
We further evaluate Tree-Ring in the identification task and surprisingly find that Tree-Ring struggles to distinguish between even dozens of different keys. Identification accuracy is reported in Tab. 2. The disappearance of the extra power of distribution shift has clearly exposed more design flaws in Tree-Ring. First, we find that tree-ring patterns are not discriminative enough to distinguish between different keys. This results in Tree-Ring’s bad accuracy under various attacks or even under the clean setting. Meanwhile, the defective imprinting process leads to the injection of broken watermarks, further weakening its distinguishing ability. In addition, without the help of distribution shift, Tree-Ring is completely incapable of handling rotation and crop/scale transformations. Identification accuracy under these attacks is close to zero.
In order to address the above issues, we propose RingID for enhanced identification capability. It bases on a novel Multi-Channel Heterogeneous watermarking framework, which can amalgamate distinctive advantages from different types of watermarks to resist to various kinds of attacks. Coupled with discretization and lossless imprinting methods, RingID gains substantial improvement on distinguishing ability. What’s more, for the specific attack such as rotation, we locate the flaws in Tree-Ring’s original design and propose systematic solutions, which improves the identification accuracy under rotation from nearly 0 to 0.86.
For the first time, we elucidate how the unintentional introduction of distribution shift substantially contributes to Tree-Ring’s robustness. We demonstrate this effect from both mathematical and empirical perspectives.
We systematically explore Tree-Ring’s performance in the identification scenario and reveal its limitations through comprehensive analysis.
We present RingID, a systematic solution that significantly enhances watermark verification and identification, delivering impressive results. The verification AUC improves from 0.975 to 0.995, while the identification accuracy rises from 0.07 to 0.82.
Related Work
Diffusion models signify a swiftly advancing category of generative models. They define a forward diffusion process that gradually add small amount of Gaussian noise to the input data and learn to recover the clean data from the distortion . Diffusion models initially demonstrate huge potential in generating high-resolution images , resulting in globally recognized products like Midjourney , DALL-E and StableDiffusion . Thereafter, they are broadly applied to modeling various types of data, e.g. video , audio , 3d object , humans , etc.
2 Watermarking Diffusion Models
When it comes to watermarking diffusion models, there are two different types of objectives: watermarking the model weights or the generated content . Model weights watermarking methods aim to protect the intellectual property of the model itself and identify whether another model is an instance of plagiarism. While content watermarking methods protect the copyright of generated images or videos by adding imperceptible perturbations onto pixels. They either imprint the watermark into the frequency via DFT , DCT and wavelet transformations , or directly onto pixels via an encoder network . Recently, Tree-Ring proposes to imprint a tree-ring watermark pattern into the initial diffusion noise. It demonstrates strong robustness to various attacks. In this work, we delve into Tree-Ring, uncovering its source of power in watermark verification and pinpointing its vulnerability in identifying multiple keys. Subsequently, we introduce RingID, a systematic solution fortified with enhanced identification capability.
Preliminaries
Model Owner: The owner of the diffusion model provides the image generation service through API access. He would like to imprint imperceptible watermarks into every generated image for copyright protection and source tracking. Given an arbitrary image, the model owner would like to check whether the image contains his own watermark (a.k.a. Watermark Verification) and which watermark it is if he distributes multiple different watermarks (a.k.a. Identification). The goal of the model owner is to ensure high verification and identification accuracy regardless of any image transformations.
Attacker: The attacker uses the model owner’s service to generate an image. He attempts to apply various image transformations such as rotation and JPEG compression to disrupt the watermark, in order to evade the detection and identification check of the model owner. Subsequently, he can claim ownership of the image and profit from it illegally
2 Notation
3 Task Formulation: Verification vs. Identification
When a set of different keys are distributed, the goal of watermark identification is to find the best match for the recovered key among all candidates in , i.e. . In practice, different users are usually assigned different keys. Thus, each generated image can be source tracked and attributed to a unique user via watermark identification.
Delve into Tree-Ring Watermarking
In this section, we will reveal why Tree-Ring is robust to various image transformations, even some of them are not considered by design. We further evaluate Tree-Ring in the multi-key identification task and pinpoint vulnerabilities in its original design.
2 Distribution Shift in Watermarking
Mathematically, we prove that the shift factor is in a general and simplified setting. Let denote the recovered watermark that has never been shifted, thus we have the following:
Detailed derivation is elaborated in Appendix A Actually, we find any operation leading to a sufficient distribution shift can be considered a good stand-alone watermarking approach. We demonstrate the operation of discarding the imaginary part performs well on its own in Appendix B.
Fig. 3 illustrates the effect of distribution shift under different tasks. In Verification, it helps distinguish between watermarked and non-watermarked images. Next, we go one step further and consider the situation when the image undergoes different attacks like JPEG or Resize. If the tree-ring pattern itself can cope with a certain attack, the improvement brought by the distribution shift will be small, because the two distributions are already separable even without the help of shift. On the contrary, if the tree-ring pattern itself cannot deal with a certain attack, the improvement brought by the distribution shift will be significant. In our ablation (Tab. 7), we find distribution shift helps a lot under Rotation and Crop/Scale attacks. This implies that the original design of tree-ring cannot handle these two attacks. When it comes to Identification, things are a little different. Since different watermarks experience the same distribution shift, discarding the imaginary part doesn’t help distinguish them.
This sparks our curiosity about Tree-Ring’s performance in identification, since distribution shift doesn’t help and identification is more difficult. We would like to know if Tree-Ring still works under Rotation and Crop/Scale attacks. In the next section, we will explore these questions.
3 Is Tree-Ring Good at Identification?
We report the performance of Tree-Ring when identifying different numbers of keys in Tab. 2. Tree-Ring struggles to identify as few as 32 keys. What’s more, identification accuracy is nearly zero under Rotation and Crop/Scale attacks.
We summarize 2 reasons as follows: (1) Distribution shift doesn’t help. As tree-ring watermark cannot handle Rotation and Crop/Scale attacks, identification accuracy under these attacks drops to nearly zero. (2) Tee-ring pattern doesn’t have enough distinguishing ability. Since identification is more difficult than verification, poor distinguishing ability results in poor performance under most attacks or even under no attack setting.
4 Why is Tree-Ring Vulnerable to Rotation and Crop/Scale?
All previous experimental results and analysis show one thing, tree-ring pattern cannot handle these two attacks. In this section, we will see what happens.
Scaling an image by a factor of corresponds to compressing the frequency spectrum in both magnitudes and positions: . Cropping an image modulates the frequency magnitudes by a sinc function. It is straightforward that these transformations easily disrupt the tree-ring pattern on frequency domain. As shown in Fig. 4(e), the pattern is completely disrupted.
4.2 Vulnerable to Rotation
Although Tree-Ring is designed as ring patterns to seek rotation robustness, we find the existing design ignores many details, leading to no rotation robustness. In short, a combination of corner cropping when rotating, a lossy injection process, and the rotationally asymmetric shape cause the problem. We will explain in detail and propose solutions in Sec. 5.2.
RingID: Enhanced Watermarking for Identification
In this section, we show how our proposed method RingID advances in identifying multiple distinct keys. RingID originates from Tree-Ring but incorporates systematic enhancements. We introduce a multi-channel heterogeneous watermarking framework, a discretization approach and a method of lossless imprinting to generally improve its distinguishing ability. We also discuss a series of designs regarding stronger rotation robustness and increased capacity.
where denotes the channel index and traverses all watermarked channels . denotes the index of candidate reference watermarks. is a channel-wise normalizing factor.
Eq. 2 bases on the intuition that if a watermark is robust to a certain attack, then its distance from the ground-truth reference is smaller than that of other watermarks that are not robust. Ideally, it enables adaptive selection of the most robust watermark under different attacks. Notably, the total capacity of MCH is the minimum of each watermark. Therefore, it’s recommended to use a watermark pattern with sufficient capacity on each channel. At the same time, in order to reduce the impact on the generation quality, we need to carefully select the type of watermarks for combination. We find the Gaussian noise pattern is a perfect candidate. It has infinite capacity and the same distribution as the initial noise. Meanwhile, it has good robustness to non-geometric attacks and is complementary to the tree-ring pattern. So RingID explores the combination of a Gaussian noise watermark and a tree-ring watermark. Empirical study in Sec. 6.3 demonstrates that this combination is able to perfectly amalgamate unique advantages from the noise and ring watermark.
2 Being Invariant to Rotation
Although the tree-ring pattern is designed to be rotationally symmetric, it exhibits unexpected vulnerability to rotation transformation as shown in Tab. 2. Below, we locate the causes of the problems and propose a series of improvements to solve them.
We visualize the imprinted tree-ring watermark in spatial domain in Fig. 4(c). Due to the properties of FFT, tree-ring patterns in the frequency tends to exhibit concentrated energy in the four corners in the spatial domain, rendering them vulnerable to cropping during rotation. We propose to circularly shift the watermark pattern in the spatial domain to the center of the image, thus the majority of watermark information is kept while rotating. For an latent image, we circularly shift the pattern by pixels in both width and height dimension. This is equivalent to directly multiplying the frequency watermark with a chessboard pattern given by . The change of the watermark can be clearly observed in Fig. 4(c). Although spatial-shifting helps a lot for rotation invariance, we noticed that this often leads to the generation of circular artifacts at the center of the image. Keeping generation quality in mind, we further multiply the shifted pattern by a factor to suppress the peak at the center. trades-off the robustness, as smaller weakens the watermark pattern. In practice, we find strikes a good balance between image quality and robustness.
2.2 Lossless Imprinting
2.3 Make A Rounder Ring
Rings are originally determined by locating pixels equidistant from a give circle center : . We observed that this simple approach does not yield a well-rounded circular ring, particularly in images with lower resolutions like initial noise of shape . Artifacts and alias are shown in Fig. 4(b). Although it is centrosymmetric, it lacks rotational symmetry. Well, the solution is quite straightforward: drawing a rounder ring. There are a number of tools that can help draw a much rounder ring. Here, we employ a very simple idea to get a rounder ring. We place a white pixel on a black background at a distance from the rotation center. By rotating the low resolution image 360 degrees and recording the trajectory of this pixel, we obtain a rounder ring with a radius of . Despite its simplicity, empirical results in Tab. 3 show that it is crucial for rotational invariance.
3 Discretization for Enhanced Distinguishability
Tree-Ring samples values for each ring from a Gaussian distribution to keep consistency with the distribution of initial noise. We find that this random sampling makes it exceedingly challenging to distinguish between two keys. To address this issue, we propose a discretization approach, specifying that the value for each ring can only be either or , illustrated in Fig. 4(a). Consequently, the watermark pattern for rings has a capacity of . While this seemingly reduces the upper limit of watermark capacity, it significantly enhances its valid capacity, ensuring optimal utilization of all available slots. Our experiments reveal that setting as the standard deviation of the initial noise ensures substantial distinguishablity between different keys while reducing the impact on generation quality to a small level.
4 Increasing Capacity
We consider two simple approaches to increase capacity: adjusting the number of rings in a single channel or imprinting rings in extra channels. We evaluate the advantages and drawbacks of both methods in Sec. 6.4. In summary, increasing the number of rings in a single channel can effectively boost capacity, leading to keys. However, excessive ring count results in a noticeable decline in robustness and generation quality. Imprinting watermarks in multiple channels exponentially increases capacity by , where is number of channels and is the capacity of a single channel. However, it’s likely to introduce ring-like artifacts in the center of generated images. Therefore, we opt to place the watermark in a single channel and enhance capacity by adjusting the range of rings.
Experiment
Following Tree-Ring , we experiment with the state-of-the-art opensource diffusion model StableDiffusion-V2 . By default, RingID imprints a tree-ring watermark with radius 3-14 on the channel 3 of the initial noise and a Gaussian noise watermark on channels 0. We adopt discretization (to 64 and -64), spatial shift, rounder drawing, and lossless imprinting as ordered in Fig. 2. We report the ROC-AUC between 1000 watermarked and 1000 unwatermarked images as the verification metric. In identification, we report identification accuracy with different numbers of keys. Under both verification and identification settings, the evaluations are performed across various image distortions, including rotation, JPEG compression, cropping and scaling (C&S), blurring, noising, and brightness adjustment. We assess the generation quality with CLIP scores by OpenCLIP-ViT/G between the generated image and the prompts to measure text-image alignment, and Frechet Inception Distance (FID) to quantify the similarity between generated and natural images. we generate 5000 images and evaluate FID on the COCO 2017 training dataset . We set both generation and inversion steps to 50.
2 Comparison with Baselines
In verification (Tab. 1), both RingID and Tree-Ring perfectly distinguish between watermarked and unwatermarked images when no attacks are performed, achieving AUC score of 1. However, RingID exhibit stronger robustness to various image distortions than Tree-Ring, achieving near-perfect average AUC score, even though it is not optimized for verification. In the more challenging identification task, (Tab. 2), RingID significantly outperforms Tree-Ring in all settings. When the number of keys to identify is 32, Tree-Ring’s average identification accuracy (excluding C&S) is 0.465, while RingID’s accuracy reaches up to 0.992. As the number of keys increases to 2048, this gap widens drastically (0.077 v.s. 0.942), further highlighting RingID’s superior robustness and larger valid key capacity.
Notably, both methods struggle with the C&S distortion in identification (Tab. 2). This aligns with our analysis in Sec. 4.2. Scaling in the spatial domain causes inverse scaling in the frequency domain, leading to pattern disruption and rendering identification challenging under such distortions. However, in verification, with the help of distribution shift, watermarked images can still be effectively distinguished from unwatermarked ones under C&S. Both Tree-Ring and RingID achieve average AUC higher than 0.97.
2.2 Image Quality
RingID achieves strong performance without compromising the generation quality. It maintains comparable text-image alignments with Tree-Ring. Both methods achieve similar CLIP scores (0.364 v.s. 0.365). RingID obtains an FID of 26.13, which is also comparable to Tree-Ring’s FID of 25.93.
3 Ablation Study
We study the effectiveness of all proposed components of RingID. From Tab. 3, we can find that all designs are crucial for the final success of RingID. Spatial shift, lossless imprinting, and discretization play a significant role in the robustness of rotation. While the use of rounder rings also increase the accuracy by approximately 40% (from 0.620 to 0.860). Notably, spatial shift sacrifices the robustness to blur and brightness attacks for rotation robustness.
3.2 Heterogeneous Watermarking
RingID employs the combination of a Gaussian noise watermark and a ring watermark. Here, we ablate the choice of imprinting the noise watermark on different channels. By comparing the first 3 rows of Tab. 4, we find that heterogeneous watermarking can effectively combine the advantage from both the noise watermark and the ring watermark. Among single-channel configurations, imprinting noise watermark on channel 0 (0.819) achieves the best performance. The combination of channels 0 and 1 yields the highest accuracy among all trials. Interestingly, imprinting noise watermarks on all available channels (0, 1, 2) decrease the accuracy from 0.830 to 0.807.
4 Discussion
In this section, we will discuss where is the best place to put rings. We could imprint all rings on one channel or distribute them evenly across multiple channels. So what is the better practice? In Tab. 5, we fix the total number of rings to 12 and distribute them to 1-4 channels. we can find that performance drops as we distribute rings to more channels. What’s more, imprinting rings on multiple channels results in a greater likelihood of generating artifacts (shown in Appendix C). So given fixed capacity, it’s better to imprint all rings on a single channel.
When we put all rings on a single channel, what’s the best inner radius and outer radius? From Tab. 6 we can see that 3-14 forms a good choice for 11 rings with capacity 2048. The optimal ring radius also varies for different ring numbers.
4.2 More Keys
We would like to explore whether RingID can handle a larger number of keys. Fig. 5 shows the average identification accuracy under various attacks v.s. number of keys. We observe that RingID maintains good accuracy even with a large number of keys. And identification accuracy linearly decreases as the number of keys exponentially increase.
5 Qualitative Results
As shown in Fig. 6, both RingID and Tree-Ring watermarking can generate images of good visual quality. Note that unlike low-level watermarking techniques, both RingID and Tree-Ring cause subtle changes in the image layout. However, RingID generates images with layout that is closer to the unwatermarked image. We owe this to avoiding watermarking on the low-frequency part (radius 0-3).
Conclusion and Limitation
In this work, we conduct a comprehensive investigation on Tree-Ring Watermarking. We reveal that the distribution shift introduced in the watermarking process contributes to the detection of the watermark. Further, we systematically evaluate Tree-Ring Watermarking in the identification task and pinpoint its limitations. Based on the findings and analysis, we propose RingID, which demonstrates enhanced identification capability.
Limitation: Under multiple-key identifications, one limitation of both Tree-Ring and our RingID is the vulnerability to cropping and scaling attacks due to ring pattern disruption. Future work may explore other transform domains to resist these transformations.
Ethics: Like any watermarking methods, RingID could be misused by malicious actors or authoritarian regimes to track, monitor, or persecute individuals based on their creative outputs, which could lead to infringements of privacy.
References
Appendix A Distribution Shift in Watermarking
We following the notation convention of , representing 2D spatial domain signals using lower case letters indexed by (e.g. ) and 2D frequency domain signals using upper case letters indexed by (e.g. ). We use and to denote DFT and inverse DFT, respectively. The energy of a signal is defined as
A signal can be represented as the sum of a real-valued signal and a complex-valued signal, . It can also be represented as the sum of a conjugate symmetric signal and a conjugate asymmetric signal, . If a spatial domain signal is real-valued, then its frequency-domain counterpart is conjugate symmetric :
where is the conjugate of .
A.2 Tree-Ring’s Pipeline
As visualized in Figure Fig. 7, Tree-Ring proposes to add a watermark to a frequency domain initial latent noise (of size ) by substituting the value of pixels within a watermark region mask . For these watermarked pixels (corresponds with used in the paper main content), Tree-Ring samples their values from a circularly-symmetric complex normal distribution . Other non-watermarked pixels , individually, also follow this distribution, but the difference is that in Tree-Ring’s context, are spatially ensured to be conjugate symmetric , while does not have such a guarantee. During the analysis, we view each pixel as a random variable.
After adding the watermark to , Tree-Ring transforms it back to the spatial domain, . The obtained would typically have both the real and imaginary parts. However, since diffusion denoising would always start with a purely-real noise signal, Tree-Ring discards the imaginary part of , turning it into , whose frequency domain counterpart is the conjugate symmetric part of the original watermarked signal, .
A.3 Distribution Shift in Watermarked Region
Eq. 4 implies that the frequency domain signal satisfies:
For pixels within the watermark region , and are uncorrelated, and they both follow a circularly-symmetric complex normal distribution . Viewing as a random variable, its distribution is given by
Oppositely, pixels outside the watermark region are conjugate symmetric, which implies that . The distribution of non-watermarked pixels are unchanged.
This shows that discarding the imaginary part from the spatial domain pixels changes the distribution of the frequency domain pixels within the watermark region mask from to .
A.4 From Energy’s Viewpoint
Discarding the imaginary part of also causes watermarked region to lose half of its energy. Since , . Therefore, the expected energy of is given by
Similarly, for . Therefore, the expected energy of is given by
The fraction of energy of , compared to , is given by
So the watermarked region loses half of its energy.
: Recovered watermark that never experience imaginary part discarding.
: Recovered watermark that experienced imaginary part discarding.
: Null watermark recovered from unwatermarked images.
: The reference watermark to imprint.
Since and are both Gaussian, their combination is also Gaussian with summed variance:
And similarly for the imaginary parts. Therefore,
where the pdf of a -distributed random variable is given by:
A.6 Distribution Shift in Real Scenarios
In Control 1, we follow the original setup of Tree-Ring . The original setup aims at distinguishing between the recovered watermark (shifted) and the null watermark (not shifted) recovered from the unwatermarked images, i.e. v.s. .
In Control 2, we shift null watermark to the same extent as the operation of discarding imaginary part does. Then we get a shifted null watermark . We distinguish between and , i.e. v.s. . Here, both and are shifted to the same extent, so we eliminate the help of distribution shift.
We compare the results of Control 1 and 2 in Tab. 7. We can find general performance drop under all attacks in Control 2. The average AUC decreases from 0.975 to 0.913, making it more challenging to distinguish watermarked and non-watermarked images without the help of distribution shift. Further observation reveals that the major drop occurs in Rotation and Crop & Scale attacks. This result aligns with our identification results Tab. 2 in the main text. indicating that distribution shift contributes substantially to the robustness to Rotate and Crop & Scale attacks. This also implies that the original tree-ring watermark pattern cannot handle these attacks.
Appendix B Discarding Imaginary Part As Standalone Watermarking Approach
We demonstrate the operation of discarding the imaginary part can be used as a standalone watermarking approach. Concretely, for each initial noise instance intended for watermarking, rather than injecting a ring watermark into the frequency spectrum , we opt to inject a random Gaussian noise into the same region. Note that the injected noise are i.i.d. sampled for each case thus different from each other. Although originating from the same distribution, the newly introduced Gaussian noise typically lacks conjugate symmetry. Consequently, when transforming to the spatial domain, it necessitates discarding the excess imaginary parts. This results in a distribution shift in watermarked noise, thus the actually injected noise is from .
Relation with Tree-Ringrand provides a variant called Tree-Ringrand that also injects noise as the watermark. However, they inject the same noise pattern for all generated images and intend to rely on pattern matching for watermark verification. The proposed method in this section distinguishes itself from Tree-Ringrand by injecting i.i.d. sampled noise for each generated image. The AUC for Tree-Ringrand and ours is 0.918 and 0.901, respectively. The closely matched performances indicate that the deviation introduced by discarding imaginary part offers very robust discriminative power. This suggests that discarding imaginary part can effectively distinguish between watermarked and non-watermarked images even without relying on the specific noise pattern.
Appendix C Failure Cases of Multi-Channel Rings
As discussed in the main text, we can imprint the ring watermarks onto multiple channels to trivially increase the capacity. However, we find that this often leads to the generation of ring-like artifacts, evident in a substantial proportion of cases, illustrated in Fig. 8. So we only imprint the ring watermark on a single channel by default.
Appendix D Results on More Diffusion Models
RingID is a universal method that can be applied to different diffusion models. Tab. 10 shows the results on more diffusion models. It is worth noting that the performance of RingID gradually improves from the older version of SD to the newer version of SD.