The Stable Signature: Rooting Watermarks in Latent Diffusion Models

Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, Teddy Furon

Introduction

Recent progress in generative modeling and natural language processing enable easy creation and manipulation of photo-realistic images, such as with DALL·E 2 or Stable Diffusion . They have given birth to many image edition tools like ControlNet , Instruct-Pix2Pix , and others , that are establishing themselves as creative tools for artists, designers, and the general public.

While this is a great step forward for generative AI, it also renews concerns about undermining confidence in the authenticity or veracity of photo-realistic images. Indeed, methods to convincingly augment photo-realistic images have existed for a while, but generative AI significantly lowers the barriers to convincing synthetic image generation and edition (e.g. a generated picture recently won an art competition ). Not being able to identify that images are generated by AI makes it difficult to remove them from certain platforms and to ensure their compliance with ethical standards. It opens the doors to new risks like deep fakes, impersonation or copyright usurpation .

A baseline solution to identify generated images is forensics, i.e. passive methods to detect generated/manipulated images. On the other hand, existing watermarking methods can be added on top of image generation. They are based on the idea of invisibly embedding a secret message into the image, which can then be extracted and used to identify the image. This has several drawbacks. If the model leaks or is open-sourced, the post-generation watermarking is easy to remove. The open source Stable Diffusion is a case in point, since removing the watermark amounts to commenting out a single line in the source code.

Our Stable Signature method merges watermarking into the generation process itself, without any architectural changes. It adjusts the pre-trained generative model such that all the images it produces conceal a given watermark. There are several advantages to this approach . It protects both the generator and its productions. Besides, it does not require additional processing of the generated image, which makes the watermarking computationally lighter, straightforward, and secure. Model providers would then be able to deploy their models to different user groups with a unique watermark, and monitor that they are used in a responsible manner. They could also give art platforms, news outlets and other sharing platforms the ability to detect when an image has been generated by their AI.

We focus on Latent Diffusion Models (LDM) since they can perform a wide range of generative tasks. This work shows that simply fine-tuning a small part of the generative model – the decoder that generates images from the latent vectors – is enough to natively embed a watermark into all generated images. Stable Signature does not require any architectural change and does not modify the diffusion process. Hence it is compatible with most of the LDM-based generative methods . The fine-tuning stage is performed by back-propagating a combination of a perceptual image loss and the hidden message loss from a watermark extractor back to the LDM decoder. We pre-train the extractor with a simplified version of the deep watermarking method HiDDeN .

We create an evaluation benchmark close to real world situations where images may be edited. The tasks are: detection of AI generated images, tracing models from their generations. For instance, we detect 90%90\% of images generated with the generative model, even if they are cropped to 10%10\% of their original size, while flagging only one false positive every 10610^{6} images. To ensure that the model’s utility is not weakened, we show that the FID score of the generation is not affected and that the generated images are perceptually indistinguishable from the ones produced by the original model. This is done over several tasks involving LDM (text-to-image, inpainting, edition, etc.).

As a summary, (1) we efficiently merge watermarking into the generation process of LDMs, in a way that is compatible with most of the LDM-based generative methods; (2) we demonstrate how it can be used to detect and trace generated images, through a real-world evaluation benchmark; (3) we compare to post-hoc watermarking methods, showing that it is competitive while being more secure and efficient, and (4) evaluate robustness to intentional attacks.

Related Work

has long been dominated by GANs, still state-of-the-art on many datasets . Transformers have also been successfully used for modeling image or video distributions, providing higher diversity at the expense of increased inference time. Images are typically converted to token lists using vector-quantized architectures , relying on an image decoder.

Diffusion models have brought huge improvements in text-conditional image generation, being now able to synthesize high-resolution photo-realistic images for a wide variety of text prompts . They can also perform conditional image generation tasks – like inpainting or text-guided image editing – by fine-tuning the diffusion model with additional conditioning, e.g. masked input image, segmentation map, etc. . Because of their iterative denoising algorithm, diffusion models can also be adapted for image editing in a zero-shot fashion by guiding the generative process . All these methods, when applied on top of Stable Diffusion, operate in the latent space of images, requiring a latent decoder to produce an RGB image.

Detection of AI-generated/manipulated images

is notably active in the context of deep-fakes . Many works focus on the detection of GAN-generated images . One way is to detect inconsistencies in the generated images, via lights, perspective or physical objects . These approaches are restricted to photo-realistic images or faces but do not cover artworks where objects are not necessarily physically correct.

Other approaches use traces left by the generators in the spatial or frequency domains. They have extended to diffusion models in recent works , and showed encouraging results. However purely relying on forensics and passive detection is limiting. As an example, the best performing method to our knowledge is able to detect 50%50\% of generated images for an FPR around 11/100100. Put differently, if a user-generated content platform were to receive 11 billion images every day, it would need to wrongly flag 1010 million images to detect only half of the generated images. Besides, passive techniques cannot trace images from different versions of the same model, conversely to active ones like watermarking.

Image watermarking

has long been studied in the context of tracing and intellectual property protection . More recently, deep learning encoder/extractor alternatives like HiDDeN or iterative methods by Vukotić et al. showed competitive results in terms of robustness to a wide range of transformations, namely geometric ones.

In the specific case of generative models, some works deal with watemarking the training set on which the generative model is learned . It is highly inefficient since every new message to be embedded requires a new training pipeline. Merging the watermarking and the generative process is a recent idea , that is closer to the model watermarking litterature . They suffer from two strong limitations. First, these methods only apply to GAN, while LDM are beginning to replace them in most applications. Second, watermarking is incorporated in the training process of the GAN from the start. This strategy is very risky because the generative model training is more and more costlyStable Diffusion training costs ∼\sim$600k of cloud compute (Wikipedia).. Our work shows that a quick fine-tuning of the latent decoder part of the generative model is enough to achieve a good watermarking performance, provided that the watermark extractor is well chosen.

Problem Statement & Background

Figure 1 shows a model provider Alice who deploys a latent diffusion model to users Bobs. Stable Signature embeds a binary signature into the generated images. This section derives how Alice can use this signature for two scenarios:

Detection: “Is it generated by my model?”. Alice detects if an image was generated by her model. As many generations as possible should be flagged, while controlling the probability of flagging a natural image.

Identification: “Who generated this image?”. Alice monitors who created each image, while avoiding to mistakenly identifying a Bob who did not generate the image.

Alice embeds a kk-bit binary signature into the generated images. The watermark extractor then decodes messages from the images it receives and detects when the message is close to Alice’s signature. An example application is to block AI-generated images on a content sharing platform.

Let m∈{0,1}km\in\{0,1\}^{k} be Alice’s signature. We extract the message m′m^{\prime} from an image xx and compare it to mm. As done in previous works , the detection test relies on the number of matching bits M(m,m′)M(m,m^{\prime}): if

then the image is flagged. This provides a level of robustness to imperfections of the watermarking.

Formally, we test the statistical hypothesis H1H_{1}: “xx was generated by Alice’s model” against the null hypothesis H0H_{0}: “xx was not generated by Alice’s model”. Under H0H_{0} (i.e. for vanilla images), we assume that bits m1′,…,mk′m^{\prime}_{1},\ldots,m^{\prime}_{k} are (i.i.d.) Bernoulli random variables with parameter 0.50.5. Then M(m,m′)M(m,m^{\prime}) follows a binomial distribution with parameters (kk, 0.50.5). We verify this assumption experimentally in App. B.5. The False Positive Rate (FPR) is the probability that M(m,m′)M(m,m^{\prime}) takes a value bigger than the threshold τ\tau. It is obtained from the CDF of the binomial distribution, and a closed-form can be written with the regularized incomplete beta function Ix(a;b)I_{x}(a;b):

2 Image watermarking for identification

Alice now embeds a signature m(i)m^{(i)} drawn randomly from {0,1}k\{0,1\}^{k} into the model distributed to Bob(i) (for i=1⋯Ni=1\cdots N, with NN the number of Bobs). Alice can trace any misuse of her model: generated images violating her policy (gore content, deepfakes) are linked back to the specific Bob by comparing the extracted message to Bobs’ signatures.

We compare the message m′m^{\prime} from the watermark extractor to (m(1),…,m(N))\left(m^{(1)},\dots,m^{(N)}\right). There are now NN detection hypotheses to test. If the NN hypotheses are rejected, we conclude that the image was not generated by any of the models. Otherwise, we attribute the image to argmaxi=1..NM(m′,m(i))\textrm{argmax}_{i=1..N}M\left(m^{\prime},m^{(i)}\right). With regards to the detection task, false positives are more likely since there are NN tests. The global FPR at a given threshold τ\tau is:

Equation (3) (resp. (2)), is used reversely: we find threshold τ\tau to achieve a required FPR for identification (resp. detection). Note that these formulae hold only under the assumption of i.i.d. Bernoulli bits extracted from vanilla images. This crucial point is enforced in the next section.

Method

Stable Signature modifies the generative network so that the generated images have a given signature through a fixed watermark extractor. It is trained in two phases. First, we create the watermark extractor network W\mathcal{W}. Then, we fine-tune the Latent Diffusion Model (LDM) decoder D\mathcal{D}, such that all generated images have a given signature through W\mathcal{W}.

We use HiDDeN , a classical method in the deep watermarking literature. It jointly optimizes the parameters of watermark encoder WE\mathcal{W}_{E} and extractor network W\mathcal{W} to embed kk-bit messages into images, robustly to transformations that are applied during training. We discard WE\mathcal{W}_{E} after training, since only W\mathcal{W} serves our purpose.

The network architectures are kept simple to ease the LDM fine-tuning in the second phase. They are the same as HiDDeN (see App. A.1) with two changes.

Second, we observed that W\mathcal{W}’s output bits for vanilla images are correlated and highly biased, which violates the assumptions of Sec. 3.1. Therefore we remove the bias and decorrelate the outputs of W\mathcal{W} by applying a PCA whitening transformation (more details in App. A.1).

2 Fine-tuning the generative model

In LDM, the diffusion happens in the latent space of an auto-encoder. The latent vector zz obtained at the end of the diffusion is input to decoder D\mathcal{D} to produce an image. Here we fine-tune D\mathcal{D} such that the image contains a given message mm that can be extracted by W\mathcal{W}. Stable Signature is compatible with many generative tasks, since modifying only D\mathcal{D} does not affect the diffusion process.

First, we fix the signature m=(m1,…,mk)∈{0,1}km=(m_{1},\ldots,m_{k})\in\{0,1\}^{k}. The fine-tuning of D\mathcal{D} into Dm\mathcal{D}_{m} is inspired by the original training of the auto-encoder in LDM .

The weights of Dm\mathcal{D}_{m} are optimized in a few backpropagation steps to minimize

Text-to-Image Watermarking Performance

This section shows the potential of our method for detection and identification or images generated by a Stable-Diffusion-like model We refrain from experimenting with pre-existing third-party generative models, such as Stable Diffusion or LDMs, and instead use a large diffusion model (2.2B parameters) trained on an internal dataset of 330M licensed image-text pairs. . We apply generative models watermarked with 4848-bit signatures on prompts of the MS-COCO validation set. We evaluate detection and identification on the outputs, as illustrated in Figure 1.

We evaluate their robustness to different transformations applied to generated images: strong cropping (10%10\% of the image remaining), brightness shift (strength factor 2.02.0), as well as a combination of crop 50%50\%, brightness shift 1.51.5 and JPEG 8080. This covers typical geometric and photometric edits (see Fig. 5 for visual examples).

The performance is partly obtained from experiments and partly by extrapolating small-scale measurements.

For detection, we fine-tune the decoder of the LDM with a random key mm, generate 10001000 images and use the test of Eq. (1). We report the tradeoff between True Positive Rate (TPR), i.e. the probability of flagging a generated image and the FPR, while varying τ∈{0,..,48}\tau\in\{0,..,48\}. For instance, for τ=0\tau=0, we flag all images so FPR=1\textrm{FPR}=1, and TPR=1\textrm{TPR}=1. The TPR is measured directly. In contrast the FPR is inferred from Eq. (2), because it would otherwise be too small to be measured on reasonably sized problems (this approximation is validated experimentally in App. B.6). The experiment is run on 1010 random signatures and we report averaged results.

Figure 5 shows the tradeoff under image transformations. For example, when the generated images are not modified, Stable Signature detects 99%99\% of them, while only 11 vanilla image out of 10910^{9} is flagged. At the same FPR=10−9\textrm{FPR}=10^{-9}, Stable Signature detects 84%84\% of generated images for a crop that keeps 10%10\% of the image, and 65%65\% for a transformation that combines a crop, a color shift, and a JPEG compression. For comparison, we report results of a state-of-the-art passive method , applied on resized and compressed images. As to be expected, we observe that these baseline results have orders of magnitudes larger FPR than Stable Signature, which actively marks the content.

2 Identification results

Each Bob has its own copy of the generative model. Given an image, the goal is to find if any of the NN Bobs created it (detection) and if so, which one (identification). There are 33 types of error: false positive: flag a vanilla image; false negative: miss a generated image; false accusation: flag a generated image but identify the wrong user.

For evaluation, we fine-tune N′=1000N^{\prime}=1000 models with random signatures. Each model generates 100100 images. For each of these 100100k watermarked images, we extract the Stable Signature message, compute the matching score with all NN signatures and select the user with the highest score. The image is predicted to be generated by that user if this score is above threshold τ\tau. We determined τ\tau such that FPR=10−6\textrm{FPR}=10^{-6}, see Eq. (3). For example, for N=1N=1, τ=41\tau=41 and for N=1000N=1000, τ=44\tau=44. Accuracy is extrapolated beyond the N′N^{\prime} users by adding additional signatures and having N>N′N>N^{\prime} (e.g. users that have not generated any images).

Figure 5 reports the per-transformation identification accuracy. For example, we identify a user among NN=10510^{5} with 98%98\% accuracy when the image is not modified. Note that for the combined edit, this becomes 40%40\%. This may still be dissuasive: if a user generates 33 images, he will be identified 80%80\% of the time. We observe that at this scale, the false accusation rate is zero, i.e. we never identify the wrong user. This is because τ\tau is set high to avoid FPs, which also makes false accusations unlikely. We observe that the identification accuracy decreases when NN increases, because the threshold τ\tau required to avoid false positives is higher when NN increases, as pointed out by the approximation in (3). In a nutshell, by distributing more models, Alice trades some accuracy of detection against the ability to identify users.

Experimental Results

We presented in the previous section how to leverage watermarks for detection and identification of images generated from text prompts. We now present more general results on robustness and image quality for different generative tasks. We also compare Stable Signature to other watermarking algorithms applied post-generation.

Since our method only involves the LDM decoder, it makes it compatible with many generative tasks. We evaluate text-to-image generation and image edition on the validation set of MS-COCO , super-resolution and inpainting on the validation set of ImageNet (all evaluation details are available in App. A.3).

2 Image generation quality

Figure 6 shows qualitative examples of how the image generation is altered by the latent decoder’s fine-tuning. The difference is very hard to perceive even for a trained eye. This is surprising for such a low PSNR, especially since the watermark embedding is not constrained by any Human Visual System like in professional watermarking techniques. Most interestingly, the LDM decoder has indeed learned to add the watermark signal only over textured areas where the human eyes are not sensitive, while the uniform backgrounds are kept intact (see the pixel-wise difference).

Table 1 presents a quantitative evaluation of image generation quality on the different tasks. We report the FID, and the average PSNR and SSIM that are computed between the images generated by the fine-tuned LDM and the original one. The results show that no matter the task, the watermarking has very small impact on the FID of the generation.

The average PSNR is around 3030 dB and SSIM around 0.90.9 between images generated by the original and a watermarked model. They are a bit low from a watermarking perspective because we do not explicitly optimize for them. Indeed, in a real world scenario, one would only have the watermarked version of the image. Therefore we don’t need to be as close as possible to the original image but only want to generate artifacts-free images. Without access to the image generated by the original LDM, it is very hard to tell whether a watermark is present or not.

3 Watermark robustness

We evaluate the robustness of the watermark to different image transformations applied before extraction. For each task, we generate 11k images with 1010 models fine-tuned for different messages, and report the average bit accuracy in Table 1. Additionally, Table 2 reports results on more image transformations for images generated from COCO prompts. The main evaluated transformations are presented in Fig. 5 (more evaluation details are available in App. A.3).

We see that the watermark is indeed robust for several tasks and across transformations. The bit accuracy is always above 0.90.9, except for inpainting, when replacing only the masked region of the image (between 1−501-50% of the image, with an average of 27%27\% across masks). Besides, the bit accuracy is not perfect even without edition, mainly because there are images that are harder to watermark (e.g. the ones that are very uniform, like the background in Fig. 6) and for which the accuracy is lower.

Note that the robustness comes even without any transformation during the LDM fine-tuning phase: it is due to the watermark extractor. If the watermark embedding pipeline is learned to be robust against an augmentation, then the LDM will learn how to produce watermarks that are robust against it during fine-tuning.

4 Comparison to post-hoc watermarking

An alternative way to watermark generated images is to process them after the generation (post-hoc). This may be simpler, but less secure and efficient than Stable Signature. We compare our method to a frequency based method, DCT-DWT , iterative approaches (SSL Watermark and FNNS ), and an encoder/decoder one like HiDDeN . We choose DCT-DWT since it is employed by the original open source release of Stable Diffusion , and the other methods because of their performance and their ability to handle arbitrary image sizes and number of bits. We use our implementations (see details in App. A.4).

Table 1 compares the generation quality and the robustness over 55k generated images. Overall, Stable Signature achieves comparable results in terms of robustness. HiDDeN’s performance is a bit higher but its output bits are not i.i.d. meaning that it cannot be used with the same guarantees as the other methods. We also observe that post-hoc generation gives worse qualitative results, images tend to present artifacts (see Fig. 13 in the supplement). One explanation is that Stable Signature is merged into the high-quality generation process with the LDM auto-encoder model, which is able to modify images in a more subtle way.

5 Can we trade image quality for robustness?

We can choose to maximize the image quality or the robustness of the watermark thanks to the weight λi\lambda_{i} of the perceptual loss in (4). We report the average PSNR of 11k generated images, as well as the bit accuracy obtained on the extracted message for the ‘Combined’ editing applied before detection (qualitative results are in App. B.1). A higher λi\lambda_{i} leads to an image closer to the original one, but to lower bit accuracies on the extracted message, see Table 3.

6 Attack simulation layer

Watermark robustness against image transformations depends solely on the watermark extractor. here, we pre-train them with or without specific transformations in the simulation layer, on a shorter schedule of 5050 epochs, with 128×128128\times 128 images and 1616-bits messages. From there, we plug them in the LDM fine-tuning stage and we generate 11k images from text prompts. We report the bit accuracy of the extracted watermarks in Table 4. The extractor is naturally robust to some transformations, such as crops or brightness, without being trained with them, while others, like rotations or JPEG, require simulation during training for the watermark to be recovered at test time. Empirically we observed that adding a transformation improves results for the latter, but makes training more challenging.

Attacks on Stable Signature’s Watermarks

We examine the watermark’s resistance to intentional tampering, as opposed to distortions that happen without bad intentions like crops or compression (discussed in Sec. 5). We consider two threat models: one is typical for many image watermarking methods and operates at the image level, and another targets the generative model level. For image-level attacks, we evaluate on 55k images generated from COCO prompts. Full details on the following experiments can be found in Appendix A.5.

Bob alters the image to remove the watermark with deep learning techniques, like methods used for adversarial purification or neural auto-encoders . Note that this kind of attacks has not been explored in the image watermarking literature to our knowledge. Figure 7 evaluates the robustness of the watermark against neural auto-encoders at different compression rates. To reduce the bit accuracy closer to random (50%), the image distortion needs to be strong (PSNR<<26). However, assuming the attack is informed on the generative model, i.e. the auto-encoder is the same as the one used to generate the images, the attack becomes much more effective. It erases the watermark while achieving high quality (PSNR>>29). This is because the image is modified precisely in the bandwidth where the watermark is embedded. Note that this assumption is strong, because Alice does not need to distribute the original generator.

Watermark removal & embedding (white-box).

Instead of removing the watermark, an attacker could embed a signature into vanilla images (unauthorized embedding ) to impersonate another Bob of whom they have a generated image. It highlights the importance of keeping the watermark extractor private.

2 Network-level attacks

Bob gets Alice’s generative model and uses a fine-tuning process akin to Sec. 4.2 to eliminate the watermark embedding – that we coin model purification. This involves removing the message loss Lm\mathcal{L}_{m}, and shifting the focus to the perceptual loss Li\mathcal{L}_{i} between the original image and the one reconstructed by the LDM auto-encoder.

Figure 8 shows the results of this attack for the MSE loss. The PSNR between the watermarked and purified images is plotted at various stages of fine-tuning. Empirically, it is difficult to significantly reduce the bit accuracy without compromising the image quality: artifacts start to appear during the purification.

Model collusion.

This so-called marking assumption plays a crucial role in traitor tracing literature . Surprisingly, it holds even though our watermarking process is not explicitly designed for it. The study has room for improvement, such as creating user identifiers with more powerful traitor tracing codes and using more powerful traitor accusation algorithms . Importantly, we found the precedent remarks also hold if the colluders operate at the image level.

Conclusion & Discussion

By a quick fine-tuning of the decoder of Latent Diffusion Models, we can embed watermarks in all the images they generate. This does not alter the diffusion process, making it compatible with most of LDM-based generative models. These watermarks are robust, invisible to the human eye and can be employed to detect generated images and identify the user that generated it, with very high performance.

The public release of image generative models has an important societal impact. With this work, we put to light the usefulness of using watermarking instead of relying on passive detection methods. We hope it will encourage researchers and practitioners to employ similar approaches before making their models publicly available.

Although the diffusion-based generative model has been trained on an internal dataset of licensed images, we use the KL auto-encoder from LDM with compression factor f=8f=8. This is the one used by open-source alternatives. Code is available at github.com/facebookresearch/stable_signature.

Environmental Impact.

We do not expect any environmental impact specific from this work. The cost of the experiments and the method is high, though order of magnitudes less than other computer vision fields. We roughly estimated that the total GPU-days used for running all our experiments to 20002000, or ≈50000\approx 50000 GPU-hours. This amounts to total emissions in the order of 10 tons of CO2eq. This is excluding the training of the generative model itself, since we did not perform that training. Estimations are conducted using the Machine Learning Impact calculator presented by Lacoste et al. . We do not consider in this approximation: memory storage, CPU-hours, production cost of GPUs/ CPUs, etc.

References

Appendix A Implementation Details & Parameters

We keep the same architecture as in HiDDeN , which is a simple convolutional encoder and extractor. The encoder consist of 44 Conv-BN-ReLU blocks, with 6464 output filters, 3×33\times 3 kernels, stride 11 and padding 11. The extractor has 77 blocks, followed by a block with kk output filters (kk being the number of bits to hide), an average pooling layer, and a k×kk\times k linear layer. For more details, we refer the reader to the original paper .

Optimization.

We train on the MS-COCO dataset , with 256×256256\times 256 images. The number of bits is k=48k=48, and the scaling factor is α=0.3\alpha=0.3. The optimization is carried out for 300300 epochs on 88 GPUs, with the Lamb optimizer (it takes around a day). The learning rate follows a cosine annealing schedule with 55 epochs of linear warmup to 10−210^{-2}, and decays to 10−610^{-6}. The batch size per GPU is 6464.

Attack simulation layer.

Whitening.

At the end of the training, we whiten the output of the watermark extractor to make the hard thresholded bits independently and identically Bernoulli distributed on vanilla images (so that the assumption of 3.1 holds better, see App. B.5). We perform the PCA of the output of the watermark extractor on a set of 1010k vanilla images, and get the mean μ\mu and eigendecomposition of the covariance matrix Σ=UΛUT\Sigma=U\Lambda U^{T}. The whitening is applied with a linear layer with bias −Λ−1/2UTμ-\Lambda^{-1/2}U^{T}\mu and weight Λ−1/2UT\Lambda^{-1/2}U^{T}, appended to the extractor.

A.2 Image transformations

We evaluate the robustness of the watermark to a set of transformations in sections 5, 6 and B.2. They simulate image processing steps that are commonly used in image editing software. We illustrate them in Figure 9. For crop and resize, the parameter is the ratio of the new area to the original area. For rotation, the parameter is the angle in degrees. For JPEG compression, the parameter is the quality factor (in general 90% or higher is considered high quality, 80%-90% is medium, and 70%-80% is low). For brightness, contrast, saturation, and sharpness, the parameter is the default factor used in the PIL and Torchvision libraries. The text overlay is made through the AugLy library , and adds a text at a random position in the image. The combined transformation is a combination of a crop 0.50.5, a brightness change 1.51.5, and a JPEG 8080 compression.

A.3 Generative tasks

In text-to-image generation, the diffusion process is guided by a text prompt. We follow the standard protocol in the literature and evaluate the generation on prompts from the validation set of MS-COCO . To do so, we first retrieve all the captions from the validation set, keep only the first one for each image, and select the first 10001000 or 50005000 captions depending on the evaluation protocol. We use guidance scale 3.03.0 and 5050 diffusion steps. If not specified, the generation is done for 50005000 images. The FID is computed over the validation set of MS-COCO, resized to 512×512512\times 512.

Image edition.

DiffEdit takes as input an image, a text describing the image and a novel description that the edited image should match. First, a mask is computed to identify which regions of the image should be edited. Then, mask-based generation is performed in the latent space, before converting the output back to RGB space with the image decoder. We use the default parameters used in the original paper, with an encoding ratio of 90%, and compute a set of 50005000 images from the COCO dataset, edited with the same prompts as the paper . The FID is computed over the validation set of MS-COCO, resized to 512×512512\times 512.

Inpainting.

We follow the protocol of LaMa , and generate 50005000 masks with the “thick” setting, at resolution 512×512512\times 512, each mask covering 1−50%1-50\% of the initial image (with an average of 27%27\%). For the diffusion-based inpainting, we use the inference-time algorithm presented in , also used in Glide , which corrects intermediate estimations of the final generated image with the ground truth pixel values outside the inpainting mask. For latent diffusion models, the same algorithm can be applied in latent space, by encoding the image to be inpainted and downsampling the inpainting mask. In this case, we consider 22 different variations: (1) inpainting is performed in the latent space and the final image is obtained by simply decoding the latent image; and (2) the same procedure is applied, but after decoding, ground truth pixel values from outside the inpainting mask are copy-pasted from the original image. The latter allows to keep the rest of the image perfectly identical to the original one, at the cost of introducing copy-paste artifacts, visible in the borders. Image quality is measured with an FID score, computed over the validation set of ImageNet , resized to 512×512512\times 512.

Super-resolution.

We follow the protocol suggested by Saharia et al. . We first resize 50005000 random images from the validation set of ImageNet to 128×128128\times 128 using bicubic interpolation, and upscale them to 512×512512\times 512. The FID is computed over the validation set of ImageNet, cropped and resized to 512×512512\times 512.

A.4 Watermarking methods

For Dct-Dwt, we use the implementation of https://github.com/ShieldMnt/invisible-watermark (the one used in Stable Diffusion). For SSL Watermark and FNNS the watermark is embedded by optimizing the image, such that the output of a pre-trained model is close to the given key (like in adversarial examples ). The difference between the two is that in SSL Watermark we use a model pre-trained with DINO , while FNNS uses a watermark or stenography model. For SSL Watermark we use the default pre-trained model of the original paper. For FNNS we use the HiDDeN extractor used in all our experiments, and not SteganoGan as in the original paper, because we want to extract watermarks from images of different sizes. We use the image optimization scheme of Active Indexing , i.e. we optimize the distortion image for 1010 iterations, and modulate it with a perceptual just noticeable difference (JND) mask. This avoids visible artifacts and gives a PSNR comparable with our method (≈30\approx 30dB). For HiDDeN, we use the watermark encoder and extractor from our pre-training phase, but the extractor is not whitened and we modulate the encoder output with the same JND mask. Note that in all cases we watermark images one by one for simplicity. In practice the watermarking could be done by batch, which would be more efficient.

A.5 Attacks

The perceptual auto-encoders aim to create compressed latent representations of images. We select 22 state-of-the-art auto-encoders from the CompressAI library zoo : the factorized prior model and the anchor model variant . We also select the auto-encoders from Esser et al. and Rombach et al. . For all models, we use different compression factors to observe the trade-off between quality degradation and removal robustness. For bmshj2018: 11, 44 and 88, for cheng2020: 11, 33 and 66, for esser2021: VQ-44, 88 and 1616, for rombach2022 KL-44, 88, 1616 and 3232 (KL-88 being the one used by SD v1.4). We generate 11k images from text prompts with our LDM watermarked with a 4848-bits key. We then try to remove the watermark using the auto-encoders, and compute the bit accuracy on the extracted watermark. The PSNR is computed between the original image and the reconstructed one, which explains why the PSNR does not exceed 3030dB (since the watermarked image already has a PNSR of 3030dB). If we compared between the watermarked image and the image reconstructed by the auto-encoder instead, the curves would show the same trend but the PSNR would be 22-33 points higher.

Watermark removal (white-box).

In the white-box case, we assume have access to the extractor model. The adversarial attack is performed by optimizing the image in the same manner as . The objective is a MSE loss between the output of the extractor and a random binary message fixed beforehand. The attack is performed for 1010 iterations with the Adam optimizer with learning rate 0.10.1.

Watermark removal (network-level).

We use the same fine-tuning procedure as in Sec. 4.2. This is done for different numbers of steps, namely 100100, 200200, and every multiple of 200200 up to 16001600. The bit accuracy and the reported PSNR are computed on 11k images of the validation set of COCO, for the auto-encoding task.

Model collusion.

The goal is to observe the decoded watermarks on the generation when 22 models are averaged together. We fine-tune the LDM decoder for 1010 different 4848-bits keys (representing 1010 Bobs). We then randomly sample a pair of Bobs and average the 22 models, with which we generate 100100 images. We then extract the watermark from the generated images and compare them to the 22 original keys. We repeat this experiment 1010 times, meaning that we observe 10×100×48=4800010\times 100\times 48=48000 decoded bits.

In the inline figure, the rightmost skewed normal is fitted with the Scipy library and the corresponding parameters are a:6.96,e:0.06,w:0.38a:6.96,e:0.06,w:0.38. This done over all bits where Bobs both have a 11. The same observation holds when there is no collusion, with approximately the same parameters. When the bit is not the same between Bobs, we denote by m1(i)m_{1}^{(i)} the random variable representing the output of the extractor in the case where the generative model only comes from Bob(i), and by m2m_{2} the random variable representing the output of the extractor in the case where the generative model comes from the average of the two Bobs. Then in our model m2=0.5⋅(m1(i)+m1(j))m_{2}=0.5\cdot(m_{1}^{(i)}+m_{1}^{(j)}), and the pdf of m2m_{2} is the convolution of the pdf of m1(i)m_{1}^{(i)} and the pdf of m1(j)m_{1}^{(j)}, rescaled in the x axis because of the factor 0.50.5.

Appendix B Additional Experiments

The perceptual loss of (4) affects the image quality. Figure 10 shows how the parameter λi\lambda_{i} affects the image quality. For high values, the image quality is very good. For low values, artifacts mainly appear in textured area of the image. It is interesting to note that this begins to be problematic only for low PSNR values (around 25 dB).

Figure 10 shows an example of a watermarked image for different perceptual losses: Watson-VGG , Watson-DFT , LPIPS , MSE, and LPIPS+MSE. We set the weight λi\lambda_{i} of the perceptual loss so that the watermark performance is approximately the same for all types of loss, and such that the degradation of the image quality is strong enough to be seen. Overall, we observe that the Watson-VGG loss gave the most eye-pleasing results, closely followed by the LPIPS. When using the MSE, images are blurry and artifacts appear more easily, even though the PSNR is higher.

B.2 Additional results on watermarks robustness

In Table 5, we report the same table as in Table 1 that evaluates the watermark robustness in bit accuracy on different tasks, with additional image transformations. They are detailed and illustrated in App. A.3. As a reminder, the watermark is a 4848-bit binary key. It is robust to a wide range transformations, and most often yields above 0.90.9 bit accuracy. The resize and JPEG 5050 transformations seems to be the most challenging ones, and sometimes get bellow 0.90.9. Note that the crop location is not important but the visual content of the crop is, e.g. there is no way to decode the watermark on crops of blue sky (this is the reason we only show center crop).

B.3 Additional network level attacks

Tab. 6 reports robustness of the watermarks to different quantization and pruning levels for the LDM decoder. Quantization is performed naively, by rounding the weights to the closest quantized value in the min-max range of every weight matrix. Pruning is done using PyTorch pruning API, with the L1 norm as criterion. We observe that the network generation quality degrades faster than WM robustness. To reduce bit accuracy lower than 98%, quantization degrades the PSNR <<25dB, and pruning <<20dB.

B.4 Scaling factor at pre-training.

The watermark encoder does not need to be perceptually good and it is beneficial to degrade image quality during pre-training. In the following, ablations are conducted on a shorter schedule of 5050 epochs, on 128×128128\times 128 images and 1616-bits messages. In Table 7, we train watermark encoders/extractors for different scaling factor α\alpha (see Sec. 4.1), and observe that α\alpha strongly affects the bit accuracy of the method. When it is too high, the LDM needs to generate low quality images for the same performance because the distortions seen at pre-training by the extractor are too strong. When it is too low, they are not strong enough for the watermarks to be robust: the LDM will learn how to generate watermarked images, but the extractor won’t be able to extract them on edited images.

B.5 Are the decoded bits i.i.d. Bernoulli random variables?

The FPR and the pp-value (2) are computed with the assumption that, for vanilla images (not watermarked), the bits output by the watermark decoder W\mathcal{W} are independent and identically distributed (i.i.d.) Bernoulli random variables with parameter 0.50.5. This assumption is not true in practice, even when we tried using regularizing losses in the training at phase one . This is why we whiten the output at the end of the pre-training.

Figure 12 shows the covariance matrix of the hard bits output by W\mathcal{W} before and after whitening. They are computed over 55k vanilla images, generated with our LDM at resolution 512×512512\times 512 (as a reminder the whitening is performed on 11k vanilla images from COCO at 256×256256\times 256). We compare them to the covariance matrix of a Bernoulli simulation, where we simulate 5k5k random messages of 4848 Bernoulli variables. We observe the strong influence of the whitening on the covariance matrix, although it still differs a little from the Bernoulli simulation. We also compute the bit-wise mean and observe that for un-whitened output bits, some bits are very biased. For instance, before whitening, one bit had an average value of 0.950.95 (meaning that it almost always outputs 11). After whitening, the maximum average value of a bit is 0.680.68. For the sake of comparison, the maximum average value of a bit in the Bernoulli simulation was 0.520.52. It seems to indicate that the distribution of the generated images are different than the one of vanilla images, and that it impacts the output bits. Therefore, the bits are not perfectly i.i.d. Bernoulli random variables. We however found they are close enough for the theoretical FPR computation to match the empirical one (see next section) – which was what we wanted to achieve.

B.6 Empirical check of the FPR

In Figure 5, we plotted the TPR against a theoretical value for the FPR, with the i.i.d. Bernoulli assumption. The FPR was computed theoretically with (2). Here, we empirically check on smaller values of the FPR (up to 10−710^{-7}) that the empirical FPR matches the theoretical one (higher values would be too computationally costly). To do so, we use the 1.41.4 million vanilla images from the training set of ImageNet resized and cropped to 512×512512\times 512, and perform the watermark extraction with W\mathcal{W}. We then fix 1010 random 4848-bits key m(1),⋯ ,m(10)m^{(1)},\cdots,m^{(10)}, and, for each image, we compute the number of matching bits d(m′,m(i))d(m^{\prime},m^{(i)}) between the extracted message m′m^{\prime} and the key m(i)m^{(i)}, and flag the image if d(m′,m(i))≥τd(m^{\prime},m^{(i)})\geq\tau.

Figure 12 plots the FPR averaged over the 1010 keys, as a function of the threshold τ\tau. We compare it to the theoretical one obtained with (2). As it can be seen, they match almost perfectly for high FPR values. For lower ones (<10−6<10^{-6}), the theoretical FPR is slightly higher than the empirical one. This is a good thing since it means that if we fixed the FPR at a certain value, we would observe a lower one in practice.

Appendix C Additional Qualitative Results