Attributing Image Generative Models using Latent Fingerprints

Guangyu Nie, Changhoon Kim, Yezhou Yang, Yi Ren

Introduction

Generative models can now create synthetic content such as images and audio that are indistinguishable from those captured in nature (Karras et al., 2020; Rombach et al., 2022; Ramesh et al., 2022; Hawthorne et al., 2022). This poses a serious threat when used for malicious purposes, such as disinformation (Breland, 2019) and malicious impersonation (Satter, 2019). Such potential threats have slowed down the industrialization process of generative models, as conservative model inventors hesitate to release their source code (Yu et al., 2020). For example, in 2020, OpenAI refused to release the source code of their GPT-2 (Radford et al., 2019) model due to concerns over potential malicious use (Brockman et al., 2020). Additionally, the source code for DALL-E (Ramesh et al., 2021) and DALL-E 2 (Ramesh et al., 2022) has not been released for the same reason (Mishkin & Ahmad, 2022).

One potential solution is model attribution (Yu et al., 2018; Kim et al., 2021; Yu et al., 2020), where the model distributor tweaks each user-end model to generate content with model-specific fingerprints. In practice, we consider a scenario where the model distributor or regulator maintains a database of user-specific keys corresponding to each user’s downloaded model. In the event of a malicious attempt, the regulator can identify the user responsible for the attempt by using attribution.

The generation quality of gig_{i} measures the difference between px,ip_{x,i} and the data distribution used for learning G\mathcal{G}, e.g., the Fréchet Inception Distance (FID) score (Heusel et al., 2017) for images. Inception score (IS) (Salimans et al., 2016) is also measured for px,ip_{x,i} as additional generation quality metrics. Fingerprint secrecy is measured by the mean structural similarity index measure (SSIM) of individual images drawn from px,ip_{x,i}. Compared with generation quality, this metric focuses on how obvious fingerprints are rather than how well two content distributions match. Lastly, the fingerprint capacity is n=2dϕn=2^{d_{\phi}}.

Existing model attribution methods exhibit a significant tradeoff between attribution accuracy, generation quality, and secrecy, particularly when countermeasures against deattribution attempts, e.g., image postprocesses, are considered. For example, Kim et al. (2021) use shallow fingerprints for image generators in the form of gi(z)=g0(z)+ϕig_{i}(z)=g_{0}(z)+\phi_{i} where g0(⋅)g_{0}(\cdot) is an unfingerprinted model and show that ϕi\phi_{i}s have to significantly alter the original contents to achieve good attribution accuracy against image blurring, causing an unfavorable drop in generation quality and secrecy (Fig. 1(a)).

Contributions. (1) We propose a novel fingerprinting strategy that directly embeds the fingerprints into pretrained generative model, as a mean to achieve responsible white-box model distribution. (2) We prove and empirically verify that there exists an intrinsic tradeoff between attribution accuracy and generation quality. This tradeoff is affected by fingerprint variables including the choice of the fingerprinting space, the fingerprint strength, and its capacity. Parametric studies on these variables for StyleGAN2 (SG2) and a Latent Diffusion Model(LDM) lead to improved accuracy-quality tradeoff from the previous SOTA. In addition, our method requires negligible computation compared with previous SOTA, rendering it more applicable to popular large-scale models, including latent diffusion ones. (3) We show that using a postprocess-specific LPIPS metric for model attribution leads to further improved attribution accuracy against image postprocesses.

Related Work

Certifiable model attribution through shallow fingerprints. Kim et al. (2021) propose shallow fingerprints gi(z)=g0(z)+ϕig_{i}(z)=g_{0}(z)+\phi_{i} and linear classifiers for attribution. These simplifications allow the derivation of sufficient conditions of Φ\Phi to achieve certifiable attribution of G\mathcal{G}. Since the fingerprints are added as noises rather than semantic changes coherent with the generated contents, increased noise strength becomes necessary to maintain attribution accuracy under postprocesses. While this paper does not provide attribution certification for latent fingerprint, we discuss the technical feasibility and challenges in achieving this goal.

GAN inversion. The model attribution problem can be formulated as a GAN inversion problem. A learning-based inversion (Perarnau et al., 2016; Bau et al., 2019) optimizes parameters of the encoder network which map an image to latent code zz. On the other hand, optimization-based inversion (Abdal et al., 2019; Huh et al., 2020) solve for latent code zz that minimizes distance metric between a given image and generated image g(z)g(z). The learning-based method is computationally more efficient in the inference stage compared to optimization-based method. However, optimization-based GAN inversion achieves a superior quality of latent interpretation, which can be referred to as the quality-time tradeoff (Xia et al., 2022). In our method, we utilized the optimization-based inversion, as faithful latent interpretation is critical in our application. To further enforce faithful latent interpretation, we incorporate existing techniques, e.g., parallel search, to solve this non-convex problem, but uniquely exploit the fact that fingerprints are small latent perturbations to enable analysis on the accuracy-quality tradeoff.

Methods

where α∼pα\alpha\sim p_{\alpha} and pαp_{\alpha} is induced by pwp_{w}. Then, the user can generate fingerprinted images, g(wϕ(α))g(w_{\phi}(\alpha)). The choice of (U,V)(U,V) and σ\sigma affects the attribution accuracy and generalization performance, which we analyze in Sec. 3.2.

Attribution.To decode user-specific key from the image g(wϕ(α))g(w_{\phi}(\alpha)), we formulate an optimization problem:

While ll is l2l_{2} norm for analysis in Sec.3.2, here we minimize LPIPS (Zhang et al., 2018) which measures the perceptual difference between two images. Through experiments, we discovered that attribution accuracy can be improved by constraining α\alpha. Here the upper and lower bounds of α\alpha are chosen based on the empirical limits observed from pαp_{\alpha}. In practice, we introduce a penalty on α^\hat{\alpha} with large enough Lagrange multipliers and solve the resulting unconstrained problem. To avoid convergence to unfavorable local solutions, we also employ parallel search with n initial guesses of α^\hat{\alpha} drawn through Latin hypercube sampling (LHS).

2 Accuracy-quality tradeoff

∃ c>0\exists~{}c>0 such that if σ≤c\sigma\leq c and ∣∣ϵα∣∣2≤c||\epsilon_{\alpha}||_{2}\leq c, the fingerprint estimation problem

has an error ϵϕ=−(σ2VTHˉϕV)−1VTHˉϕUϵα\epsilon_{\phi}=-(\sigma^{2}V^{T}\bar{H}_{\phi}V)^{-1}V^{T}\bar{H}_{\phi}U\epsilon_{\alpha}.

Remarks: (1) Similar to the classic design of the experiment, one can reduce ∣∣ϵϕ∣∣||\epsilon_{\phi}|| by maximizing det(VTHˉϕV)det(V^{T}\bar{H}_{\phi}V), which sets columns of VV as the eigenvectors associated with the largest dϕd_{\phi} eigenvalues of Hˉϕ\bar{H}_{\phi}. However, Hˉϕ\bar{H}_{\phi} is neither computable because ϕ\phi is unknown during the estimation, nor is it tractable because Jwϕ(α)J_{w_{\phi}(\alpha)} is large in practice. To this end, we propose to use the covariance of pwp_{w}, denoted by Σw\Sigma_{w}, to replace Hˉϕ\bar{H}_{\phi} in experiments. In Appendix C, we support this approximation empirically by showing that Σw\Sigma_{w} and Hˉw\bar{H}_{w} (the non-fingerprinted mean Gram matrix) are qualitatively similar in that the principal components of both matrices offer disentangled semantic dimensions. (2) Let the kkth largest eigenvalue of Hˉw\bar{H}_{w} be γk\gamma_{k}. By setting columns of VV as the eigenvectors of Hˉw\bar{H}_{w} associated with the largest dϕd_{\phi} eigenvalues, and by noting that ϕ^\hat{\phi} is accurate only when all of its elements match with ϕ\phi ((1)), the worst-case estimation error is governed by γdϕ−1\gamma_{d_{\phi}}^{-1}. This means that higher key capacity, i.e., larger dϕd_{\phi}, leads to worse attribution accuracy. (3) From the proposition, ϵϕ=0\epsilon_{\phi}=0 if VV and UU are complementary sets of eigenvectors of Hˉϕ\bar{H}_{\phi}. In practice this decoupling between ϵϕ\epsilon_{\phi} and ϵα\epsilon_{\alpha} cannot be achieved due to the assumptions and approximations we made.

For any τ>0\tau>0 and η∈(0,1)\eta\in(0,1), ∃ c(τ,η)>0\exists~{}c(\tau,\eta)>0 and ν>0\nu>0, such that if σ≤c(τ,η)\sigma\leq c(\tau,\eta) and λV,i≤c(τ,η)\lambda_{V,i}\leq c(\tau,\eta) for all i=1,...,dϕi=1,...,d_{\phi}, then ∥μ0−μ1∥22≤σ2γU,maxdϕ+τ\|\mu_{0}-\mu_{1}\|_{2}^{2}\leq\sigma^{2}\gamma_{U,max}d_{\phi}+\tau and ∣tr(Σ0−Σ1)∣≤λV,maxγU,maxdϕ+2νσdϕ+τ|tr(\Sigma_{0}-\Sigma_{1})|\leq\lambda_{V,max}\gamma_{U,max}d_{\phi}+2\nu\sigma\sqrt{d_{\phi}}+\tau with probability at least 1−η1-\eta.

Remarks: Recall that for improving attribution accuracy, a practical approach is to choose VV as eigenvectors associated with the largest eigenvalues of Σw\Sigma_{w}. Notices that with the approximated distribution with α∼N(0,diag(λU))\alpha\sim\mathcal{N}(0,diag(\lambda_{U})) and β∼N(0,diag(λV))\beta\sim\mathcal{N}(0,diag(\lambda_{V})), Σw=diag([λUT,λVT]T)\Sigma_{w}=diag([\lambda_{U}^{T},\lambda_{V}^{T}]^{T}). On the other hand, from Proposition 3.2, generation quality improves if we minimize λV,max\lambda_{V,max} by choosing VV according to the smallest eigenvalues of Σw\Sigma_{w}. In addition, smaller key capacity (dϕd_{\phi}) and lower strength (σ\sigma) also improve the generation quality. Propositions 3.1 and 3.2 together reveal the intrinsic accuracy-quality tradeoff.

Experiments

In this section, we present empirical evidence of the accuracy-quality tradeoff and show an improved tradeoff from the previous SOTA by using latent fingerprints. Experiments are conducted for both with and without a combination of postprocesses including (image noising, blurring, and JPEG compression, and their combination).

Models, data, and metrics. We conduct experiments on SG2 (Karras et al., 2020) and LDM (Rombach et al., 2022) models trained on various datasets including FFHQ (Karras et al., 2019), AFHQ-Cat, and AFHQ-Dog (Choi et al., 2020). Generation quality is measured by the Frechet Inception distance (FID) (Heusel et al., 2017) and inception score (IS) (Salimans et al., 2016), attribution accuracy by (1), and fingerprint secrecy by SSIM.

Latent fingerprint dimensions. To approximate Σw\Sigma_{w}, we drew 10K samples from pwp_{w} for SG2, which has a semantic latent space dimensionality of 512, and 50K samples from pwp_{w} for LDM, which has a semantic latent space dimensionality of 12,288. We define fingerprint dimensions VV as a subset eigenvectors of Σw\Sigma_{w} associated with consecutive eigenvalues: V:=PC[i:j]V:=PC[i:j], where PCPC is the full set of principal components of Σw\Sigma_{w} sorted by their variances in the descending order while ii and jj represent the starting and ending indices of the subset.

Attribution. To compute the empirical accuracy (Eq. (1)), we use 1K samples drawn from pzp_{z} for each fingerprint ϕ\phi, and use 1K fingerprints where each bit is drawn independently from a Bernoulli distribution with p=0.5p=0.5. In Table 1, we show that both constraints on α^\hat{\alpha} and parallel search with 20 initial guesses improve the empirical attribution accuracy across models and datasets. Notably, constrained estimation is essential for the successful attribution of LDMs. In these experiments, VV is chosen as the eigenvectors associated with the 64 smallest eigenvalues of Σw\Sigma_{w} as a worst-case scenario for attribution accuracy.

2 Attribution performance without postprocessing

We present generation quality results in Table 1. Since the least variant principal components are used as fingerprints, generation quality (FID) and fingerprint secrecy (SSIM) are preserved.

The results suggest that the attribution accuracy, generation quality, capacity (2642^{64}), and fingerprint secrecy are all acceptable using the proposed method. Fig. 2 visualizes and compares latent fingerprints generated from small vs. large eigenvalues of Σw\Sigma_{w}. Fingerprints corresponding to small eigenvalues are non-semantic, while those to large eigenvalues create semantic changes. We will later show that semantic yet subtle (perceptually insignificant) fingerprints are necessary to counter image postprocesses.

Accuracy-quality tradeoff. Table 2 summarizes the tradeoff when we vary the choice of VV and the fingerprint strength σ\sigma while fixing the fingerprint length dϕd_{\phi} to 64. Then in Table 3 we sweep dϕd_{\phi} while keeping VV as PCs associated with smallest eigenvalues of Σw\Sigma_{w} and σ=1\sigma=1. The experiments are conducted on SG2 and LDM on the FFHQ dataset. The empirical results in Table 2 are consistent with our analysis: Accuracy decreases while generation quality improves when VV is moved from major to minor principal components. For fingerprint strength, however, we observe that the positive effect of strength on the accuracy, as predicted by Proposition 3.1, is only limited to small σ\sigma. This is because larger σ\sigma causes pixel values to go out of bounds, causing loss of information. In Table 3, we summarize the attribution accuracy, FID, and SSIM score under 32- to 128-bit keys. Accuracy and generation quality, in particular the latter, are both affected by dϕd_{\phi} as predicted.

3 Fingerprint performance with postprocessing

We now consider more realistic scenarios where generated images are postprocessed, either maliciously as an attempt to remove the fingerprints or unintentionally, before they are attributed. Under this setting, our method achieves better accuracy-quality tradeoff than shallow fingerprinting under two realistic settings: (1) when noising and JPEG compression are used as unknown postprocesses, and (2) when the set of postprocesses, rather than the ones that are actually chosen, is known.

Postprocesses. To keep our solution realistic, we solve the attribution problem by assuming that the potential postprocesses are unknown:

where T:Rdx→RdxT:R^{d_{x}}\rightarrow R^{d_{x}} is a postprocess function, and T(g(wϕ(α)))T(g(w_{\phi}(\alpha))) is a given image from which the fingerprint is to be estimated. We assume that TT does not change the image in a semantically meaningful way, because otherwise the value of the image for either an attacker or a benign user will be lost. Since our method adds semantically meaningful perturbations to the images, we expect such latent fingerprints to be more robust to postprocesses than shallow ones (Kim et al., 2021) added directly to images and will lead to improved attribution accuracy. To test this hypothesis, we consider four types of postprocesses: Noising, Blurring, JPEG, and Combo. Noising adds a Gaussian white noise of standard deviation randomly sample from UU[0, 0.1]. Blurring uses a randomly selected Gaussian kernel size from and a standard deviation of [0.5, 1.0, 1.5, 2.0]. We randomly sample the JPEG quality from . These parameters are chosen to be mild so that images do not lose their semantic contents. And Combo randomly chooses a subset of the three through a binomial distribution with p=0.5p=0.5 and uses the same postprocess parameters.

Modified LPIPS metric. In addition to testing the worst-case scenario where postprocesses are completely unknown, we also consider cases where they are known. While this is unrealistic for individual postprocesses, it is worth investigating when we assume that the set of postprocesses, rather than the ones that are actually chosen, is known. Within this scenario, we show that modifying LPIPS according to the postprocess improves the attribution accuracy. To explain, LPIPS is originally trained on a so-called “two alternative forced choice” (2AFC) dataset. Each data point of 2AFC contains three images: a reference, p0p_{0}, and p1p_{1}, where p0p_{0} and p1p_{1} are distorted in different ways based on the reference. A human evaluator then ranks p0p_{0} and p1p_{1} by their similarity to the reference. Here we propose the following modification to the dataset for training postprocess-specific metrics: Similar to 2AFC, for each data point, we draw a reference image xx from the default generative model to be fingerprinted and define p0p_{0} as the fingerprinted version of xx. p1p_{1} is then a postprocessed version of xx given a specific TT (or random combinations for Combo). To match with the setting of 2AFC, we sample 64 ×\times 64 patches from xx, p0p_{0}, and p1p_{1} as training samples. We then rank patches from p1p_{1} as being more similar to those of xx than p0p_{0}. With this setting, the resulting LPIPS metric becomes more sensitive to fingerprints than mild postprocesses. The detailed training of the modified LPIPS follows the vgg-lin configuration in (Zhang et al., 2018). It should be noted that, unlike previous SOTA where shallow fingerprint (Kim et al., 2021) or encoder-decoder models (Yu et al., 2020) are retrained based on the known attacks, our fingerprinting mechanism, and therefore generalization performance, are agnostic to postprocesses.

Accuracy-quality tradeoff. We summarize fingerprint performance metrics on SG2 and FFHQ in Table 4. The attribution accuracies reported here are estimated using the strongest parameters of each attack. For Combo, we use sequentially apply Blurring+Noising+JPEG as a deterministic worst-case attack. To estimate attribution accuracy, we solved the estimation problem in (4.3) where postprocesses are applied. The proposed method: We choose VV as a subset of 32 consecutive eigenvectors of Σw\Sigma_{w} starting from the 1th, 17th, and 33th eigenvectors, denoted respectively by PC[0:32], PC[16:48], and PC[32:64] in the table. fingerprinting strength σ\sigma is set to 3. Attribution results from both a standard and a postprocess-specific LPIPS metric are reported in the UK (unknown) and KN (known) columns, respectively. Accuracies for our method are computed based on 100 random fingerprint samples from 2322^{32}, each with 100 random generations. The baseline: We compare with a shallow fingerprinting method from (Kim et al., 2021) (denoted by BL). When the postprocesses are known, BL performs postprocess-specific computation to derive shallow fingerprints that are optimally robust against the known postprocess. Results in UK and KN columns for BL are respectively without and with the postprocess-specific fingerprint computation. BL accuracies are computed based on 10 fingerprints, each with 100 random generations.

It is worth noting that the shallow fingerprinting method is not as scalable as ours (See Appendix: D, and increasing the key capacity decreases the overall attribution accuracy (see (Kim et al., 2021)). Also, recall that the key length affects attribution accuracy (Proposition 1). Therefore, we conduct a fairer comparison to highlight the advantage of our method. Here we choose a subset of fingerprints PC[32:40]PC[32:40] (256 fingerprints) and report performance in Table 5, where accuracies are computed using the same settings as before. Visual comparisons between our method (PC[32:40]PC[32:40]) and the baseline can be found in Fig. 3: To maintain attribution accuracy, high-strength shallow fingerprints, in the form of color patches, are needed around eyes and noses, and significantly lower the generation quality. In comparison, our method uses semantic changes that are robust to postprocesses. The choice of semantic dimensions, however, needs to be carefully chosen for the fingerprint to be perceptually subtle.

Fingerprint secrecy. The secrecy of the fingerprint is evaluated through the SSIM which is designed to measure the similarity between two images by taking into account three components: loss of correlation, luminance distortion, and contrast distortion (Wang et al., 2004). To facilitate a more equitable comparison between our methodology and the baseline, we select a subset of fingerprints presented in Table 5 for evaluation. As shown in Table 5, the robustly fingerprinted model for PC[32:40]PC[32:40] exhibits superior fingerprint secrecy compared to the baseline, while simultaneously outperforming it in attribution accuracy. Besides these quantitative measures, our approach also demonstrates a qualitative advantage in terms of secrecy when compared to the baseline (see Fig. 3). This is largely attributed to the fact that subtle semantic variations across images are more challenging to visually detect and thus being removed compared to typical artifacts introduced by shallow fingerprinting.

Conclusion

This paper investigated latent fingerprint as a solution to enable the attribution of generative models. Our solution achieved a better tradeoff between attribution accuracy and generation quality than the previous SOTA that uses shallow fingerprints, and also has extremely low computational cost compared to SOTA methods that require encoder-decoder training with high data complexity, rendering our method more scalable to attributing large models with high-dimensional latent spaces. Limitations and future directions: (1) There is currently a lack of certification on attribution accuracy due to the nonlinear nature of both the fingerprinting and the fingerprint estimation processes. Formally, by considering both the generation and estimation processes as discrete-time dynamics, such certification would require forward reachability analysis of fingerprinted contents and backward reachability analysis of the fingerprint, e.g., convex approximation of the support of px,ip_{x,i} and ϕ^\hat{\phi}. It is worth investigating whether existing neural net certification methods can be applied. (2) Our method extracts fingerprints from the training data. Even with feature decomposition, the amount of features that can be used as fingerprints is limited. Thus the accuracy-quality tradeoff is governed by the data. It would be interesting to see if auxiliary datasets can help to learn novel and perceptually insignificant fingerprints that are robust against postprocesses, e.g., background patterns.

Acknowledgment

This work is partially supported by the National Science Foundation under Grant No. 2038666 and No. 2101052 and by an Amazon AWS Machine Learning Research Award (MLRA). Any opinions, findings, and conclusions expressed in this material are those of the author(s) and do not reflect the views of the funding entities.

References

Appendix A Proof of Propositions

Proposition 1. ∃c>0\exists c>0 such that if σ≤c\sigma\leq c and ∣∣ϵα∣∣2≤c||\epsilon_{\alpha}||_{2}\leq c, the fingerprint estimation problem

Let x^:=g(Uα^ϕ(α)+σVϕ^)\hat{x}:=g(U\hat{\alpha}_{\phi}(\alpha)+\sigma V\hat{\phi}), we have

Ignoring higher-order terms and , we then have

For any τ>0\tau>0, there exists cc, such that if σ≤c\sigma\leq c and ∣∣ϵα∣∣2≤c||\epsilon_{\alpha}||_{2}\leq c,

Removing terms independent from ϵϕ\epsilon_{\phi} to reformulate (3) as

A.2 Proposition 2

Proposition 2. For any τ>0\tau>0 and η∈(0,1)\eta\in(0,1), there exists c(τ,η)>0c(\tau,\eta)>0 and ν>0\nu>0, such that if σ≤c(τ,η)\sigma\leq c(\tau,\eta) and λV,i≤c(τ,η)\lambda_{V,i}\leq c(\tau,\eta) for all i=1,...,dϕi=1,...,d_{\phi}, ∥μ0−μ1∥22≤σ2γU,maxdϕ+τ\|\mu_{0}-\mu_{1}\|_{2}^{2}\leq\sigma^{2}\gamma_{U,max}d_{\phi}+\tau and ∣tr(Σ0−Σ1)∣≤λV,maxγU,maxdϕ+2νσdϕ+τ|tr(\Sigma_{0}-\Sigma_{1})|\leq\lambda_{V,max}\gamma_{U,max}d_{\phi}+2\nu\sigma\sqrt{d_{\phi}}+\tau with probability at least 1−η1-\eta.

We start with ∥μ0−μ1∥22\|\mu_{0}-\mu_{1}\|_{2}^{2}. From Taylor’s expansion and using the independence between α\alpha and β\beta, we have

Let v=Vϕv=V\phi. With orthonormal VV and binary-coded ϕ\phi, we have

Then combining (11), (5), (6), (7), we have with probability at least 1−η1-\eta

For covariances, let ΣU=Cov(g(μ+Uα))\Sigma_{U}=Cov(g(\mu+U\alpha)). We have

For tr(Σ0)tr(\Sigma_{0}), using the same treatment for the residual, we have for any τ>0\tau>0 and η∈(0,1)\eta\in(0,1), there exists c(τ,η)>0c(\tau,\eta)>0, such that if λV,i≤c(τ,η)\lambda_{V,i}\leq c(\tau,\eta) for all i=1,...,dϕi=1,...,d_{\phi}, the following upper bound applies with at least probability 1−η1-\eta:

For the lower bound, we have tr(Σ0)≥tr(ΣU)tr(\Sigma_{0})\geq tr(\Sigma_{U}).

For tr(Σ1)tr(\Sigma_{1}), we first denote by JiTJ_{i}^{T} the iith row of Jμ+UαJ_{\mu+U\alpha}, ΣJi\Sigma_{J_{i}} its covariance matrix, and σi2\sigma_{i}^{2} the maximum eigenvalue of ΣJi\Sigma_{J_{i}}. Then with binary-coded ϕ\phi, we have

Then let gig_{i} (resp. viv_{i}) be the iith element of g(μ+Uα)g(\mu+U\alpha) (resp. Jμ+UαVϕJ_{\mu+U\alpha}V\phi), and σU,i2\sigma_{U,i}^{2} be the iith diagonal element of ΣU\Sigma_{U}. Using (11), we have the following bound on the trace of the covariance between g(μ+Uα)g(\mu+U\alpha) and Jμ+UαVϕJ_{\mu+U\alpha}V\phi:

Lastly, by ignoring σ2\sigma^{2} terms and borrowing the same τ\tau, η\eta, and c(τ,η)c(\tau,\eta), we have with probability at least 1−η1-\eta:

Therefore, with probability at least 1−η1-\eta

Appendix B Convergence on α𝛼\alpha

In the proofs, we assume that ∥ϵα∥2\|\epsilon_{\alpha}\|_{2} is small and constant. Here we show empirical estimation results on SG2 and on FFHQ, AFHQ-DOG, AFHQ-CAT datasets. The results in Fig. 4(a) are averaged over 100 random α\alpha and 100 random ϕ\phi, and uses parallel search on α\alpha during the estimation.

Since computing Hˉw\bar{H}_{w} for large models is intractable, here we train a SG2 on MNIST to estimate Hˉw\bar{H}_{w}. Fig. 4(b) summarizes perturbed images from a randomly chosen reference along principal components of Hˉw\bar{H}_{w} and Σw\Sigma_{w}. Note that both have quickly diminishing eigenvalues. Therefore most components other than the few major ones lead to imperceptible changes in the image space.

Appendix D Computational Complexity and Efficiency Comparison of Proposed and Baseline Methods

To comprehensively evaluate the computational costs of the proposed and baseline methods, we analyzed them from two perspectives: fingerprint generation and attribution. Our proposed method demonstrates a significant increase in efficiency compared to existing methods with regards to fingerprint generation. However, current methods, such as those proposed in (Kim et al., 2021), often necessitate a considerable amount of time for fine-tuning the model before generating a fingerprinted image. For instance, the baseline method requires approximately one hour to fine-tune the model for each individual user. For a key capacity of 2322^{32} users, it would take an estimated 2322^{32} hours to train on an NVIDIA V100 GPU. In contrast, our proposed method does not require the model to be fine-tuned, and only requires Principal Component Analysis (PCA) to be performed on the latent space of each pre-trained model to identify the editing direction. This process takes approximately three seconds for a key capacity of 2322^{32}.

Regarding attribution time, the baseline approach performs attribution through a pre-trained network, resulting in low computation cost during inference. On the other hand, the proposed method uses optimization techniques to perform attribution, resulting in a longer computation time. Specifically, the average attribution time for the baseline method is approximately two seconds, while the proposed method takes approximately 126 seconds on average for 1k optimization trials (in parallel) using an NVIDIA V100 GPU.

It is important to acknowledge that there exists a trade-off between computational costs during generation and attribution. The choice of which aspect is more critical may depend on the specific application requirements.

Appendix E Ablation Study

In this section, we estimated attribution accuracy based on various attack parameters with multiple editing directions (see Tab.6,7,8,9). The image quality evaluation is available in Tab.10 and more visualizations can be found in Fig.5.