SVDiff: Compact Parameter Space for Diffusion Fine-Tuning

Ligong Han, Yinxiao Li, Han Zhang, Peyman Milanfar, Dimitris Metaxas, Feng Yang

Introduction

Recent years have witnessed the rapid advancement of diffusion-based text-to-image generative models , which have enabled the generation of high-quality images through simple text prompts. These models are capable of generating a wide range of objects, styles, and scenes with remarkable realism and diversity. These models, with their exceptional results, have inspired researchers to investigate various ways to harness their power for image editing .

In the pursuit of model personalization and customization, some recent works such as Textual-Inversion , DreamBooth , and Custom Diffusion have further unleashed the potential of large-scale text-to-image diffusion models. By fine-tuning the parameters of the pre-trained models, these methods allow the diffusion models to be adapted to specific tasks or individual user preferences.

Despite their promising results, there are still some limitations associated with fine-tuning large-scale text-to-image diffusion models. One limitation is the large parameter space, which can lead to overfitting or drifting from the original generalization ability . Another challenge is the difficulty in learning multiple personalized concepts especially when they are of similar categories .

To alleviate overfitting, we draw inspiration from the efficient parameter space in the GAN literature and propose a compact yet efficient parameter space, spectral shift, for diffusion model by only fine-tuning the singular values of the weight matrices of the model. This approach is inspired by prior work in GAN adaptation showing that constraining the space of trainable parameters can lead to improved performance on target domain . Comparing with another popular low-rank constraint , the spectral shifts utilize the full representation power of the weight matrix while being more compact (e.g. 1.7MB for StableDiffusion , full weight checkpoint consumes 3.66GB of storage). The compact parameter space allows us to combat overfitting and language-drifting issues, especially when prior-preservation loss is not applicable. We demonstrate this use case by presenting a simple DreamBooth-based single-image editing framework.

To further enhance the ability of the model to learn multiple personalized concepts, we propose a simple Cut-Mix-Unmix data-augmentation technique. This technique, together with our proposed spectral shift parameter space, enables us to learn multiple personalized concepts even for semantically similar categories (e.g. a “cat” and a “dog”).

We present a compact (≈\approx2,200×\times fewer parameters compared with vanilla DreamBooth , measured on StableDiffusion ) yet efficient parameter space for diffusion model fine-tuning based on singular-value decomposition of weight kernels.

We present a text-based single-image editing framework and demonstrate its use case with our proposed spectral shift parameter space.

We present a generic Cut-Mix-Unmix method for data-augmentation to enhance the ability of the model to learn multiple personalized concepts.

This work opens up new avenues for the efficient and effective fine-tuning large-scale text-to-image diffusion models for personalization and customization. Our proposed method provides a promising starting point for further research in this direction.

Related Work

Text-to-image diffusion models Diffusion models have proven to be highly effective in learning data distributions and have shown impressive results in image synthesis, leading to various applications . Recent advancements have also explored transformer-based architectures . In particular, the field of text-guided image synthesis has seen significant growth with the introduction of diffusion models, achieving state-of-the-art results in large-scale text-to-image synthesis tasks . Our main experiments were conducted using StableDiffusion , which is a popular variant of latent diffusion models (LDMs) that operates on a latent space of a pre-trained autoencoder to reduce the dimensionality of the data samples, allowing the diffusion model to utilize the well-compressed semantic features and visual patterns learned by the encoder.

Fine-tuning generative models for personalization Recent works have focused on customizing and personalizing text-to-image diffusion models by fine-tuning the text embedding , full weights , cross-attention layers , or adapters using a few personalized images. Other works have also investigated training-free approaches for fast adaptation . The idea of fine-tuning only the singular values of weight matrices was introduced by FSGAN in the GAN literature and further advanced by NaviGAN with an unsupervised method for discovering semantic directions in this compact parameter space. Our method, SVDiff, introduces this concept to the fine-tuning of diffusion models and is designed for few-shot adaptation. A similar approach, LoRA , explores low-rank adaptation for text-to-image diffusion fine-tuning, while our proposed SVDiff optimizes all singular values of the weight matrix, leading to an even smaller model checkpoint. Similar idea has also been explored in few-shot segmentation .

Diffusion-based image editing Diffusion models have also shown great potential for semantic editing . These methods typically focus on inversion and reconstruction by optimizing the null-text embedding or overfitting to the given image . Our proposed method, SVDiff, presents a simple DreamBooth-based single-image editing framework that demonstrates the potential of SVDiff in single image editing and mitigating overfitting.

Method

Diffusion models StableDiffusion , the model we experiment with, is a variant of latent diffusion models (LDMs) . LDMs transform the input images x\mathbf{x} into a latent code z\mathbf{z} through an encoder E\mathcal{E}, where z=E(x)\mathbf{z}=\mathcal{E}(\mathbf{x}), and perform the denoising process in the latent space Z\mathcal{Z}. Briefly, a LDM ϵ^θ\hat{{\boldsymbol{\epsilon}}}_{\theta} is trained with a denoising objective:

where (z,c)(\mathbf{z},\mathbf{c}) are data-conditioning pairs (image latents and text embeddings), ϵ∼N(0,I){\boldsymbol{\epsilon}}\sim\mathcal{N}(\mathbf{0},\mathbf{I}), t∼Uniform(1,T)t\sim\text{Uniform}(1,T), and θ\theta represents the model parameters. We omit tt in the following for brevity.

2 Compact Parameter Space for Diffusion Fine-tuning

Instead of fine-tuning the full weight matrix, we only update the weight matrix by optimizing the spectral shift , δ{\boldsymbol{\delta}}, which is defined as the difference between the singular values of the updated weight matrix and the original weight matrix. The updated weight matrix can be re-assembled by

Training loss The fine-tuning is performed using the same loss function that was used for training the diffusion model, with a weighted prior-preservation loss :

where (z∗,c∗)(\mathbf{z}^{*},\mathbf{c}^{*}) represents the target data-conditioning pairs that the model is being adapted to, and (zpr,cpr)(\mathbf{z}^{pr},\mathbf{c}^{pr}) represents the prior data-conditioning pairs generated by the pretrained model. This loss function extends the one proposed by Model Rewriting for GANs to the context of diffusion models, with the prior-preservation loss serving as the smoothing term. In the case of single image editing, where the prior-preservation loss cannot be utilized, we set λ=0\lambda=0.

Combining spectral shifts Moreover, the individually trained spectral shifts can be combined into a new model to create novel renderings. This can enable applications including interpolation, style mixing (Fig. 9), or multi-subject generation (Fig. 8). Here we consider two common strategies, addition and interpolation. To add δ1{\boldsymbol{\delta}}_{1} and δ2{\boldsymbol{\delta}}_{2} into δ′{\boldsymbol{\delta}}^{\prime},

For interpolation between two models with 0≤α≤10\leq\alpha\leq 1,

This allows for smooth transitions between models and the ability to interpolate between different image styles.

3 Cut-Mix-Unmix for Multi-Subject Generation

We discovered that when training the StableDiffusion model with multiple concepts simultaneously (randomly choosing one concept at each data sampling iteration), the model tends to mix their styles when rendering them in one image for difficult compositions or subjects of similar categories (as shown in Fig. 6). To explicitly guide the model not to mix personalized styles, we propose a simple technique called Cut-Mix-Unmix. By constructing and presenting the model with “correctly” cut-and-mixed image samples (as shown in Fig. 3), we instruct the model to unmix styles. In this method, we manually create CutMix-like image samples and corresponding prompts (e.g. “photo of a [V1V_{1}] dog on the left and a [V2V_{2}] sculpture on the right” or “photo of a [V2V_{2}] sculpture and a [V1V_{1}] dog” as illustrated in Fig. 3). The prior loss samples are generated in a similar manner. During training, Cut-Mix-Unmix data augmentation is applied with a pre-defined probability (usually set to 0.6). This probability is not set to 1, as doing so would make it challenging for the model to differentiate between subjects. During inference, we use a different prompt from the one used during training, such as “a [V1V_{1}] dog sitting beside a [V2V_{2}] sculpture”. However, if the model overfits to the Cut-Mix-Unmix samples, it may generate samples with stitching artifacts even with a different prompt. We found that using negative prompts can sometimes alleviate these artifacts, as detailed in appendix.

We further present an extension to our fine-tuning approach by incorporating an “unmix” regularization on the cross-attention maps. This is motivated by our observation that in fine-tuned models, the dog’s special token (“sks”) attends largely to the panda, as depicted in Fig. 23. To enforce separation between the two subjects, we use MSE on the non-corresponding regions of the cross-attention maps. This loss encourages the dog’s special token to focus solely on the dog and vice versa for the panda. The results of this extension show a significant reduction in stitching artifact.

4 Single-Image Editing

In this section, we present a framework for single image editing, called CoSINE (Compact parameter space for SINgle image Editing), by fine-tuning a diffusion model with an image-prompt pair. The procedure is outlined in Fig. 4. The desired edits can be obtained at inference time by modifying the prompt. For example, we fine-tune the model with the input image and text description “photo of a crown with a blue diamond and a golden eagle on it”, and at inference time if we want to remove the eagle, we simply sample from the fine-tuned model with text “photo of a crown with a blue diamond on it”. To mitigate overfitting during fine-tuning, CoSINE uses the spectral shift parameter space instead of full weights, reducing the risk of overfitting and language drifting. The trade-off between faithful reconstruction and editability, as discussed in , is acknowledged, and the purpose of CoSINE is to allow more flexible edits rather than exact reconstructions.

For edits that do not require large structural changes (like repose, “standing” →\rightarrow “lying down” or “zoom in”), results can be improved with DDIM inversion . Before sampling, we run DDIM inversion with classifier-free guidance scale 1 conditioned on the target text prompt c\mathbf{c} and encode the input image z∗\mathbf{z}^{*} to a latent noise map,

(θ′\theta^{\prime} denotes the fine-tuned model parameters) from which the inference pipeline starts. As expected, large structural changes may still require more noise being injected in the denoising process. Here we consider two types of noise injection: i) setting η>0\eta>0 (as defined in DDIM , and ii) perturbing zT\mathbf{z}_{T}. For the latter, we interpolate between zT\mathbf{z}_{T} and a random noise ϵ∼N(0,I){\boldsymbol{\epsilon}}\sim\mathcal{N}(0,\mathbf{I}) with spherical linear interpolation ,

with ϕ=arccos⁡(cos⁡(zT,ϵ))\phi=\arccos{(\cos(\mathbf{z}_{T},{\boldsymbol{\epsilon}}))}. For more results and analysis, please see the experimental section.

Other approaches, such as Imagic , have been proposed to address overfitting and language drifting in fine-tuning-based single-image editing. Imagic fine-tunes the diffusion model on the input image and target text description, and then interpolates between the optimized and target text embedding to avoid overfitting. However, Imagic requires fine-tuning on each target text prompt at test time.

Experiment

The experiments evaluate SVDiff on various tasks such as single-/multi-subject generation, single image editing, and ablations. The DDIM sampler with η=0\eta=0 is used for all generated samples, unless specified otherwise.

In this section, we present the results of our proposed SVDiff for customized single-subject generation proposed in DreamBooth , which involves fine-tuning the pretrained text-to-image diffusion model on a single object or concept (using 3-5 images). The original DreamBooth was implemented on Imagen and we conduct our experiments based on its StableDiffusion implementation . We provide visual comparisons of 5 examples in Fig. 5. All baselines were trained for 500 or 1000 steps with batch size 1 (except for Custom Diffusion , which used a default batch size of 2), and the best model was selected for fair comparison. As Fig. 5 shows, SVDiff produces similar results to DreamBooth (which fine-tunes the full model weights) despite having a much smaller parameter space. Custom Diffusion, on the other hand, tends to underfit the training images as seen in rows 2, 3, and 5 of Fig. 5. We assess the text and image alignment in Fig. 10. The results show that the performance of SVDiff is similar to that of DreamBooth, while Custom Diffusion tends to underfit as seen from its position in the upper left corner of the plot.

2 Multi-Subject Generation

In this section, we present the multi-subject generation results to illustrate the advantage of our proposed “Cut-Mix-Unmix” data augmentation technique. When enabled, we perform Cut-Mix-Unmix data-augmentation with probability of 0.6 in each data sampling iteration and two subjects are randomly selected without replacement. A comparison between using “Cut-Mix-Unmix” (marked as “w/ Cut-Mix-Unmix”) and not using it (marked as “w/o Cut-Mix-Unmix”, performing augmentation with probability 0) are shown in Fig. 6. Each row of images are generated using the same text prompt displayed below the images. Note that the Cut-Mix-Unmix data augmentation technique is generic and can be applied to fine-tuning full weights as well.

To assess the visual quality of images generated using the “Cut-Mix-Unmix” method with either SVD or full weights, we conducted a user study using Amazon MTurk with 400 generated image pairs . The participants were presented with an image pair generated using the same random seed, and were asked to identify the better image by answering the question, “Which image contains both objects from the two input images with a consistent background?” Each image pair was evaluated by 10 different raters, and the aggregated results showed that SVD was favored over full weights 60.9% of the time, with a standard deviation of 6.9%. More details and analysis will be provided in the appendix.

Additionally, we also conducted experiments that involve training on three concepts simultaneously. During training, we still construct Cut-Mix samples with probability 0.6 by randomly sample two subjects. Interestingly, we observe that for concepts that are already semantically well-separated, e.g. “dog/building” or “sculpture/building”, the model can successfully generate desired results even without using Cut-Mix-Unmix. However, it fails to disentangle semantically more similar concepts, e.g. “dog/panda” as shown in Fig. 6-g.

3 Single Image Editing

In this section, we present results for the single image editing application. As depicted in Fig. 7, each row presents three edits with fine-tuning of both spectral shifts (marked as “Ours”) and full weights (marked as “Full”). The text prompts for the corresponding edited images are given below the images. The aim of this experiment is to demonstrate that regularizing the parameter space with spectral shifts effectively mitigates the language drift issue, as defined in (the model overfits to a single image and loses its ability to generalize and perform desired edits).

As previously discussed, when DDIM inversion is not employed, fine-tuning with spectral shifts can lead to sometimes over-creative results. We show examples and comparisons of editing results with and without DDIM inversion in the appendix (Fig. 21). Our results show that DDIM inversion improves the editing quality and alignment with the input image for non-structural edits when using our spectral shift parameter space, but may worsen the results for full weight fine-tuning. For example, in Fig. 7, we use DDIM inversion for the edits in (a,c,e) and the first edit in (d). The second edit in (d) presents an interesting example where our method can actually make the statue hold an apple with its hand. Additionally, our fine-tuning approach still produces the desired edit of an empty room even with DDIM inversion, as seen in the third edit of Fig. 7-a. Overall, we see that SVDiff can still perform desired edits when full model fine-tuning exhibits language drift, i.e. it fails to remove the picture in the second edit of (a), change the pose of the dog in the second edit of (c), and zoom-in view in (d).

4 Analysis and Ablation

Due to space limitations, we present parameter subsets, weight combination, interpolation and style mixing analysis in this section and provide further analysis including rank, scaling, and correlation in the appendix.

Parameter subsets We explore the fine-tuning of spectral shifts within a subset of parameters in UNet. We consider 12 distinct subsets for our ablation study, as outlined in Tab. 1. Due to space limitations, we provide the visual samples and text-/image-alignment scores for each subset on 5 subjects in appendix Fig. 27 and Fig. 14, respectively. Our findings are as follows: (1) Optimizing the cross-attention (CA) layers generally results in better preservation of subject identity compared to optimizing key and value projections. (2) Optimizing the up-, down-, or mid-blocks of UNet alone is insufficient to maintain identity, which is why we did not further isolate subsets of each part. However, it appears that the up-blocks exhibit the best preservation of identity. (3) In terms of dimensionality, the 2D weights demonstrate the most influence, and offer better identity preservation than UNet-CA.

Weight combination We analyze the effects of weight combination by Eq. 4. Fig. 8 shows a comparison between combining only spectral shifts (marked in “SVD”) and combining the full weights (marked in “Full”). The combined model in both cases retains unique features for individual subjects, but may blend their styles for similar concepts (as seen in (e)). For dissimilar concepts (such as the [V2V_{2}] sculpture and [V3V_{3}] building in (j)), the models can still produce separate representations of each subject. Interestingly, combining full weight deltas can sometimes result in better preservation of individual concepts, as seen in the clear building feature in (j). We posit that this is due to the fact that SVDiff limits update directions to the eigenvectors, which are identical for different subjects. As a result, summing individually trained spectral shifts tends to create more “interference” than summing full weight deltas.

Style transfer and mixing We demonstrate the capability of style transfer using our proposed method. We show that by using a single fine-tuned model, the personalized style can be transferred to a different class by changing the class word during inference, or by adding a prompt such as “in style of”. We also show that by summing two sets of spectral shifts (as discussed above), their styles can be mixed. The results show different outcomes of different style-mixing strategies, with changes to both the class and personalized style. We further explore a more challenging and controllable approach for style-mixing. Inspired by the disentangling property observed in StyleGAN , we hypothesize that a similar property applies in our context. Following Extended Textual Inversion (XTI ), we conducted a style mixing experiment, as illustrated in Fig. 12. For this experiment, we fine-tuned SVDiff on the UNet-2D subset and employed the geometry information provided by (16, down’, 1) - (8, down’, 0) (as described in XTI, Section 8.1). We observe that our spectral shift parameter space allows us to achieve a similar disentangled style-mixing effect, comparable to the P+\mathcal{P}+ space in XTI.

Interpolation Fig. 11 shows the results of weight interpolation for both spectral shifts and full weights. The models are marked as “SVD” and “Full”, respectively. The first two rows of the figure demonstrate interpolating between two different classes, such as “dog” and “sculpture”, using the same abstract class word “thing” for training. Each column shows the sample from α\alpha-interpolated models. For spectral shifts (“SVD”), we use Eq. 5 and for full weights, we use W′=W+αΔW1+(1−α)ΔW2=αW1+(1−α)W2W^{\prime}=W+\alpha\Delta W_{1}+(1-\alpha)\Delta W_{2}=\alpha W_{1}+(1-\alpha)W_{2}. The images in each row are generated using the same random seed with the deterministic DDIM sampler (η=0\eta=0). As seen from the results, both spectral shift and full weight interpolation are capable of generating intermediate concepts between the two original classes.

5 Comparison with LoRA

In our comparison of SVDiff and LoRA for single image editing, we find that while LoRA tends to underfit, SVDiff provides a balanced trade-off between faithfulness and realism. Additionally, SVDiff results in a significantly smaller delta checkpoint size, being 1/2 to 1/3 that of LoRA. However, in cases where the model requires extensive fine-tuning or learning of new concepts, LoRA’s flexibility to adjust its capability by changing the rank may be beneficial. Further research is needed to explore the potential benefits of combining these approaches. A comparison can be found in the appendix Fig. 21.

It is noteworthy that, with rank one, the storage and update requirements for the WW matrix of shape M×NM\times N in SVDiff are min⁡(M,N)\min(M,N) floats, compared to (M+N)(M+N) floats for LoRA. This may be useful for amortizing or developing training-free approaches for DreamBooth . Additionally, exploring functional forms of spectral shifts is an interesting avenue for future research.

Conclusion and Limitation

In conclusion, we have proposed a compact parameter space, spectral shift, for diffusion model fine-tuning. The results of our experiments show that fine-tuning in this parameter space achieves similar or even better results compared to full weight fine-tuning in both single- and multi-subject generation. Our proposed Cut-Mix-Unmix data-augmentation technique also improves the quality of multi-subject generation, making it possible to handle cases where subjects are of similar categories. Additionally, spectral shift serves as a regularization method, enabling new use cases like single image editing.

Limitations Our method has certain limitations, including the decrease in performance of Cut-Mix-Unmix as more subjects are added and the possibility of an inadequately-preserved background in single image editing. Despite these limitations, we see great potential in our approach for fine-tuning diffusion models and look forward to exploring its capabilities further in future research, such as combining spectral shifts with LoRA or developing training-free approaches for fast personalizing concepts.

References

Appendix

Appendix A Implementation Details

Implementation The original DreamBooth was implemented on Imagen and we conduct our experiments based on its StableDiffusion implementation . DreamBooth and Custom Diffusion are implemented in StableDiffusion with Diffusers library . For LoRA , we use our own implementation for fair comparison, in which we also fine-tune the 1-D weight kernels, and use rank-1 for 2-D and 4-D weight kernels. This results in a slightly larger delta checkpoint (of size 5.62MB) than the official LoRA implementation .

Learning rate Our experiments show that the learning rate for these spectral shifts needs to be much larger (1,000 times, e.g. 10−310^{-3}) than the learning rate used for fine-tuning the full weights. For 1-D weights that are not decomposed, we use either the original learning rate of 10−610^{-6} to prevent overfitting or a larger learning rate to allow for a more rapid adaptation of the model, depending on the desired trade-off between stability and speed of adaptation.

Appendix B Single Image Editing

We show comparisons of with and without DDIM inversion using ours (“SVD”), LoRA (“LoRA”), and DreamBooth (“Full”) on single-image editing in Fig. 21. If inversion is not used, DDIM sampler with η=0\eta=0 is applied. If inversion is employed, we use DDIM sampler with η=0.5\eta=0.5 and α=0\alpha=0, except for edits in Fig. 21-(d,f,h) where η=0.9\eta=0.9 and α=0.9\alpha=0.9. Interestingly, for the chair example (row 2) we need to inject large amount of noise to get desired edits. For other edits (a,b,e,g,i,j) DDIM inversion improves editing quality and alignment with input images for “SVD”, but makes results worse for “Full” in edits (b,g,i) and for “LoRA” in edits (b,i). We can conclude that DDIM inversion improves editing quality and alignment with input images for non-structural edits when using our spectral shift parameter space. We also observe that LoRA in general tends to underfit the input image, as shown in (c,d,e,i) (without inversion).

B.2 Comparison with Other Methods

Furthermore, we compare our method with the popular Instruct-Pix2Pix in Fig. 22 (marked as “ip2p”). The comparison is not entirely fair as Instruct-Pix2Pix does not require fine-tuning on individual images. Nevertheless, it is worth investigating fast personalized adaptation and avoiding per-image fine-tuning in future work.

Appendix C Multi-Subject Generation

In Tab. 2, we present the results of human evaluation comparing our method (“SVD”) and the full weight fine-tuning method (“Full”). For each of the four subject combinations, 1000 ratings were collected. Participants were shown two generated images side-by-side and were asked to choose their preferred image or indicate that it was “hard to decide” (4.1%, 1.2%, 2.1%, and 2.2% respectively). Visual examples are given in Fig. 13.

C.2 Analysis of Cut-Mix-Unmix

In this section, we present additional analysis of the Cut-Mix-Unmix data augmentation technique (without unmix regularization on the cross-attention maps). Fig. 17 illustrates the results of the default “left and right” augmentations, which still generate meaningful relations such as (a) “wear”, (b) “in”, and (c) “ride”. In the case of (a), “full” overfits to the augmentation layout. In our initial experiments, we also randomly split the left and right images and observed similar results as with a fixed 1:1 ratio. In (d), the “up and down” augmentation exhibits similar behavior to “left and right”. Nevertheless, we concur that introducing a random layout (particularly with our proposed cross-attention regularization) could further mitigate overfitting. We leave this study for future work.

C.3 Negative Prompt

To perform negative prompting, we repurpose the prior-preservation prompts as negative prompts cneg\mathbf{c}^{neg}. Recall that Classifier-free guidance (CFG) extrapolates the conditional score by a scale factor s>1s>1,

where 0<β<10<\beta<1. This can be easily extended to including multiple negative prompts. Fig. 20 shows a few examples of using negative prompts to remove the stitching artifacts introduced by Cut-Mix-Unmix. We hypothesize that this is because the model is trained to associate the prior prompt to the stitching style so negative prompting can help removing the stitching edges. However, we observe that negative prompting may not always help.

C.4 Extensions

We show a preliminary extension of our Cut-Mix-Unmix to Attend-and-Excite . As shown in Fig. 16, Cut-Mix-Unmix helps better disentangle respective visual features of the dog and the cat. It is also possible to extend and integrate our method to other attention-based methods .

Appendix D Single-Subject Generation

As previously discussed in the context of multi-subject generation, adding regularization terms on the cross-attention (CA) maps can help to enforce separation between the subjects. Our observations also show that the cross-attention map associated with the special token may attend to unwanted areas, even in the case of single-subject generation. For instance, as shown in Fig. 19, the attention of the special token “[V1V_{1}]” leaks to the background (whereas the attention of “[V2V_{2}]” does not). To address this issue, we explore the use of regularization on cross-attention maps to improve single-subject generation. The main idea is to limit the attention of the special token to be no more spread-out than that of the coarse class token (e.g. “dog”). To achieve this, we first obtain a binary mask MtM_{t} indicating the subject by thresholding the coarse class token’s attention map. Then, we add a L2 regularization loss on the special token’s attention map AtVA_{t}^{V}, as follows:

where ⊙\odot denotes elementwise multiplication and sg is a stop gradient operator. The results of using this CA regularization are compared to the case without regularization in Fig. 18, and as expected, the regularization reduces overfitting to the background.

D.2 Fine-Tuning with Fewer Steps

Here we show results of fast adaptation for single subject generation in Fig. 26. This setting is slightly different from the experiments in the main text since we limit the fine-tuning steps as 100 without prior-preservation loss (for main results we fine-tune 500-1000 steps with prior-preservation loss). Thus we tune the learning rate for each method to balance between faithfullness and realism . The learning rates we used are as follows:

SVDiff: 1-D weights 2×10−32\times 10^{-3}, 2-D and 4-D weights 5×10−35\times 10^{-3}

LoRA : 1-D weights 2×10−32\times 10^{-3}, 2-D and 4-D weights 1×10−41\times 10^{-4}

DreamBooth : 1-D weights 1×10−31\times 10^{-3}, 2-D and 4-D weights 5×10−65\times 10^{-6}

In Fig. 26, the performance comparison of our method, LoRA and DreamBooth is shown under fast fine-tuning setting. The results indicate that all three methods perform similarly, except for the “No-Face” sculpture in (c) where LoRA shows underfitting and DreamBooth exhibits overfitting. In (e), SVDiff also shows overfitting, which could be a result of the large learning rate used.

Appendix E Analysis on Spectral Shifts

Fig. 24 shows the results of limiting the rank of the spectral shifts of 2-D and 4-D weight kernels during training. Two examples are shown for each of the three subjects, one with the training prompt (to “reconstruct” the subject) and one with an edited prompt. Results show that the model can still reconstruct the subject with rank 1, but may struggle to capture details with an edited prompt when the rank of spectral shift is low. The visual differences between reconstructed and edited samples are smaller for the Teddybear than the building and panda sculpture, potentially because the pre-trained model already understands the concept of a Teddybear.

E.2 Correlations

We present the results of the correlation analysis of individually learned spectral shifts for each subject in Fig. 15. Each entry in the figure represents the average cosine similarities between the spectral shifts of two subjects, computed across all layers. The diagonal entries show the average cosine similarities between two runs with the learning rate of 1-D weights set to 10−310^{-3} and 10−610^{-6}, respectively. The results indicate that the similarity between conceptually similar subjects is relatively high, such as between the “panda” and “No-Face” sculptures or between the “Teddybear” and “Tortoise” plushies.

E.3 Scaling

Fig. 25 demonstrates the effect of scaling spectral shifts (labeled as “SVD”, Σδ′=diag(ReLU(σ+sδ))\Sigma_{{\boldsymbol{\delta}}^{\prime}}=\text{diag}(\text{ReLU}({\boldsymbol{\sigma}}+s{\boldsymbol{\delta}})) with scale ss) and weight deltas (marked as “full”, W′=W+sΔWW^{\prime}=W+s\Delta W with scale ss). Samples are generated using the same random seed. Scaling both the spectral shift and full weight delta affects the presence of personalized attributes and features. The results show that scaling the weight delta also influences attribute strength. However, a scale value that is too large (e.g. s=2s=2) can cause deviation from the text prompt and result in dark samples.

Appendix F Image Attribution

Avocado plushy: https://unsplash.com/photos/8V4y-XXT3MQ.

Pink chair: https://unsplash.com/photos/1JJJIHh7-Mk.

Brown and white puppy: https://unsplash.com/photos/brFsZ7qszSY, https://unsplash.com/photos/eoqnr8ikwFE, https://unsplash.com/photos/LHeDYF6az38, and https://unsplash.com/photos/9M0tSjb-cpA.

Crown: https://unsplash.com/photos/8Dpi2Mb1-PM.

Bedroom: https://unsplash.com/photos/x53OUnxwynQ.

Dog with flower: https://unsplash.com/photos/Sg3XwuEpybU.

Statue-of-Liberty: https://unsplash.com/photos/s0di82cRiUQ.

Beetle car: https://unsplash.com/photos/YEPDV3T8Vi8.

Building: https://finmath.rutgers.edu/admissions/how-to-apply and luvemakphoto/ Getty Images.

Teddybear, tortoise plushy, grey dog, and cat images are taken from Custom Diffusion : https://www.cs.cmu.edu/~custom-diffusion/assets/data.zip.

Panda and “No-Face” sculpture images were captured and collected by the authors.