HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models

Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Wei Wei, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, Kfir Aberman

Introduction

Recent work on text-to-image (T2I) personalization has opened the door for a new class of creative applications. Specifically, for face personalization, it allows generation of new images of a specific face or person in different styles. The impressive diversity of styles is owed to the strong prior of pre-trained diffusion model, and one of the key properties of works such as DreamBooth , is the ability to implant a new subject into the model without damaging the model’s prior. Another key feature of this type of method is that subject’s essence and details are conserved even when applying vastly different styles. For example, when training on photographs of a person’s face, one is able to generate new images of that person in animated cartoon styles, where a part of that person’s essence is preserved and represented in the animated cartoon figure - suggesting some amount of visual semantic understanding in the diffusion model. These are two core characteristics of DreamBooth and related methods, that we would like to leave untouched. Nevertheless, DreamBooth has some shortcomings: size and speed. For size, the original DreamBooth paper finetunes all of the weights of the UNet and Text Encoder of the diffusion model, which amount to more than 1GB for Stable Diffusion. In terms of speed, notwithstanding inference speed issues of diffusion models, training a DreamBooth model takes about 5 minutes for Stable Diffusion (1,000 iterations of training). This limits the potential impact of the work. In this work, we want to address these shortcomings, without altering the impressive key properties of DreamBooth, namely style diversity and subject fidelity, as depctied in Figure 1. Specifically, we want to conserve model integrity and closely approximate subject essence in a fast manner with a small model.

Our work proposes to tackle the problems of size and speed of DreamBooth, while preserving model integrity, editability and subject fidelity. We propose the following contributions:

Lighweight DreamBooth (LiDB) - a personalized text-to-image model, where the customized part is roughly 100KB of size. This is achieved by training a DreamBooth model in a low-dimensional weight-space generated by a random orthogonal incomplete basis inside of a low-rank adaptation weight space.

New HyperNetwork architecture that leverages the Lightweight DreamBooth configuration and generates the customized part of the weights for a given subject in a text-to-image diffusion model. These provide a strong directional initialization that allows us to further finetune the model in order to achieve strong subject fidelity within a few iteration. Our method is 25x faster than DreamBooth while achieving similar performances.

We propose the technique of rank-relaxed finetuning, where the rank of a LoRA DreamBooth model is relaxed during optimization in order to achieve higher subject fidelity, allowing us to initialize the personalized model with an initial approximation using our HyperNetwork, and then approximate the high-level subject details using rank-relaxed finetuning.

One key aspect that leads us to investigate a HyperNetwork approach is the realization that in order to be able to synthesize specific subjects with high fidelity, using a given generative model, we have to “modify" its output domain, and insert knowledge about the subject into the model, namely by modifying the network weights.

Related Work

Several recent models such as Imagen , DALL-E2 , Stable Diffusion (SD) , Muse , Parti etc. demonstrate excellent image generation capabilities given a text prompt. Some Text-to-Image (T2I) models such as Stable Diffusion and Muse also allows conditioning the generation with a given image via an encoder network. Techniques such as ControlNet propose ways to incorporate new input conditioning such as depth. Test text and image based conditioning in these models do not capture sufficient subject details. Given the relatively small size of SD, for the ease of experimentation, we demonstrate our HyperDreamBooth on SD model. But the proposed technique is generic and can be applicable to any T2I model.

Personalization of Generative Models

Given one or few subject images, the aim of personalized generation is to generate images of that particular subject in various contexts. Earlier works in this space use GANs to edit a given subject image into new contexts. Pivotal tuning proposes to finetune a GAN with an inverted latent code. The work of proposes to finetune StyleGAN using around 100 face images to obtain a personalized generative prior. Casanova et al. proposes to condition a GAN using an input image to generate variations of that input image. All these GAN based techniques suffer from either poor subject fidelity or a lack of context diversity in the generated images.

HyperNetworks were introduced as an idea of using an auxiliary neural network to predict network weights in order to change the functioning of a specific neural network . Since then, they have been used for tasks in image generation that are close to personalization, such as inversion for StyleGAN , similar to work that seeks to invert the latent code of an image in order to edit that image in the GAN latent space .

T2I Personalization via Finetuning

More recently, several works propose techniques for personalizing T2I models resulting in higher subject fidelity and versatile text based recontextualization of a given subject. Textual Inversion proposes to optimize an input text embedding on the few subject images and use that optimized text embedding to generate subject images. propose a richer textual inversion space capturing more subject details. DreamBooth proposes to optimize the entire T2I network weights to adapt to a given subject resulting in higher subject fidelity in output images. Several works propose ways to optimize compact weight spaces instead of the entire network as in DreamBooth. CustomDiffusion proposes to only optimize cross-attention layers. SVDiff proposes to optimize singular values of weights. LoRa proposes to optimize low-rank approximations of weight residuals. StyleDrop proposes to use adapter tuning and finetunes a small set of adapter weights for style personalization. DreamArtist proposes a one-shot personalization techniques by employing a positive-negative prompt tuning strategy. Most of these finetuning techniques, despite generating high-quality subject-driven generations, are slow and can take several minutes for every subject.

Fast T2I Personalization

Several concurrent works propose ways for faster personalization of T2I models. The works of and propose to learn encoders that predicts initial text embeddings following by complete network finetuning for better subject fidelity. In contrast, our hypernetwork directly predicts low-rank network residuals. SuTI proposes to first create a large paired dataset of input images and the corresponding recontexualized images generated using standard DreamBooth. It then uses this dataset to train a separate network that can perform personalized image generation in a feed-forward manner. Despite mitigating the need for finetuning, the inference model in SuTI does not conserve the original T2I model’s integrity and also suffers from a lack of high subject fidelity. InstantBooth and Taming Encoder create a new conditioning branch for the diffusion model, which can be conditioned using a small set of images, or a single image, in order to generate personalized outputs in different styles. Both methods need to train the diffusion model, or the conditioning branch, to achieve this task. These methods are trained on large datasets of images (InstantBooth 1.3M samples of bodies from a proprietary dataset, Taming Encoder on CelebA and Getty ). FastComposer proposes to use image encoder to predict subject-specific embeddings and focus on the problem of identity blending in multi-subject generation. The work of propose to guide the diffusion process using face recognition loss to generate specific subject images. In such guidance techniques, it is usually difficult to balance diversity in recontextualizations and subject fidelity while also keeping the generations within the image distribution. Face0 proposes to condition a T2I model on face embeddings so that one can generate subject-specific images in a feedforward manner without any test-time optimization. Celeb-basis proposes to learn PCA basis of celebrity name embeddings which are then used for efficient personalization of T2I models. In contrast to these existing techniques, we propose a novel hypernetwork based approach to directly predict low-rank network residuals for a given subject.

Preliminaries

DreamBooth provides a network fine-tuning strategy to adapt a given T2I denoising network Dθ\mathcal{D}_{\theta} to generate images of a specific subject. At a high-level, DreamBooth optimizes all the diffusion network weights θ\theta on a few given subject images while also retaining the generalization ability of the original model with class-specific prior preservation loss . In the case of Stable Diffusion , this amounts to finetuning the entire denoising UNet has over 1GB of parameters. In addition, DreamBooth on a single subject takes about 5 minutes with 1K training iterations.

Method

Our approach consists of 3 core elements which we explain in this section. We begin by introducing the concept of the Lightweight DreamBooth (LiDB) and demonstrate how the Low-Rank decomposition (LoRa) of the weights can be further decomposed to effectively minimize the number of personalized weights within the model. Next, we discuss the HyperNetwork training and the architecture the model entails, which enables us to predict the LiDB weights from a single image. Lastly, we present the concept of rank-relaxed fast fine-tuning, a technique that enables us to significantly amplify the fidelity of the output subject within a few seconds. Fig. 2 shows the overview of hypernetwork training followed by fast fine-tuning strategy in our HyperDreamBooth technique.

Given our objective of generating the personalized subset of weights directly using a HyperNetwork, it would be beneficial to reduce their number to a minimum while maintaining strong results for subject fidelity, editability and style diversity. To this end, we propose a new low-dimensional weight space for model personalization which allows for personalized diffusion models that are 10,000 times smaller than a DreamBooth model and more than 10 times smaller than a LoRA DreamBooth model. Our final version has only 30K variables and takes up only 120 KB of storage space.

where r<<min(n,m)r<<\text{min}(n,m), a<na<n and b<mb<m. AauxA_{\text{aux}} and BauxB_{\text{aux}} are randomly initialized with orthogonal row vectors with constant magnitude - and frozen, and BtrainB_{\text{train}} and AtrainA_{\text{train}} are learnable. Surprisingly, we find that with a=100a=100 and b=50b=50, which yields models that have only 30K trainable variables and are 120 KB in size, personalization results are strong and maintain subject fidelity, editability and style diversity. We show results for personalization using LiDB in the experiments section.

2 HyperNetwork for Fast Personalization of Text-to-Image Models

where x\mathbf{x} is the reference image, θ\theta are the pre-optimized weight parameters of the personalized model for image x\mathbf{x}, Dθ\mathcal{D}_{\theta} is the diffusion model (with weights θ\theta) conditioned on the noisy image x+ϵ\mathbf{x}+\epsilon and the supervisory text-prompt c\mathbf{c}, and finally α\alpha and β\beta are hyperparameters that control for the relative weight of each loss. Fig. 2 (top) illustrates the hypernetwork training.

We propose to eschew any type of learned token embedding for this task, and our hypernetwork acts solely to predict the LiDB weights of the diffusion model. We simply propose to condition the learning process “a [V] face” for all samples, where [V] is a rare identifier described in . At inference time variations of this prompt can be used, to insert semantic modifications, for example “a [V] face in impressionist style”.

HyperNetwork Architecture

Concretely, as illustrated in Fig. 4, we separate the HyperNetwork architecture into two parts: a ViT image encoder and a transformer decoder. We use a ViT-H for the encoder architecture and a 2-hidden layer transformer decoder for the decoder architecture. The transformer decoder is a strong fit for this type of weight prediction task, since the output of a diffusion UNet or Text Encoder is sequentially dependent on the weights of the layers, thus in order to personalize a model there is interdependence of the weights from different layers. In previous work , this dependency is not rigorously modeled in the HyperNetwork, whereas with a transformer decoder with a positional embedding, this positional dependency is modeled - similar to dependencies between words in a language model transformer. To the best of our knowledge this is the first use of a transformer decoder as a HyperNetwork.

Iterative Prediction

We find that the HyperNetwork achieves better and more confident predictions given an iterative learning and prediction scenario , where intermediate weight predictions are fed to the HyperNetwork and the network’s task is to improve that initial prediction. We only perform the image encoding once, and these extracted features f\mathbf{f} are then used for all rounds of iterative prediction for the HyperNetwork decoding transformer T\mathcal{T}. This speeds up training and inference, and we find that it does not affect the quality of results. Specifically, the forward pass of T\mathcal{T} becomes:

where kk is the current iteration of weight prediction, and terminates once k=sk=s, where ss is a hyperparameter controlling the maximum amount of iterations. Weights θ\theta are initialized to zero for k=0k=0. Trainable linear layers are used to convert the decoder outputs into the final layer weights. We use the CelebAHQ dataset for training the HyperNetwork, and find that we only need 15K identities to achieve strong results, much less data than other concurrent methods.

3 Rank-Relaxed Fast Finetuning

We find that the initial HyperNetwork prediction is in great measure directionally correct and generates faces with similar semantic attributes (gender, facial hair, hair color, skin color, etc.) as the target face consistently. Nevertheless, fine details are not sufficiently captured. We propose a final fast finetuning step in order to capture such details, which is magnitudes faster than DreamBooth, but achieves virtually identical results with strong subject fidelity, editability and style diversity. Specifically, we first predict personalized diffusion model weights θ^=H(x)\hat{\theta}=\mathcal{H}(\mathbf{x}) and then subsequently finetune the weights using the diffusion denoising loss L(x)=∣∣Dθ^(x+ϵ,c)−x∣∣22L(\mathbf{x})=||\mathcal{D}_{\hat{\theta}}(\mathbf{x}+\epsilon,\mathbf{c})-\mathbf{x}||_{2}^{2}. A key contribution of our work is the idea of rank-relaxed finetuning, where we relax the rank of the LoRA model from r=1r=1 to r>1r>1 before fast finetuning. Specifically, we add the predicted HyperNetwork weights to the overall weights of the model, and then perform LoRA finetuning with a new higher rank. This expands the capability of our method of approximating high-frequency details of the subject, giving higher subject fidelity than methods that are locked to lower ranks of weight updates. To the best of our knowledge we are the first to propose such rank-relaxed LoRA models.

We use the same supervision text prompt “a [V] face” this fast finetuning step. We find that given the HyperNetwork initialization, fast finetuning can be done in 40 iterations, which is 25x faster than DreamBooth and LoRA DreamBooth . We show an example of initial, intermediate and final results in Figure 5.

Experiments

We implement our HyperDreamBooth on the Stable Diffusion v1.5 diffusion model and we predict the LoRa weights for all cross and self-attention layers of the diffusion UNet as well as the CLIP text encoder. For privacy reasons, all face images used for visuals are synthetic, from the SFHQ dataset . For training, we use 15K images from CelebA-HQ .

Our method achieves strong personalization results for widely diverse faces, with performance that is identically or surpasses that of the state-of-the art optimization driven methods . Moreover, we achieve very strong editability, with semantic transformations of face identities into highly different domains such as figurines and animated characters, and we conserve the strong style prior of the model which allows for a wide variety of style generations. We show results in Figure 6.

Given the statistical nature of HyperNetwork prediction, some samples that are OOD for the HyperNetwork due to lighting, pose, or other reasons, can yield subotpimal results. Specifically, we identity three types of errors that can occur. There can be (1) a semantic directional error in the HyperNetwork’s initial prediction which can yield erroneous semantic information of a subject (wrong eye color, wrong hair type, wrong gender, etc.) (2) incorrect subject detail capture during the fast finetuning phase, which yields samples that are close to the reference identity but not similar enough and (3) underfitting of both HyperNetwork and fast finetuning, which can yield low editability with respect to some styles.

2 Comparisons

We compare our method to both Textual Inversion and DreamBooth using the parameters proposed in both works, with the exception that we increase the number of iterations of DreamBooth to 1,200 in order to achieve improved personalization and facial details. Results are shown in Figure 7. We observe that our method outperforms both Textual Inversion and DreamBooth generally, in the one-input-image regime.

Quantitative Comparisons and Ablations

We compare our method to Textual Inversion and DreamBooth using a face recognition metric (“Face Rec.” using an Inception ResNet, trained on VGGFace2), and the DINO, CLIP-I and CLIP-T metrics proposed in . We use 100 identities from CelebAHQ , and 30 prompts, including both simple and complex style-modification and recontextualization prompts for a total of 30,000 samples. We show in Table 1 that our approach obtains the highest scores for all metrics. One thing to note is that face recognition metrics are relatively weak in this specific scenario, given that face recognition networks are only trained on real images and are not trained to recognize the same person in different styles. In order to compensate for this, we conduct a user study described further below.

We also conduct comparisons to more aggressive DreamBooth training, with lower number of iterations and higher learning rate. Specifically, we use 400 iterations for DreamBooth-Agg-1 and 40 iterations for DreamBooth-Agg-2 instead of 1200 for DreamBooth. We increase the learning rate and tune the weight decay to compensate for the change in number of iterations. Note that DreamBooth-Agg-2 is roughly equivalent to only doing fast finetuning without the hypernetwork component of our work. We show in Table 2 that more aggressive training of DreamBooth generally degrades results when not using our method, which includes a HyperNetwork initialization of the diffusion model weights.

Finally, we show an ablation study of our method. We remove the HyperNetwork (No Hyper), only use the HyperNetwork without finetuning (Only Hyper) and also use our full setup without iterative HyperNetwork predictions (k=1). We show results in Table 3 and find that our full setup with iterative prediction achieves best subject fidelity, with a slightly lower prompt following metric.

User Study

We conduct a user study for face identity preservation of outputs and compare our method to DreamBooth and Textual Inversion. Specifically, we present the reference face image and two random generations using the same prompt from our method and the baseline, and ask the user to rate which one has most similar face identity to the reference face image. We test a total of 25 identities, and query 5 users per question, with a total of 1,000 sample pairs evaluated. We take the majority vote for each pair. We present our results in Table 4, where we show a strong preference for face identity preservation of our method.

Societal Impact

This work aims to empower users with a tool for augmenting their creativity and ability to express themselves through creations in an intuitive manner. However, advanced methods for image generation can affect society in complex ways . Our proposed method inherits many possible concerns that affect this class of image generation, including altering sensitive personal characteristics such as skin color, age and gender, as well as reproducing unfair bias that can already be found in pre-trained model’s training data. The underlying open source pre-trained model used in our work, Stable Diffusion, exhibits some of these concerns. All concerns related to our work have been present in the litany of recent personalization work, and the only augmented risk is that our method is more efficient and faster than previous work. In particular, we haven’t found in our experiments any difference with respect to previous work on bias, or harmful content, and we have qualitatively found that our method works equally well across different ethnicities, ages, and other important personal characteristics. Nevertheless, future research in generative modeling and model personalization must continue investigating and revalidating these concerns.

Conclusion

In this work, we have presented HyperDreamBooth a novel method for fast and lightweight subject-driven personalization of text-to-image diffusion models. Our method leverages a HyperNetwork to generate Lightweight DreamBooth (LiDB) parameters for a diffusion model with a subsequent fast rank-relaxed finetuning that achieves a significant reduction in size and speed compared to DreamBooth and other optimization-based personalization work. We have demonstrated that our method can produce high-quality and diverse images of faces in different styles and with different semantic modifications, while preserving subject details and model integrity.

References