Subject-driven Text-to-Image Generation via Apprenticeship Learning

Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, William W. Cohen

Introduction

Recent text-to-image generation models have shown great progress in generating highly realistic, accurate, and diverse images from a given text prompt. These models are pre-trained on web-crawled image-text pairs like LAION with autoregressive backend models or diffusion backend models . Though achieving unprecedented success in generating highly accurate images, these models are not able to customize to a given subject, like a specific dog, shoe, backpack, etc. Therefore, subject-driven text-to-image generation, the task of generating highly customized images with respect to a target subject, has attracted significant attention from the community. Subject-driven image generation is related to text-driven image editing but often needs to perform more sophisticated transformations to source images (e.g., rotating the view, zooming in/out, changing the pose of subject, etc.) so existing image editing methods are generally not suitable for this new task.

Current subject-driven text-to-image generation approaches are slow and expensive. While different approaches like DreamBooth , Imagic , and Textual Inversion have been proposed, they all require fine-tuning specific models for a given subject on one or a few demonstrated examples, which typically takes at least 10-20 minutesRunning on A100 according to public colab: https://huggingface.co/sd-dreambooth-library and https://huggingface.co/docs/diffusers/training/text_inversion. to specialize the text-to-image model checkpoint for the given subjects. These approaches are time-consuming as they require back-propagating gradients over the entire model for hundreds or even thousands of steps per customization. Moreover, they are space-consuming as they require storing a subject-specific checkpoint per subject. To avoid the excessive cost, Re-Imagen proposed a retrieval-augmented text-to-image framework to train a subject-driven generation model in a weakly-supervised fashion. Since the retrieved neighbor images are not guaranteed to contain the same subjects, the model does not perform as good as DreamBooth for the task of subject-driven image generation.

To avoid excessive computation and memory costs, we propose to train a single subject-driven text-to-image generation model that can perform on-the-fly subject customization. Our method is dubbed Subject-driven Text-to-Image generator (SuTI), which is trained with a novel apprenticeship learning algorithm. Unlike standard apprenticeship learning which only focuses on learning from one expert, our apprentice model imitates the behaviors of a massive number of specialized expert models. After such training, SuTI can instantly adapt to unseen subjects and unseen or even compositional descriptions with only 3-5 in-context demonstrations within 30 seconds (on a Cloud TPU v4).

Figure 2 presents a conceptual diagram of the learning and data preparation pipeline. We first group the images in WebLI by their source URL to form tiny image clusters because images from the same URL are likely to contain the same subject. We then performed extensive image-to-image and image-to-text similarity filtering to retain image clusters that contain highly similar content. For each subject image cluster, we fine-tuned an expert model to specialize in the given subject. Then, we use the fine-tuned experts to synthesize new images given unseen creative captions proposed by large language models. However, the tuned expert models are not perfect and prone to errors, therefore, we adopt a quality validation metric to filter out a large portion of degraded outputs. The remaining high-quality images are provided as a training signal to teach the apprentice model SuTI to perform subject-driven image generation with high fidelity. During inference, the trained SuTI can attend to a few in-context demonstrations to synthesize new images on the fly.

We evaluate SuTI on various tasks such as subject re-contextualization, attribute editing, artistic style transfer, and accessorization. We compare SuTI with existing models on DreamBench , which contains diverse subjects from wide categories accompanied by some prompt templates. We compute the CLIP-I/CLIP-T and DINO scores of SuTI’s generated images on this dataset and compare them with DreamBooth. The results indicate that SuTI can outperform DreamBooth while having 20x faster inference speed and significantly less memory footprint.

Further, we manually created 220 diverse and compositional prompts regarding the subjects in DreamBench for human evaluation, which is dubbed the DreamBench-v2 dataset. We then comprehensively compare with other baselines like InstructPix2Pix , Null-Text Inversion , Imagic , Textual Inversion , Re-Imagen , and DreamBooth on DreamBench-v2. Our human evaluation results indicate that SuTI is 5% higher than DreamBooth and at least 30% better than the other baseline in terms of human evaluation. We conduct detailed fine-grained analysis and found that SuTI’s textual alignment is significantly better than DreamBoothm while its subject alignment is slightly better than DreamBooth. However, DreamBooth’s outputs are still better in the photorealism aspect, especially in terms of fine-grained detail preservation.

We summarize our contributions in the following aspects:

We introduce the SuTI model, a subject-driven text-to-image generator that performs instant and customized generation for a visual subject with few (image, text) exemplars, all in context.

We employ the apprenticeship learning to train one single apprentice SuTI model to imitate half a million fine-tuned subject-specific experts on a large-scale seed dataset, leading to a generator model that generalizes to unseen subjects and unseen compositional descriptions.

We perform a comprehensive set of automatic and human evaluations to show the capability of our model on generating highly faithful and creative images on DreamBench and DreamBench-v2.

To facilitate the reproducibility of our model performance, we release the SuTI model API as a Google Cloud Vertex AI model service, under the production name ‘Instant tuning’Generally available at https://cloud.google.com/vertex-ai/docs/generative-ai/image/fine-tune-model.

Preliminary

In this section, we introduce the key concepts and notations about subject-driven image-text data, then discuss the basics of text-to-image diffusion models.

Diffusion models are trained to learn the image distribution by reversing the diffusion Markov chain. Theoretically, this reduces to learning to denoise xt∼q(xt∣x0)\bm{x}_{t}\sim q(\bm{x}_{t}|\bm{x}_{0}) into x0\bm{x}_{0}, with a time re-weighted square error loss—see for the complete proof:

where DD is the training dataset containing (image, condition) = (x0,c)(\bm{x}_{0},\bm{c}) pairs, the condition normally refers to the input text prompt. In practice, wtw_{t} can be simplified as 1 according to .

The customized diffusion model x^θs(xt,c)\hat{\bm{x}}_{\theta_{s}}(\bm{x}_{t},\bm{c}) has shown impressive capabilities to generate highly faithful images of the specified subject ss.

Apprenticeship Learning from Subject-specific Experts

where xt∼q(xt∣xs)\bm{x}_{t}\sim q(\bm{x}_{t}|\bm{x}_{s}). The training is similar to Eqn. 3 except that we do not have negative examples for prior preservation because finding the negative examples from the same class is expensive.

We then feed GG as a training batch to update the parameter Θ\Theta of the apprentice model using Eqn. 5. In all our experiments, we set K=400\mathtt{K}=400, with each TPU core training an expert model.

Inference. To perform subject-driven text-to-image generation, the trained SuTI takes 3-5 image-text pairs as the demonstration to generate new images based on the given text description. No optimization is needed during inference time. The only overhead of SuTI is the cost of encoding these 3-5 image-text pairs and the attention computation, which is more affordable. Our inference speed is roughly in the same order as the original text-to-image generator .

Mining and Generating Subject-driven Text-to-Image Demonstrations

Experiment

In this paper, we only train SuTI on the text →\rightarrow 64x64 diffusion model and retain the original 256x256 and 1024x1024 super-resolution as it is from Imagen .

Expert Models. The expert model is initialized from the original 2.1B Imagen 64x64 model. We tune each model on a single TPU core (32 GB) for 500 steps using Adafactor optimizer with a learning rate of 1e-5, which only takes 5 minutes to finish. We use classifier-free guidance to sample new images, where the guidance weight is set to 3030. To avoid excessive memory costs, we use fine-tuned experts to sample pseudo-target images and then write the samples as separate files. SuTI will read these files asynchronously to maximize the training speed. Our expert models have a few distinctions from the DreamBooth : 1) we adopt Adafactor instead of Adam optimizer, 2) we do not include any class word token like ‘[DOG] dog’ in the prompt. 3) we do not include in-class negatives for prior preservation. Though our expert model is weaker than DreamBooth, such design choices significantly reduce time/space costs to enable us to train millions of experts with reasonable resources.

Apprentice Model. The apprentice model contains 2.5B parameters, which is 400M parameters larger than the original 2.1B Imagen 64x64 model. The added parameters are coming from the extra attention layers over the demonstrated image-text inputs. We adopt the same architecture as Re-Imagen , where the additional image-text pairs are encoded by re-using the UNet DownStack, and the attention layers are added to the down UNet DownStack and UpStack at different resolutions.

We initialize our model from Imagen’s checkpoint. For the additional attention layers, we use random initialization. The apprentice training is performed on 128 Cloud TPU v4 chips. We train the model for a total of 150K steps. We use an Adafactor optimizer with a learning rate of 1e-4. We use 3 demonstrations during training, while the model can generalize to leverage any number of demonstrations during inference. We show our ablation studies in the following section.

Inference. We normally provide 4 demonstration image-text pairs to SuTI during inference. Increasing the number of demonstrations does not improve the generation quality much. We use a lower classifier-free guidance weight of 1515 with DDPM sampling strategy.

DreamBench. In this paper, we use the DreamBench dataset proposed by DreamBooth . The dataset contains 30 subjects like backpacks, stuffed animals, dogs, cats, clocks, etc. These images are downloaded from Unsplash. The original dataset contains 25 prompt templates covering different skills like recontextualization, property modification, accessorization, etc. In total, there are a total of 750 unique prompts generated by the template. We follow the original paper to generate 4 images for each prompt to form the 3000 images for robust evaluation. We follow DreamBooth to adopt DINO, CLIP-I to evaluate the subject fidelity, and CLIP-T to evaluate the text fidelity.

DreamBench-v2. To further increase the difficulty and diversity of DreamBench, we annotate 220 prompts for the 30 subjects in DreamBench as DreamBench-v2. We gradually increase the compositional levels of the prompt to increase the difficulty, like ‘back view of [dog]’ →\rightarrow ‘back view of [dog] watching TV’ →\rightarrow ‘back view of [dog] watching TV about birds’. This enables us to perform a breakdown analysis to understand the model’s compositional capabilities.

We use human evaluation to measure the generation quality in DreamBench-v2. Specifically, we aim at measuring the following three aspects: (1) the subject fidelity score sss_{s} measures whether the subject is being preserved, (2) the textual fidelity score sts_{t} measures whether it is aligned with the text description, (3) the photorealism score sps_{p} measures whether the image contains artifacts or blurry subjects. These are all binary scores, which are averaged over the entire dataset. We combine them as an overall score so=ss∧st∧sps_{o}=s_{s}\wedge s_{t}\wedge s_{p}, which is the most stringent score.

2 Main Results

Baselines. We provide a comprehensive list of baselines to compare with the proposed SuTI model:

InstructPix2Pix : a non-tuning method, which can generate and edit a given image really fast within a few seconds. There is no additional space consumption.

Re-Imagen : a non-tuning method, which will take a few images as input and then attend to those retrievals to generate a new image. There is no additional space consumption.

Experimental Results. We show our automatic evaluation results on the DreamBench in Table 1. We can observe that SuTI can perform better or on par with DreamBooth on all of the metrics. Specifically, SuTI outperforms DreamBooth on the DINO score by 5%, which indicates that our method is better at preserving the subject’s visual appearance. In terms of the CLIP-T score, our method is almost the same as DreamBooth, indicating an equivalent capability in terms of textual alignment. These results indicate that SuTI has achieved promising generalization to a wide variety of visual subjects, without being trained on the exact instances.

We further show our human evaluation results on the DreamBench-V2 in Table 2. It shows the related rankings for the additional storage cost and reported the average inference time measure for inferring on each subject. As can be seen, SuTI is able to outperform DreamBooth by 5% on the overall score mainly due to much higher textual alignment. In contrast, all the other existing baselines are getting much lower human evaluation score (< 42%).

Comparisons. We compare our generation results with other methods in Figure 4. As can be seen, SuTI can generate images highly faithful to the demonstrated subjects. Though SuTI is still missing some local textual (words on the bowl gets blurred) or colorization (dog hair color gets darker), the nuance is almost unperceivable for humans. The other baselines like InstructPix2Pix , and Null-Text Inversion are not able to perform very sophisticated transformations. Textual Inversion cannot achieve satisfactory results even with 30 minutes of tuning. Re-Imagen though gives reasonable outputs, the subject preservation is much weaker than SuTI. Imagic also generates reasonable outputs, however, its failure rate is still much higher than ours. DreamBooth however generates almost perfect images except for the ‘blurry’ text on the berry bowl. Through the comparison, we can observe remarkable improvement in the output image quality.

Skillset. We provide SuTI’s generation to showcase its ability in re-contextualization, novel view synthesis, art rendition, property modification, and accessorization. We demonstrate these different skills in the Appendix Figure 8. In the first row, we show that SuTI is able to synthesize the subjects with different art styles. In the second row, we show that SuTI is able to synthesize the different view angles of the given subject. In the third row, we show that SuTI can modify subjects’ facial expressions like ‘sad’, ‘screaming’, etc. In the fourth row, we show that SuTI can alter the color of a given toy. In the last two rows, we show that SuTI can add different accessories (hats, clothes, etc) to the given subjects. Further, we found that SuTI can even compose two skills together to perform highly complex image generation. As depicted in Figure 5, we show that SuTI can combine re-contextualization with editing/accessorization/stylization to generate high-quality images.

3 Model Analysis and Ablation Study

We further conducted a set of ablation studies to show factors that impact the performance of SuTI.

Quality of the expert dataset matters. We found that the Delta CLIP score is critical to ensure the quality of synthesized target images. Such a filtering mechanism is highly influential in terms of SuTI’s final performance. We evaluated several versions to increase the Δ\Delta threshold from None →\rightarrow 0.0 →\rightarrow ⋯\cdots →\rightarrow 0.025, we observe that the human evaluation can steadily increase from 0.54 to 0.82. Without such intensive filtering, the model’s overall human score can go to a very low level (54%). With an increasing Δ\Delta, although the size of the dataset GG keeps decreasing from 1.8M to around 500K, the model’s generation quality keeps improving until saturation. The empirical study indicates that Δ=0.02\Delta=0.02 strikes a good balance between the quality and quantity of the expert-generated dataset GG.

Further fine-tuning SuTI improves generation quality We note that our model is not exclusive to methods that requires further fine-tuning, such as DreamBooth . Instead, SuTI can be combined with DreamBooth naturally to achieve better quality subject-driven generation (dubbed as Dream-SuTI). Specifically, given KK reference images regarding a subject, we can randomly feed one image as the condition and use another differently sampled image as the target output. Through fine-tuning the SuTI model for 500 steps (without any auxiliary loss), the Dream-SuTI model can generate aligned and faithful results for a given subject. Table 3 shows a comparison of the Dream-SuTI, against SuTI and DreamBooth, suggesting that Dream-SuTI further improves the generation quality. Particularly, it improves the overall score from 0.82 to 0.87, yielding a 5% improvement over SuTI, and 10% improvement over DreamBooth. Since the fine-tuned Dream-SuTI model already trained on all subject images, only one subject image is needed to present during the inference time, which can further reduce the inference cost.

To gain better understanding of the quality, we show an example in Figure 7, where we pick a failure example from SuTI to investigate whether Dream-SuTI improves it. We observe that the DreamBooth does not have strong text alignment, while SuTI’s subject lacks fidelity (the generated robot uses legs instead of wheels). With further fine-tuning on subject images, Dream-SuTI is able to generate images not only faithful to the subject but also to the text description. However, we would like to note that such subject-driven fine-tuned model share the same drawback of a typical Dreambooth model, which can no longer generalize well to a general distribution objects and hence requiring a copy of model parameter per subject.

Related Work

Text-Guided Image Editing With the surge of diffusion-based models, have demonstrated the possibilities to manipulate given image without human intervention. Blended-Diffusion and SDEdit propose to blend the noise with the input image to guide the image synthesize process to maintain the original layout. Text2Live generates an edit layer on top of the original image/video input. Prompt-to-Prompt and Null-Text Inversion aims at manipulating the attention map in the diffusion model to maintain the layout of the image while changing certain subjects. Imagic propose an optimization based to achieve significant progress in manipulating visual details in a given image. InstructPix2Pix propose to distill image editing training pairs synthesized from Prompt-to-Prompt into a single diffusion model to perform instruction-driven image editing. Our method resembles InstructPix2Pix in a sense that we are training the model on expert-generated images. However, our synthesized data is generated generated by fine-tuned experts, which are mostly natural images. In contrast, the images from InstructPix2Pix are synthetic images. In the experiment section, we comprehensively compare with these existing models to show the advantage of our model, especially on more challenging prompts.

Subject-Driven Text-to-Image Generation Subject-Driven Image Generation tackles a new challenge, where the model needs to understand the visual subject contained in the demonstrations to synthesize totally new scene. Several GAN-based models pioneered to work on personalizing the image generation model to a particular instance. Later on, DreamBooth and Textual Inversion propose optimization-based approach to adapt image generation to a specific unseen subject. However, these two methods are time and space-consuming, which makes them unrealistic in real-world applications. Another line of work adopt retrieval-augmented architecture for subject-driven generation including KNN-Diffusion , Re-Imagen , however, these methods are trained with weakly-supervised data leading to much worse faithfulness. In this paper, we aim at developing an apprenticeship learning paradigm to train the image generation model with stronger supervision demonstrated by fine-tuned experts. As a result, SuTI can generate customized images about a specified subject, without requiring any test-time fine-tuning. There are some concurrent and related works focusing on specific visual domains such as human faces and / or animals. To our best knowledge, SuTI is the first subject-driven text-to-image generator that operates fully in-context, generalizing across various visual domains.

Conclusion

Our method SuTI has shown strong capabilities to generate personalized images instantly without test-time optimization. Our human evaluation indicates that SuTI is already better in the overall score than DreamBooth, however, we do identify a few weakness of our model: (1) SuTI’s generations are less diverse than DreamBooth, and our model is less inclined to transform the subjects’ poses or views in the new image. (2) SuTI is less faithful to the low-level visual details than DreamBooth, especially for more complex and often manufactured subjects such as ‘robots’ or ‘rc cars’ where the subjects contain highly sophisticated visual details that could be arbitrarily different from the examples inside the training dataset. In the future, we plan to investigate how to further improve these two aspects to make SuTI’s generation more diverse and detail-preserving.

Acknowledgement

We thank Boqing Gong, Kaifeng Chen for reviewing an early version of this paper in depth, with valuable comments and suggestions. We thank Neil Houlsby, Xiao Wang and also the PaLI-X team for providing early access to their Episodic WebLI data. We also thank Jason Baldbridge, Andrew Bunner, Nicole Brichtova for discussions and feedback on the project.

Broader Impact

Subject-driven text-to-image generation has wide downstream applications, like adapting certain given subjects into different contexts. Previously, the process was mostly done manually by experts who are specialized in photo creation software. Such manual modification process is time-consuming. We hope that our model could shed light on how to automate such a process and save huge amount of labors and training. The current model is still highly immature, which can fall into several failure modes as demonstrated in the paper. For example, the model is still prone to certain priors presented in certain subject classes. Some low-level visual details in subjects are not perfectly preserved. However, it could still be used as an intermediate form to help accelerate the creation process. On the flip side, there are risks with such models including misinformation, abuse and bias. See the discussion of broader impacts in for more discussion.

References

Appendix A Supplementary Material

To validate the effectiveness, we provide an ablation study to show that higher precision is more important than recall in training the apprentice model. Particularly, when the threshold is set to a lower number (e.g., 0.01 or 0.015), SuTI becomes less stable.

As our goal is to collect images of the same subject, we create an initial subject cluster by grouping all (image, alt-text) pairs that come from the same URL (∼\sim45M clustrers), and filter the cluster with less than 3 instances (∼\sim77.8% of the clusters). As a result, it leaves us with ∼\sim10M image clusters. We then apply the pre-trained CLIP ViT-L14 model to filter out 81.1% of clusters that has the average intra-cluster visual similarity between 0.820.82 and 0.980.98 to ensure the quality of clusters.

A.2 SuTI Skillset

We demonstrate the complete view of SuTI’s skillset in Figure 8, including styled subject generation, multi-view subject rendering, subject expression modification, subject colorization, and subject accessorization.

A.3 Failure Examples

Figure 9 show some failure examples of SuTI. We show several types of failure modes: (1) the model has a strong prior about the subject and hallucinates the visual details based on its prior knowledge. For example, the generation model believes ‘teapot’ should contain a ‘lift handle’. (2) some artifacts from the demonstration images are being transferred to the generated images. For example, the ‘bed’ from the demonstration is being brought to the generation, (3) the subject’s visual appearance is being modified through, mostly influenced by the context, like the ‘candle’ contains non-existing artifacts when contextualized in the ‘toilet’. These three failure modes constitute most of the generation errors. (4) The models are not particularly good at handling compositional prompts like the ‘bear plushie’ and ‘sunglasses’ example. In the future, we plan to work on how to improve these aspects.

A.4 More Qualitative Examples

We demonstrate more examples from DreamBench-v2 in the following figures: