One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, Jun Zhu

Introduction

Recently, we are witnessing a content-creation revolution driven by the rapid advances of generative modeling on multi-modal data. In particular, diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020; Song et al., 2021c) have shown an incredible ability to create high-fidelity and diverse data (Ramesh et al., 2022; Saharia et al., 2022; Rombach et al., 2022; Ho et al., 2022a; Popov et al., 2021), whose content aligns well with the input text condition.

However, these generative models are designed as bespoke systems, which only allow a single task. Actually, humans can generate various multi-modal content simultaneously, with arbitrary conditioning types. For example, artists can create paintings conditioned on texts, scenes, or just imagination and employ language ability to generate the caption of a photo. Toward a general generative system on multi-modal data, a unified training framework that can cover all types of multi-modal generative tasks (see Figure 1) is one of the fundamental components.

The task is solved by fitting a corresponding distribution in the view of probabilistic modeling. For instance, text-to-image generation can be formulated as learning the conditional distribution p(Image∣Text)p(\textrm{Image}|\textrm{Text}). A classical way to fit all relevant distributions is implicit – it first learns the joint distribution and then infers the marginal and conditional distributions by additional procedures (e.g., Markov Chain Monte Carlo (Srivastava & Salakhutdinov, 2012)), which is unaffordable on large-scale multi-modal data (Schuhmann et al., 2022).

In contrast, this paper presents a diffusion-based framework (dubbed UniDiffuser) that explicitly fits all relevant distributions in one model without introducing additional training or inference overhead. Our key insight is – learning diffusion models for all distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. For instance, a zero level indicates conditional generation given the corresponding modality, and a maximum level indicates unconditional generation of other modalities by ignoring the corresponding modality. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model (Ho et al., 2020) (see Figure 2) – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. Naturally, UniDiffuser is able to perform all kinds of generation (see Figure 1) in the same way as bespoken diffusion models. Moreover, UniDiffuser can perform the classifier-free guidance (Ho & Salimans, 2021) for free to improve the sample quality in both conditional and joint generation because UniDiffuser already models marginal distributions.

Besides the probabilistic modeling framework, a unified architecture that can handle input types of different modalities is another fundamental component in a general generative system. Notably, the emergence of Transformer (Vaswani et al., 2017; Dosovitskiy et al., 2021) and its applications on generative modeling (Bao et al., 2023a) provide a promising solution to capture interactions between modalities. Naturally, UniDiffuser employs a transformer-based backbone.

We implement UniDiffuser in the latent space (Rombach et al., 2022) with an additional CLIP encoder (Radford et al., 2021) for images and GPT-2 (Radford et al., 2019) decoder for texts on large-scale image-text data (Schuhmann et al., 2022). UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the corresponding bespoken models (e.g., Stable Diffusion and DALL⋅\cdotE 2) in representative tasks (e.g., text-to-image generation).

Background

Diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020) perturb the data by gradually injecting noise to data x0∼q(x0){\bm{x}}_{0}\sim q({\bm{x}}_{0}), which is formalized by a Markov chain:

where βt\beta_{t} is the noise schedule and αt=1−βt\alpha_{t}=1-\beta_{t}.

The data can be generated by reversing this process, where the reverse transition q(xt−1∣xt)q({\bm{x}}_{t-1}|{\bm{x}}_{t}) is approximated by a Gaussian model p(xt−1∣xt)=N(xt−1∣μt(xt),σt2I)p({\bm{x}}_{t-1}|{\bm{x}}_{t})={\mathcal{N}}({\bm{x}}_{t-1}|{\bm{\mu}}_{t}({\bm{x}}_{t}),\sigma_{t}^{2}{\bm{I}}). As shown by Bao et al. (2022b), the optimal mean under maximal likelihood estimation is

Conditional generation with diffusion models. In the case of conditional generation, we have paired data (x0,y0)∼q(x0,y0)({\bm{x}}_{0},{\bm{y}}_{0})\sim q({\bm{x}}_{0},{\bm{y}}_{0}), and we want to model the conditional data distribution q(x0∣y0)q({\bm{x}}_{0}|{\bm{y}}_{0}). The Gaussian model of the reverse process conditioned on y0{\bm{y}}_{0} is p(xt−1∣xt,y0)=N(xt−1∣μt(xt,y0),σt2I)p({\bm{x}}_{t-1}|{\bm{x}}_{t},{\bm{y}}_{0})={\mathcal{N}}({\bm{x}}_{t-1}|{\bm{\mu}}_{t}({\bm{x}}_{t},{\bm{y}}_{0}),\sigma_{t}^{2}{\bm{I}}). Similarly to Eq. (1), the optimal mean under maximal likelihood estimation is

Classifier-free guidance (CFG) (Ho & Salimans, 2021) is proposed to improve the sample quality of a conditional diffusion model. Specifically, it samples by linearly combining a conditional model and an unconditional one:

where ss is the guidance scale. The conditional and unconditional models share parameters by introducing a null token ∅\varnothing, i.e., ϵθ(xt,t)=ϵθ(xt,y0=∅,t){\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},t)={\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},{\bm{y}}_{0}=\varnothing,t).

Method

Section 3.1 presents UniDiffuser, a single diffusion model to capture the marginal, conditional, and joint distributions determined by multi-modal data simultaneously. Section 3.2 demonstrates how to perform classifier-free guidance (CFG) for free in conditional and joint sampling of UniDiffuser. For simplicity, we focus on two-modal data in this paper but UniDiffuser can be easily extended to more modalities.

Formally, suppose we have two modalities of data sampled from distribution q(x0,y0)q({\bm{x}}_{0},{\bm{y}}_{0}). We aim to design a diffusion-based model that is able to capture all relevant distributions determined by q(x0,y0)q({\bm{x}}_{0},{\bm{y}}_{0}), i.e., the marginal distributions q(x0)q({\bm{x}}_{0}) and q(y0)q({\bm{y}}_{0}), the conditional distributions q(x0∣y0)q({\bm{x}}_{0}|{\bm{y}}_{0}) and q(y0∣x0)q({\bm{y}}_{0}|{\bm{x}}_{0}), and the joint distribution q(x0,y0)q({\bm{x}}_{0},{\bm{y}}_{0}).

where (x0,y0)({\bm{x}}_{0},{\bm{y}}_{0}) is a random data point, $denotesconcatenation,denotes concatenation,{\bm{\epsilon}}^{x}andand{\bm{\epsilon}}^{y}aresampledfromstandardGaussiandistributions,andare sampled from standard Gaussian distributions, andt^{x}andandt^{y}areuniformlysampledfromare uniformly sampled from\{1,2,\dots,T\}$ independently. We call our method UniDiffuser because it captures multiple distributions in a unified way. We present the training algorithm in Appendix B.

The objective in Eq. (5) is as simple as the original DDPM in Eq. (2). Besides, for a single update of parameters, UniDiffuser only requires a single forward-backward calculation for multiple tasks (i.e., distributions), which is as efficient as the original DDPM. Although the gradient estimate of UniDiffuser has a slightly higher variance than the original DDPM due to two independent timesteps, we do not observe that UniDiffuser suffers from slower convergence.

UniDiffuser attempts to fit all distributions by one joint noise prediction network, requiring that the backbone can handle the mutual interaction between modalities and is scalable for large-scale data and multiple tasks. Inspired by the excellent performance of transformers on multi-modal representation learning at scale (Kim et al., 2021; Wang et al., 2022), we employ a transformer-based network in UniDiffuser, as detailed in Section 4.2.

Given a single joint noise prediction network, UniDiffuser can perform unconditional, conditional, and joint sampling according to a certain sampler (see Appendix B for the sampling algorithm). Notably, by setting the timesteps properly, the inference procedure of UniDiffuser is the same as the bespoken models. In comparison, learning a single joint distribution (Srivastava & Salakhutdinov, 2012; Hu et al., 2022) over multi-modal data requires additional procedures (e.g., Markov Chain Monte Carlo) to sample from the marginal or conditional distributions, which is unaffordable on large-scale multi-modal data (Schuhmann et al., 2022).

2 Classifier-Free Guidance for Free

Classifier-free guidance (CFG) (Ho & Salimans, 2021) combines a conditional and an unconditional model linearly during sampling (see Eq. (4)). It is simple yet effective to improve the sample quality and image-text alignment in diffusion models. Notably, CFG is directly applicable to the conditional and joint sampling of UniDiffuser without modifying the training process (see Figure 3 for results).

Formally, we denote the output of ϵθ{\bm{\epsilon}}_{\bm{\theta}} as the concatenation of ϵθx{\bm{\epsilon}}_{\bm{\theta}}^{x} and ϵθy{\bm{\epsilon}}_{\bm{\theta}}^{y}, i.e. ϵθ=[ϵθx,ϵθy]{\bm{\epsilon}}_{\bm{\theta}}=[{\bm{\epsilon}}_{\bm{\theta}}^{x},{\bm{\epsilon}}_{\bm{\theta}}^{y}], where we omit the input for simplicity. UniDiffuser can perform CFG for free in conditional sampling because it captures both the conditional and unconditional models. For example, we can generate x0{\bm{x}}_{0} conditioned on y0{\bm{y}}_{0} similarly to Eq. (4) as follows:

where ϵθx(xt,y0,t,0){\bm{\epsilon}}_{\bm{\theta}}^{x}({\bm{x}}_{t},{\bm{y}}_{0},t,0) and ϵθx(xt,ϵy,t,T){\bm{\epsilon}}_{\bm{\theta}}^{x}({\bm{x}}_{t},{\bm{\epsilon}}^{y},t,T) represent the conditional and unconditional models respectively, and ss is the guidance scale. In contrast to the original CFG, UniDiffuser does not need to specify a null token for parameter sharing.

CFG is also applicable to joint sampling. By setting tx=ty=tt^{x}=t^{y}=t, note that the joint score model can be equivalently expressed in the form of conditional models as follows:

where q(xt,yt)q({\bm{x}}_{t},{\bm{y}}_{t}) is the joint distribution of perturbed data at the same noisy level tt. Inspired by the above relationship between score functions, ϵθ(xt,yt,t,t){\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},{\bm{y}}_{t},t,t) can be viewed as approximating a pair conditional scores ∇xtlog⁡q(xt∣yt)\nabla_{{\bm{x}}_{t}}\log q({\bm{x}}_{t}|{\bm{y}}_{t}) and ∇ytlog⁡q(yt∣xt)\nabla_{{\bm{y}}_{t}}\log q({\bm{y}}_{t}|{\bm{x}}_{t}). In the same spirit of CFG, we can replace each conditional score by interpolating the joint model with the corresponding unconditional model as follows:

where ϵθx(xt,ϵy,t,T){\bm{\epsilon}}_{\bm{\theta}}^{x}({\bm{x}}_{t},{\bm{\epsilon}}^{y},t,T) and ϵθy(ϵx,yt,T,t){\bm{\epsilon}}_{\bm{\theta}}^{y}({\bm{\epsilon}}^{x},{\bm{y}}_{t},T,t) represent unconditional models. We summarize the formulation of CFG in UniDiffuser for all tasks in Appendix C.

UniDiffuser on Images and Texts

Images and texts are two of the most common modalities in daily life. Thus, it is representative to validate the effectiveness of UniDiffuser on the two modalities.

Our implementation is two-staged following (Rombach et al., 2022) (see Figure 4). First, we convert images and texts to continuous latent embeddings x0{\bm{x}}_{0} and y0{\bm{y}}_{0} via image and text encoders and introduce two decoders for reconstruction, as presented in Section 4.1. Second, we train UniDiffuser parameterized by a transformer on the latent embeddings x0{\bm{x}}_{0} and y0{\bm{y}}_{0}, as presented in Section 4.2.

The image and text encoder-decoders are illustrated in Figure 4 (a). Below we provide their details.

Remark. We observe that the latent embeddings of both image and text already have similar and reasonable numerical ranges. Specifically, they are concentrated within the range of $andexhibitapproximatelynormaldistributionswithcomparablemeanandvariancevalues(imagemodality:mean=and exhibit approximately normal distributions with comparable mean and variance values (image modality: mean =0.0269,standarddeviation=, standard deviation =0.7919;textmodality:mean=; text modality: mean =0.0127,standarddeviation=, standard deviation =0.5957$). As a result, we did not apply additional normalization to them. For more modalities, we can similarly convert them to continuous latent features through encoders that have regularization on the latent space. This makes it easy for all modalities to have similar ranges after normalization. Besides, obtaining high-quality encoders and decoders is relatively straightforward and can be achieved with a smaller amount of data. For example, the dataset size of the image encoder and decoder is less than 1% of UniDiffuser’s. Therefore, in practice, we can efficiently train high-quality encoders and decoders for each modality at a modest cost if needed.

2 Transformer as Joint Noise Prediction Network

We train a joint noise prediction network on the embeddings obtained in Section 4.1 according to Eq. (5). It is natural to employ a transformer-based backbone in UniDiffuser to handle inputs from different modalities. In particular, we adopt U-ViT (Bao et al., 2023a), a recently proposed transformer for conditional diffusion models. The original U-ViT is characterized by treating all inputs including the data, the condition, and the timestep as tokens, and employing long skip connections between shallow and deep layers. In UniDiffuser, we slightly modify U-ViT by treating the two modalities of data and their corresponding timesteps as tokens. Besides, we empirically find that the pre-layer normalization (Xiong et al., 2020) in the original U-ViT causes overflow easily when trained with mixed precision. A simple solution is to use the post-layer normalization (Vaswani et al., 2017) and add a layer normalization after concatenating a long skip connection, which stabilizes the training of UniDiffuser. We illustrate the backbone in Figure 4 (b) and present more details in Appendix D.

Related Work

Multi-modal generative modeling. Many prior work on multi-modal generative modeling can be formalized as learning a conditional distribution. Representative applications include text-to-image generation (Ramesh et al., 2021; Ding et al., 2021; Ramesh et al., 2022; Nichol et al., 2022; Saharia et al., 2022; Yu et al., 2022; Gu et al., 2022; Xu et al., 2018; Rombach et al., 2022), text-to-video generation (Ho et al., 2022a), text-to-speech generation (Chen et al., 2021; Popov et al., 2021) and image captioning (i.e., image-to-text generation) (Mokady et al., 2021; Chen et al., 2022). Such models are specially designed for a single task. In addition to learning a conditional distribution, Hu et al. (2022) aims to learn the joint distribution of image and text data via a discrete diffusion model (Gu et al., 2022). However, its scalability is unexplored.

The most related prior work is Versatile Diffusion (VD) (Xu et al., 2022), which employs a multi-flow architecture and is trained for multiple generation tasks in the traditional multi-task framework, which requires multiple feed-forward to compute losses for all tasks and carefully tuned gradient multipliers for different layers during training. In contrast, UniDiffuser provides an elegant solution based on the insightful unified view of training diffusion models. As a result, UniDiffuser is simpler (with a single training loss), more efficient to train (with a single forward-backward per update), and can handle more tasks (able to perform joint sampling) without the need for complex tricks. Besides, UniDiffuser outperforms VD in both image-to-text and text-to-image generation tasks in terms of the FID and CLIP scores in our experiments (see Section 6), suggesting that the time-condition strategy in UniDiffuser is statistically more efficient than the multi-task one in VD.

Multi-modal representation learning aims to learn features for different modalities that can be transferred to downstream tasks. Vision-and-language pretraining (VLP) is at the front. VLP can employ different strategies, such as contrastive learning (Radford et al., 2021), masked data modeling (Wang et al., 2022), and a combination of multiple losses (Kim et al., 2021; Li et al., 2022; Bao et al., 2022c). A transformer is often employed to fuse the two modalities. This work implies that a transformer is also effective for multi-modal generative modeling.

Diffusion models are initially proposed by Sohl-Dickstein et al. (2015). Recently, Ho et al. (2020) introduce a noise prediction formulation, and Song et al. (2021c) introduce a stochastic differential equation formulation for learning diffusion models. Diffusion models are able to generate high-quality images (Dhariwal & Nichol, 2021),audios (Chen et al., 2021; Kong et al., 2021), videos (Ho et al., 2022b), point clouds (Luo & Hu, 2021) and molecular conformations (Hoogeboom et al., 2022; Bao et al., 2023b). Other improvements in diffusion models include fast sampling (Song et al., 2021a; Bao et al., 2022b; Salimans & Ho, 2022; Lu et al., 2022b, c) and improved training and sampling techniques (Nichol & Dhariwal, 2021; Song et al., 2021b; Kingma et al., 2021; Vahdat et al., 2021; Zhao et al., 2022; Bao et al., 2022a; Lu et al., 2022a; Karras et al., 2022).

Experiments

We present the experimental setup in Section 6.1. We show the ability of UniDiffuser to perform multiple generation tasks and directly compare it with existing large models in Section 6.2. We further demonstrate that UniDiffuser naturally supports applications like data variation, blocked Gibbs sampling between modalities (see Section 6.3), and interpolation between images in the wild (see Section 6.4).

Dataset. We use three subsets of LAION-5B (Schuhmann et al., 2022) following Stable Diffusion (Rombach et al., 2022). The first one is laion2B-en, which contains around 2B image-text pairs with English captions. The second one is laion-high-resolution, which contains around 170M image-text pairs with image resolution ≥\geq1024 and multilingual captions. The third one is laion-aesthetics v2 5+, which is a subset of laion2b-en containing around 600M image-text pairs with high visual quality. Following Stable Diffusion, we additionally filter laion-aesthetics v2 5+ to images with resolution ≥\geq512 and an estimated watermark probability <<0.5, leading to around 193M preserved pairs. For image normalization, we follow the standard practice in diffusion models by normalizing the image values from the range of toto. Since the texts in LAION-5B are quite noisy, we further clean texts in the laion-aesthetics v2 5+ subset by removing URLs, HTML tags, emails, contents in brackets, quotes except ’s, and symbols except , . ? !. Before inputting the text into CLIP, we tokenize the preprocessed text using CLIP’s built-in tokenizer, which is based on byte-level Byte-Pair-Encoding (Radford et al., 2021).

Training and Sampling. The training is multiple-staged following Stable Diffusion (Rombach et al., 2022). In the first stage, we train 250K steps at 256×\times256 resolution on laion2B-en with a batch size of 11264 and 5K warm-up steps. In the second stage, we fine-tune the model with 200K steps at 512×\times512 resolution on laion-high-resolution with a batch size of 2112 and 5K warm-up steps. In the last stage, we resume from the last checkpoint of the second stage (including both weights of the model and states of the optimizer), and train 220K steps at 512×\times512 resolution on laion-aesthetics v2 5+ with a batch size of 2112. Following Bao et al. (2023a), we use the AdamW optimizer (Loshchilov & Hutter, 2019) with a learning rate of 2e-4, a weight decay of 0.03 and running coefficients of (β1,β2)=(0.9,0.9)(\beta_{1},\beta_{2})=(0.9,0.9) in all stages. We reduce the learning rate by a factor of 10 and continue training whenever the validation loss does not decrease. We train with mixed precision for efficiency. When U-ViT is trained at 256×\times256 resolution, we interpolate the positional embeddings related to images via bilinear interpolation. The training takes around 28 days on 88 A100 (80GB) GPUs. We use DPM-Solver (Lu et al., 2022b, c) with 50 steps in all experiments.

Baseline. To our knowledge, Versatile Diffusion (VD) (Xu et al., 2022) is the most direct competitor for general-purpose multi-modal generation (see details in Section 5). We directly compare to VD in all experiments if possible. The results of VD are reproduced by us upon official code because there is no quantitative result in the original paper.

Evaluation. For text-to-image generation, we report the FID (Heusel et al., 2017) and CLIP score (Radford et al., 2021) on the MS-COCO validation set (Lin et al., 2014) to measure the image fidelity and image-text alignment respectively. Following the literature, we randomly draw 10K and 30K prompts from the MS-COCO validation set to calculate FID and the CLIP score on generated images. For image-to-text generation, we report the CLIP score to measure the image-text alignment. Specifically, we randomly draw 10K images to calculate the score on generated texts.

2 Main Results

We first systematically compare with the most direct baseline Versatile Diffusion (VD), which is a general-purpose generative model, in both text-to-image and image-to-text generation. Quantitatively, UniDiffuser outperforms VD consistently in both tasks under all metrics and guidance CFG scales, as presented in Figure 5 and Figure 6. The empirical results demonstrate the effectiveness (in addition to the simplicity, efficiency, and generality) of UniDiffuser compared to VD (see details in Section 5). Qualitatively, Figure 7 presents samples in text-to-image generation, and UniDiffuser aligns image and text better than VD. See more results including image-to-text generation in Appendix G.

We also compare with bespoken systems designed for text-to-image generation w.r.t. zero-shot FID on MS-COCO in Table 1. Although UniDiffuser is designed to handle multiple generation tasks, its performance on the single text-to-image generation task is comparable to bespoken diffusion models such as Stable Diffusion and outperforms famous diffusion models like DALL⋅\cdotE 2.

Finally, we present examples of joint, conditional, and unconditional generation in Figure 1 (a-e) to show the generality of UniDiffuser. See more examples in Appendix A.

3 Data Variation and Gibbs Sampling

UniDiffuser naturally supports applications such as image variation and text variation. For example, given a source image, we can firstly perform image-to-text generation to obtain a description of the image, and then perform text-to-image generation with this description as input to obtain a new image with similar semantics but different contents. In Figure 1 (f-g), we present examples on image and text variation. Furthermore, we can perform blocked Gibbs sampling to see how images and texts are translated to each other by chaining conditional distributions modeled by UniDiffuser. We present examples in Figure 1 (h). More samples on data variation and blocked Gibbs can be found in Appendix A.

4 Interpolation between Two Images in the Wild

UniDiffuser can also perform interpolation between two images in the wild. Specifically, we firstly perform image-to-text generation to obtain the latent text embeddings of the two images via the deterministic DPM-Solver with the same Gaussian noise as the initial state for both images. Then we perform a noise injection process via DPM-Solver to get a noisy version of the latent image embeddings given the two latent text embeddings. We perform spherical linear interpolation between the latent text embeddings and the noisy version of the latent image embeddings to obtain intermediate states. Finally, with the text intermediate states as the condition and the image intermediate states as the initial state, we generate the final images by DPM-solver. See Appendix F for a formalized algorithm of the interpolation procedure. We present examples in Figure 1 (i) and more examples can be found in Appendix A.

Conclusion

We propose UniDiffuser, a general-purpose multi-modal probabilistic framework based on insights of unifying training of diffusion models for different distributions. UniDiffuser is able to perform various generation tasks via one model with minimal modification of the original diffusion models. Empirical results on image-text data show the effectiveness of UniDiffuser compared to large existing models. UniDiffuser also enables semi-supervised learning and learning on more modalities, which are left as future work. Currently, the text generated by our implementation is not that smooth, mainly because the text data is noisy.

UniDiffuser has high potential to improve multiple tasks: by fitting multiple tasks with one single transformer network, UniDiffuser can be much easier to further improve all tasks simultaneously (e.g., by increasing parameter scale and data scale) and maintain under the large-scale pre-training regime. Any further improvement/optimization of the underlying single network can seamlessly benefit all tasks.

Social Impact: We believe UniDiffuser can advance real-world applications with generated content due to its generality. However, it is worth noting that large-scale multi-modal generative models may have consequences like “deep-fakes”. We watermark all images sampled from the model and will provide a systematical protocol to relieve the problem before releasing the code and model.

Acknowledgements

This work was supported by NSF of China Projects (Nos. 62061136001, 61620106010, 62076145, U19B2034, U1811461, U19A2081, 6197222); Beijing Outstanding Young Scientist Program NO. BJJWZYJH012019100020098; a grant from Tsinghua Institute for Guo Qiang; the High Performance Computing Center, Tsinghua University; the Fundamental Research Funds for the Central Universities, and the Research Funds of Renmin University of China (22XNKJ13). C. Li was also sponsored by Beijing Nova Program. J.Z was also supported by the New Cornerstone Science Foundation through the XPLORER PRIZE.

References

Appendix A More Examples

Below we present more samples of UniDiffuser on joint generation (Figure 8), text-to-image generation (Figure 9), image-to-text generation (Figure 10), unconditional image generation (Figure 11), unconditional text generation (Figure 13), image variation (Figure 12), text variation (Figure 14), blocked Gibbs sampling (Figure 15) and interpolation (Figure 16).

Appendix B The Training and Sampling Algorithms

In Algorithm 1, we present the training algorithm of UniDiffuser. In Algorithm 2,3,4, we present all sampling procedure of UniDiffuser by taking the DDPM sampler (Ho et al., 2020) as an example. Note that any other learning-free efficient sampler, such as DDIM (Song et al., 2021a), Analytic-DPM (Bao et al., 2022b) and DPM-Solver (Lu et al., 2022b, c), is directly applicable.

Appendix C Summary of Classifier-Free Guidance Models

As mentioned in Section 3.2, UniDiffuser can perform classifier-free guidance for free for conditional and joint generation. We summarize models for classifier-free guidance in Table 2. We also include the models for unconditional sampling.

Appendix D Details of the U-ViT

We train a U-ViT for the joint noise prediction network, whose detailed configuration is presented in Table 3.

Appendix E Details of the GPT-2 Text Decoders

where ϕ{\bm{\phi}} denotes the parameters of the linear layer and the GPT-2.

We finetune the 124M parameter GPT-2 text decoder on texts of LAION-2B-en dataset (Schuhmann et al., 2022), which contains 2.3B image-text pairs. Following ClipCap (Mokady et al., 2021), we use the AdamW optimizer (Loshchilov & Hutter, 2019) with a learning rate of 2e-5 and 5K warm-up steps. We train the decoder with 235K steps using a batch size of 768. When generating texts, we use the beam search strategy with a beam size of 5 and a maximum length of 67.

In our experiments, the embedding dimension of y0{\bm{y}}_{0} is set to 64, and it still reconstructs the text well. Indeed, we get a BLEU-1 (Papineni et al., 2002) score of 0.969 and a BLEU-4 score of 0.894 between reference texts and reconstructed texts on MS-COCO test set (Karpathy split (Karpathy & Fei-Fei, 2015)). We present some examples in Figure 18, where the input texts are reconstructed very well.

Appendix F The Interpolation Algorithm

We present the formalized procedure of interpolating two images using UniDiffuser in Algorithm 5.

Appendix G Comparison of Examples

In Figure 19 and Figure 20, we present more examples on text-to-image generation of UniDiffuser and Versatile Diffusion. Samples generated by UniDiffuser align better with the texts than VD. In Figure 21, we present examples on image-to-text generation of UniDiffuser and Versatile Diffusion. Samples generated by UniDiffuser align better with the images than VD.

Appendix H Efficiency Comparison

In Table 4, we compare the model size, inference time and memory (for generating 10 samples with 25 denoising steps on one A100 80GB GPU), and the training cost of UniDiffuser with other bespoke and general-purpose models. Results of other methods are obtained according to the official code or paper when they are available.

UniDiffuser is more efficient than both Stable Diffusion and Versatile Diffusion in terms of inference time and memory. Besides, compared to Stable Diffusion, UniDiffuser introduces only 10% extra parameters to support five tasks (i.e., image, text, text-to-image, image-to-text, and image-text pair generation) with comparable training cost. Compared to the general-purpose model Versatile Diffusion, UniDiffuser has fewer parameters, while achieving superior results (as presented in the main paper).

Appendix I Licences

LAION-5B (Schuhmann et al., 2022): Creative Common CC-BY 4.0 license

MS-COCO (Lin et al., 2014): Creative Commons Attribution 4.0 License

GPT-2 (Radford et al., 2019): MIT License

Image autoencoder from Stable Diffusion (Rombach et al., 2022): CreativeML Open RAIL-M License