ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis

Patrick Esser, Robin Rombach, Andreas Blattmann, Björn Ommer

Introduction

Spurred by the increasingly popular attention mechanism, a remarkably simple principle has driven progress in deep generative modeling over the past few years: Factorizing the likelihood of the data in an autoregressive (AR) fashion

and subsequently learning the conditional transition probabilities with an expressive neural network such as a transformer . The success of this approach is evident in domains as diverse as language modeling , music generation , neural machine translation , and (conditional) image synthesis . However, especially for the latter task of image synthesis, which is also the focus of this work, the high dimensionality and redundancy present in the data challenges the direct applicability of this approach.

Missing Bidirectional Context Autoregressive models which represent images as a sequence from the top-left to the bottom-right have demonstrated impressive performance in sampling novel images and completing the lower half of a given image . However, the unidirectional, fixed ordering of sequence elements not only imposes a perceptually unnatural bias to attention in images by only considering context information from left or above. It also limits practical applicability to image modification: Imagine that you only have the lower half of an image and are looking for a completion of the upper half then these models fail at this minor variation of the completion task. The importance of contextual information from both directions has also been recognized in the context of language modeling . However, simply allowing bidirectional context as in does not provide a valid factorization of the density function for a generative model. Furthermore, the sequential sampling strategy introduces a gap between training and inference, as training relies on so-called teacher-forcing (where ground truth is provided for each step) and inference is performed on previously sampled tokens. This exposure bias can introduce significant accumulations of errors during the generation process, affecting sample quality and coherence .

Global Context & Control via Multinomial Diffusion We propose a coarse-to-fine approach that addresses the unidirectional bias of generative autoregressive models and their exposure bias as well as the lacking global context. We formulate learning the data density as a hierarchical problem. A coarser stage provides compressed contextual side information about the entire image for the autoregressive process on the next finer stage. We utilize a diffusion process to gradually eliminate information and compress the data, yielding a hierarchy of increasingly abstract and compact representations. The first scale of this approach is a discrete representation learning task (cf. ). Subsequently, we further compress this learned representation via a fixed, multinomial diffusion process . We then invert this process by training a Markov chain to recover the data from this hierarchy. Each Markovian transition is modeled autoregressively but it simultaneously attends to the preceding state in the hierarchy, which provides crucial global context to each individual autoregressive step. As each of this steps can also be interpreted as learning a denoising cloze task , where missing tokens at the next finer stage are “refilled” with a bidirectional encoder and an autoregressive decoder, we dub our approach ImageBART.

Contributions of our work Our approach tackles high-fidelity image synthesis with autoregressive models by learning to invert a fixed multinomial diffusion process in a discrete space of compact image representations to successively introduce context. This reduces both the often encountered exposure bias of AR models and also enables locally controlled, user-interactive image editing. Additionally, our model effectively handles a variety of conditional synthesis tasks and our introduced hierarchy corresponds to a successively compressed image representation. We observe that our model sample visually plausible images while still enabling a trade-off between reconstruction capability and compression rate.

Related Work

Latent Variable Models Among likelihood-based approaches, latent variable models represent a data distribution with the help of unobserved latent variables. For example, Variational Autoencoders (VAEs) encode data points into a lower dimensional latent variable with a factorized distribution. This makes them easy to sample, interpolate and modify . In a conditional setting , latent variables which are independent from the conditioning lead to disentangled representations . A hierarchy of latent variables gives mutli-scale representations of the data. Unfortunately, even the deepest instantiations of these models lack in sample quality compared to other generative models and are oftentimes restricted to highly regular datasets.

Autoregressive Models AR models represent a distribution as a product of conditional, learnable factors via the chain rule of probability densities. While this makes them powerful models for density estimation , their samples often lack global consistency. Especially on image data modeled with convolutional architectures , this has been attributed to a locality bias of convolutional neural networks (CNNs) which biases the model towards strong local correlations between neighboring pixels at the expense of a proper modeling of coherence . This leads to samples resembling texture patterns without discernible global structure. Attempts to fix this properties by including explicit latent variables have not been overly successful, mainly due the expressiveness of AR models, providing little incentive for learning additional latent variables.

Generative Models on Improved Representations Another successful line of work first learn an improved image representation and subsequently learn a generative model for this representation . Most works learn a discrete representation which is subsequently modeled autoregressively but approaches using continuous representations in combination with VAEs , or normalizing flows , exist too. Learning a compact representation enables the use of transformers for autoregressive modeling , which avoids the locality bias of CNNs, can be used for the synthesis of complex scenes conditioned on text as in DALL-E , and, when combined with adversarial learning , enables sampling of coherent high-resolution images . However, AR modeling of a learned representation still limits applications compared to latent variable models. Their samples can still exert artifacts resulting from a sequential modeling of components, and, since these models are always trained by “teacher-forcing”, they are susceptible to an exposure bias .

Diffusion Probabilistic Models Diffusion probabilistic models revert a fixed, diffusion process with a learned Markov Chain . Being directly applied in pixel space, however, downstream analysis reveals that these models tend to optimize subtle details of the modeled data, which have little contribution to the sample quality , particularly hindering applications on high-resolution and -complexity datasets. By using a multinomial diffusion process (recently generalized by ) on a compressed, discrete representation of images, we circumvent these issues. Diffusion probabilistic models require a very large number of diffusion steps in order to model the reverse process with a model distribution that factorizes over components. Because our approach uses autoregressively factorized models for the reverse process, we can reduce the required number of steps and obtain significant improvements in sampling speed and the ability to model complex datasets.

Method

We use L1L_{1} to learn a compressed and discrete representation of images, such that subsequent stages of the hierarchy do not need to model redundant information (Sec. 3.2). With Lt,t>1L_{t},t>1 we learn a model that can rely on global context from a coarser representation xtx_{t} to model the representation xt−1x_{t-1} (Sec. 3.3). See Fig. 1 for an overview of the proposed model.

2 Learning a compact, discrete representation for images

Here, frecf_{rec} denotes the perceptual similarity metric (known as LPIPS) and DϕD_{\phi} denotes a patch-based adversarial discriminator . Note that, due to the deterministic training, the likelihood in Eq. (3) is likely to be degenerate. DϕD_{\phi} is optimized to differentiate original images x0x_{0} from their reconstruction Gθ(x1)G_{\theta}(x_{1}) using simultaneous gradient ascent, such that the objective for learning the optimal parameters {θ∗,ϕ∗}\{\theta^{*},\phi^{*}\} reads:

The optimization of θ\theta via this objective includes the parameters of the encoder and decoder in addition to the parameters of the learned codebook, trained via the codebook loss LcbL_{cb} as in .

3 Parallel learning of hierarchies

Under suitable choices for pθ,qθp_{\theta},q_{\theta}, one can directly optimize these chains over ∑tLt\sum_{t}L_{t}. However, the objectives LtL_{t} of the hierarchy levels are coupled through the forward chain qθq_{\theta}, which makes this optimization problem difficult. With expressive reverse models pθt−1p^{t-1}_{\theta}, the latent variables xtx_{t} are often ignored by the model and the scale of the different level-objectives can be vastly different, resulting in a lot of gradient noise that hinders the optimization . In the continuous case, reweighting schemes for the objective can be derived based on a connection to score matching models . However, since we are working with a discrete x1x_{1}, there is no analogue available.

While we could follow the approach taken for the first level and sequentially optimize over the objectives LtL_{t}, this is a rather slow process since each level t−1t-1 needs to be converged before we can start solving level tt. However, this sequential dependence is only introduced through the forward models qθtq^{t}_{\theta} and since qθ1q^{1}_{\theta} already learns a strong representation, we can choose simpler and fixed, predefined forward processes for qθt,t>1q^{t}_{\theta},t>1. The goal of these processes, i.e., generating a hierarchy of distributions by reducing information in each transition, can be readily achieved by, e.g., randomly masking , removing or replacing a fraction of the components of xt−1x_{t-1}.

This process of randomly replacing a fraction βt\beta_{t} of the components with random entries can be described as a multinomial diffusion process , a natural generalization of binomial diffusion . The only parameter θ\theta of qθtq^{t}_{\theta} is therefore βt\beta_{t}, which we consider to be fixed. Using the standard basis e(k)=(δjk)j=1Ke(k)=(\delta_{jk})_{j=1}^{K}, the forward process can be written as a product of categorical distributions C\mathcal{C} specified in terms of the probabilities over the codebook indices:

This enables computation of the posterior qθ(xt−1∣xt,x1)=qθt(xt∣xt−1)qθ(xt−1∣x1)qθ(xt∣x1)q_{\theta}(x_{t-1}|x_{t},x_{1})=\frac{q^{t}_{\theta}(x_{t}|x_{t-1})q_{\theta}(x_{t-1}|x_{1})}{q_{\theta}(x_{t}|x_{1})} for t>2t>2, and, using the fact that qθ1q^{1}_{\theta} is deterministic, we can rewrite LtL_{t} as

such that the KL term can now be computed analytically for t>2t>2. For t=2t=2, we use a single sample Monte-Carlo estimate for the maximum likelihood reformulation, i.e.

Finally, we set pθTp^{T}_{\theta} to be a uniform distribution. This completes the definition of the reverse chain pθp_{\theta}, which can now be started from a random sample for xT∼pθT(xT)x_{T}\sim p^{T}_{\theta}(x_{T}), denoised sequentially through xt−1∼pθt−1(xt−1∣xt)x_{t-1}\sim p^{t-1}_{\theta}(x_{t-1}|x_{t}) for t=T,…,2t=T,\dots,2, and finally be decoded to a data sample x0=G(x1)x_{0}=G(x_{1}).

Under what conditions can we recover the true data distribution? By rewriting ∑tLt\sum_{t}L_{t}, we can see from

that this is possible as long as all reverse models are expressive enough to represent the true reverse processes defined by qθq_{\theta}. For the first level, we can ensure this by making x1x_{1} large enough such that the reconstruction error becomes negligible. For the diffusion process, previous image models relied on the fact that, in the limit βt→0\beta_{t}\to 0, the form of the true reverse process has the same functional form as the forward diffusion process . In particular, this allows modeling of the reverse process with a distribution factorized over the components. However, to make qθT−1q^{T-1}_{\theta} close to a uniform distribution requires a very large TT (in the order of 1000 steps) with small βt\beta_{t}. Training such a large number of reverse models is only feasible with shared weights for the models, but this requires a delicate reweighting of the objective and currently no suitable reweighting is known for the discrete case considered here.

Thus, to be able to recover the true data distribution with a modest number of reverse models that can be trained fully parallel, and without weight-sharing, we model each reverse process autoregressively. We use an encoder-decoder transformer architecture , such that the decoder models the reverse process for xt−1x_{t-1} autoregressively with the help of global context obtained by cross-attending to the encoder’s representation of xtx_{t} as visualized in Fig. 1. Note that the need for autoregressive modeling gets reduced for small βt\beta_{t}, which we can adjust for by reducing the number of decoder layers compared to encoder layers. The use of the compression model described in Sec. 3.2, however, allows to utilize full-attention based transformer architectures to implement the autoregressive scales.

Experiments

Sec. 4.1 evaluates the quality ImageBART achieves in image synthesis. Since we especially want to increase the controllability of the generative process, we evaluate the performance of ImageBART on class- and text-conditional image generation in Sec. 4.2. The ability of our approach to attend to global context enables a new level of localized control which is not possible with previous, purely autoregressive approaches as demonstrated in Sec. 4.3. Finally, Sec. 4.4 presents ablations on model and architecture choices.

In this section we present qualitative and quantitative results on images synthesized by our approach. We train models at resolution 256×256256\times 256 for unconditional generation on FFHQ , LSUN -Cats, -Churches and -Bedrooms and on class-conditional synthesis on ImageNet (cIN) .

Effective Discrete Representations Learning the full hierarchy as described in Eq. (2) and without unnecessary redundancies in the data requires to first learn a strong compression model via the objective in Eq. (4). demonstrated how to effectively train such a model and we directly utilize the publicly available pretrained models. For training on LSUN, we finetune an ImageNet pretrained model for one epoch on each dataset. As the majority of codebook entries remains unused, we shrink the codebook to those entries which are actually used (evaluated on the validation split of ImageNet) and assign a random entry for eventual outliers. This procedure yields an effective, compact representation on which we subsequently train ImageBART.

Training Details As described in Sec. 3.3, we use an encoder-decoder structure to model the reverse Markov Chain pθt−1(xt−1∣xt),  t<Tp^{t-1}_{\theta}(x_{t-1}|x_{t}),\;t<T, where the encoder is a bidirectional transformer model and decoder is implemented as an AR transformer. As the context for the last scale is pure noise, we employ a decoder-only variant to model pθT−1(xT−1∣xT)p^{T-1}_{\theta}(x_{T-1}|x_{T}). Furthermore, to account for the different complexities of the datasets, we adjust the number of multinomial diffusion steps for each dataset accordingly. For FFHQ we choose a chain of length T=3T=3, such that the total model consists of (i) the compression stage and (ii) n=2n=2 transformer models trained in parallel via the objective described in Eq.(7). Similarly, we set n=3n=3 for each of the LSUN models and n=5n=5 for the ImageNet model.

Results For each of these settings, Fig. 2 depicts samples of size 256×256256\times 256 generated with ImageBART and a single pass through the learned Markov Chain, demonstrating that our model is able to produce realistic and coherent samples. This is further confirmed by a quantitative analysis in Tab. 1, where we compare FID scores of competing likelihood-based and score-based methods such as TT and DDPM . Regarding other works on diffusion models such as and operating directly in pixel space, we observe that these approaches perform roughly equivalently well in terms of FID for datasets of low complexity (e.g. LSUN-Bedrooms and-Churches). For more complex datasets (LSUN-Cats, cIN), however, our method outperforms these pixel-based approaches, which can also be seen qualitatively on the right in Tab. 1. See Fig. 20 for a comparison on ImageNet.

2 Conditional Markov Chains for Controlled Image Synthesis

Being a sequence-to-sequence model, our approach allows for flexible and arbitrary conditioning by simply preprending tokens, similar to . More specifically, each learned transition pθt−1(xt−1∣xt,c), t>1p^{t-1}_{\theta}(x_{t-1}|x_{t},c),\>t>1 of the Markov chain is then additionally conditioned on a representation cc, e.g. a single token in the case of the class-conditional model of Sec. 4.1. Note that the compression model pθ0p^{0}_{\theta} remains unchanged.

Besides class-conditional modeling on ImageNet, we also learn a text-conditional model on Conceptual Captions (CC) . We obtain cc by using the publicly available tokenizer of the CLIP model , yielding a conditioning sequence of length 77. To model the dataset, we choose T=5T=5 and thus train n=4n=4 transformer models independently. For the pθ0p^{0}_{\theta}, we directly transfer the compression model from Sec. 4.1, trained on the ImageNet dataset.

Fig. 3 visualizes synthetic samples obtained with this model for various “image-cloze” tasks. Our resulting model is able to attend to semantic variations in the conditioning sentence (e.g. a change of weather for imagery of mountains) and renders the corresponding images accordingly. In Tab. 2, we evaluate FID and Inception Scores (IS) to measure the quality of synthesized images, as well as cosine similarity between CLIP embeddings of the text prompts and the synthesized images to measure how well the image reflects the text. ImageBART improves all metrics upon . Fig. 21 in the supplement provides corresponding qualitative examples for user-defined text inputs.

Resolutions Beyond 256×256\boldsymbol{256\times 256} Pixels. Our approach is not restricted to generating images of size 256×256256\times 256 pixels. Although trained on a fixed resolution, we can apply our models in a patch-wise manner, where we use the sliding attention window of for each scale t>0t>0. As we now incorporate more and more global context while decoding with the Markov chain (which can be thought of as widening a noisy receptive field), ImageBART is able to render consistent images in the megapixel regime. See for example Fig. 4, where we use our text-conditional model to render an image of size 300×1800300\times 1800 pixel and interpolate between two different text prompts. More examples, especially also for semantically guided synthesis, can be found in Sec. A.2.

3 Beyond Conditional Models: Local Editing with Autoregressive Models

Recent autoregressive approaches, which use a CNN to learn a discrete representation , partially alleviate the issues of pixel-wise autoregressive models by working on larger image patches. However, as we show in Fig. 5, even approaches which use adversarial learning to maximize the amount of context encoded in the discrete representation cannot produce completions of the upper half of an image which are consistent with a given lower half.

While our approach also models each transition autoregressively from the top-left to the bottom-right, the ability to attend to global context from the previous scale enables consistent completions of arbitrary order, e.g. right-to-left. To achieve this, we mask the diffusion process as described in Sec. A.3. For a user-specified mask mm (e.g. the upper half of an image as in Fig. 5), this results in a forward-backward process pθt−1∣t−1,mp^{t-1|t-1,m}_{\theta}, which, by definition, leaves the unmasked context intact. The reverse process then denoises the unmasked entries to make them consistent with the given context.

Fig. 5 (bottom) visualizes this mixing process, where we use a model with T=3T=3. The first column shows the masked input. To start the process we set all masked entries to random entries. The first two columns then show (decoded) samples from the masked reverse processes pθ2,mp^{2,m}_{\theta} and pθ1,mp^{1,m}_{\theta}, which still display inconsistencies. The remaining columns show the trajectory of the process pθ1∣1,mp^{1|1,m}_{\theta}, which demonstrates how the model iteratively adjusts its samples according to the given context until it converges to a globally consistent sample. For illustration, we show the analog trajectory obtained with , but because it can only attend to unidirectional context, this trajectory is equivalent to a sequence of independent samples and therefore fails to achieve global consistency.

The masked process can be used with arbitrary masks, which enables localized image editing with free, hand-drawn masks as shown in Fig. 6. Note that our model does not need to be trained specifically for this task, which also avoids generalization problems associated with training on masks . Combining this property with the conditional models from Sec. 4.2 allows for especially interesting novel applications, where local image regions are modified based on user specified class or text prompts, as shown in Fig. 7.

4 Ablations

On the Number of Diffusion Steps In this section we analyze the effect of varying the number of diffusion steps (denoted by TT). To do so, we perform an experiment for unconditional training on the FFHQ dataset, where we train a Taming Transformers (TT) baseline (corresponding to the case T=2T=2 within our framework) with 800M parameters and three variants of ImageBART with T=3T=3 (2x400M), T=5T=5 (4x200M) and T=9T=9 (8x100M), respectively. Note that for a fair comparison, all models use the same first level for compression, and we fix the number of remaining parameters to 800M and distribute them equally across all scales. All models were trained with the same computational budget and evaluated at the best validation checkpoint.

In Tab. 3, we assess both the pure synthesis and the modification ability of ImageBART by computing FID scores on samples and modified images (in the case of upper half completion as in Fig. 5). For both tasks, we use a single pass through the reverse Markov chain. We observe that the modification performance increases monotonically with the number of scales, which highlights the improved image manipulation abilities of our approach. For unconditional generation, we observe a similar trend, although FID seems to plateau beyond T=5T=5.

Joint vs. Independent Training While it is possible to optimize Eq. (2) jointly across all scales, we found that training is more robust when training all scales independently. Besides the usual separation of training the compression model pθ0p^{0}_{\theta} and the generative model pθt≥1p^{t\geq 1}_{\theta}, training the latter in parallel over multiple scales avoids the tedious weighting of the loss contribution from different scales; an often encountered problem in other denoising diffusion probabilistic models .

Efficiency with Less Decoder Layers As we implement the conditional transition probabilities pθt−1p^{t-1}_{\theta} with an encoder-decoder transformer architecture, we are interested in the effect of altering the ratio of encoder and decoder layers in the model. Recent work has provided evidence that it is possible to significantly reduce the number of decoder layers and thus also decrease autoregressive decoding speed while maintaining high quality .

We perform an experiment on LSUN-Churches, where we analyze the effect of different layer-ratios on synthesis quality (measured by FID) and on decoding speed when fixing the total number of model parameters to 200M. The results in the left part of Fig. 8 confirms that it is indeed possible to reduce the number of decoder layers while maintaining satisfactory FID scores with higher decoding efficiency. We identity a favorable trade-off between four and six decoder layers and transfer this setting to our other experiments.

Finally, we compare our model in terms of sampling speed with the recent state-of-the-art generative diffusion and AR models . The results are summarized in Fig. 8. While consistently being faster than all pixel-based models due to training in a compressed latent space, the increase in runtime w.r.t. is moderate due to the use of encoder-decoder transformers, i.e., a a decrease in pure decoder layers. If a faster runtime is desired, the speed can be further increased by reducing the number of decoder layers even more, see also the discussion in Sec. A.5.

Conclusion

We have proposed ImageBART, a hierarchical approach to introduce bidirectional context into autoregressive transformer models for high-fidelity controllable image synthesis. We invert a multinomial diffusion process by training a Markov chain to gradually incorporate context in a coarse-to-fine manner. Our study shows that this approach (i) introduces a natural hierarchical representation of images, with consecutive levels carrying more information than previous ones. (see also Fig. 9). (ii) It alleviates the unnatural unidirectional ordering of pure autoregressive models for image representation through global context from previous levels of the hierarchy. (iii) It enables global and local manipulation of a given input, a feat previously out-of-reach for ARMs. (iv) We additionally show that our model can be efficiently conditioned on various representations, allowing for a large class of conditional image synthesis tasks such as semantically guided generation or text-to-image synthesis.

Appendix A Appendix

We follow and implement our image compression models as “VQGANs”. More specifically, we use the official implementation provided at https://github.com/CompVis/taming-transformers and fine-tune the publicly available model for experiments on LSUN. For FFHQ, we train such a compression model from scratch. See Tab. 4 for an overview. As some of the codebook entries remain unused after training, we shrink the codebook to its effective size when training a generative model on top of it. For eventual entries not detected during evaluation on the subset, we assign a random entry.

A.1.2 Hierarchical Representations via Multinomial Diffusion

Tab. 5 lists the configurations of the multinomial diffusion processes for each experiment described in this work (see also Tab. 4). Note that all representations xtx_{t} for T>1T>1 have the same spatial resolution, but since each forward diffusion process gradually removes information, we obtain a coarse-to-fine hierarchy. On average, level xtx_{t} will contain ⌊αˉt⋅N⌋\lfloor\bar{\alpha}_{t}\cdot N\rfloor valid entries, which we denote as the effective sequence length in Tab. 5. Thus, ImageBART can also be interpreted as a generative compression model as illustrated in Fig. 9: By trading perfect reconstruction quality for compression, one can obtain a significantly shorter sequence, still representing a visually plausible image. This provides the basis for learning a generative model that does not waste capacity on redundancies in the data and the compressed space significantly lowers the computational demands for training and decoding.

A.1.3 Reverse Diffusion with Transformer Models

ImageBART is a learned Markov chain, trained to reverse the multinomial diffusion process described in Eq. (5). We can efficiently model the conditionals pθtp^{t}_{\theta} with a sequence-to-sequence model and follow to implement pθtp^{t}_{\theta} with an encoder-decoder architecture. Tab. 6 summarizes the hyperparameters used to implement the conditionals for each experiment. For comparison, the models in Tab. 1 contain 115M (VDVAE), 255M (DDPM), 30M (StyleGAN2), 158M (BigGAN), 448M (DCT) and 600M (TT) parameters.

A.1.4 Hardware

All models listed in Tab. 4 and Tab. 6 were optimized on a single NVIDIA A100 GPU and using 32-bit precision. Sampling speed as reported in Fig. 8 was also measured on a NVIDIA A100.

A.2 Details on Conditional Experiments

Semantically Guided Synthesis In addition to class- and text-conditional generative modeling, we apply our model to semantically guided synthesis of landscape images . To achieve this we follow and use the discrete representation of an autoencoder model trained on segmentation masks as conditioning cc for our models pθtp^{t}_{\theta}. However, since simply prepending cc here doubles the total length, which means a fourfold increase in complexity in the attention mechanism, we exploit the fact that the segmentation masks and the images (or their representations) are aligned. More specifically, within the encoder-decoder architecture, we first produce two embeddings e1e_{1} and e2e_{2} for xtx_{t} and cc, respectively, which are subsequently concatenated channel-wise, thereby keeping the sequence length of xtx_{t}. With this modifications, we train a model with T=5T=5 and individually optimize each scale similar to the unconditional training setting. Here again, we use the compression model pθ0p^{0}_{\theta} pre-trained on ImageNet. For training, we randomly crop the images and semantic maps to size 256×256256\times 256. For testing, however, we again use the sliding window approach of (cf. Sec. 4.2), which enables us to generate high-resolution images of landscapes, as visualized in Fig. 12 and Fig. 13.

A.3 Masked Diffusion Processes for Local Editing

Previous autoregressive approaches model images directly as a sequence of pixels from the top-left to the bottom-right. Thus, when generating a pixel, only context from neighbors to the left and above can be taken into account. While more recent approaches, which use a CNN to learn a discrete representation that is subsequently modeled autoregressively , improve this situation because elements of the representation now correspond to image patches, Fig. 5 showed that these models still fail to generate completions of the upper half of an image which are consistent with a given lower half.

While our approach also models each transition autoregressively from the top-left to the bottom-right, each transition additionally has access to global context from the previous step. We aim to exploit this fact to obtain novel applications such as consistent completions of upper halfs and, more generally, completions with respect to an arbitrary mask. For any such mask, let mm denote the result of downsampling it to the size of x1x_{1} using nearest-neighbor-interpolation, such that mi=0m^{i}=0 gives the positions where context should be used, and mi=1m^{i}=1 gives the positions where new content should be generated. We then define the masked forward process,

which only diffuses masked entries, and the masked reverse process,

which only denoises masked entries. By definition, running this process forward and then backward again represents the identity on umasked entries such that the given context remains constant. We denote this forward-backward process that starts from a given xt−1x_{t-1} and produces a sample xt−1,mx_{t-1,m},

by pθt−1∣t−1,mp^{t-1|t-1,m}_{\theta} and use it to sample with spatial conditioning information. Since it always leaves the unmasked context intact, the reverse process denoises the unmasked entries to make them consistent with the given context.

Besides Fig. 5, 6 additional visualizations of this process can be found in Fig. 14. The top shows masked inputs (left), final results of upper completions obtained by (middle) and by pθ1∣1,mp^{1|1,m}_{\theta} (right). The bottom visualizes the trajectory of the masked process, showing the masked input (leftmost column), denoised samples from pθ2,mp^{2,m}_{\theta} (first column) and pθ1,mp^{1,m}_{\theta} (second column), and every other sample from the forward-backward model pθ1∣1,mp^{1|1,m}_{\theta}. It demonstrates how the model iteratively incorporates global context from the previous scale to converge to a globally consistent sample at the very right. A visualization of the process on the class conditional ImageNet model is shown in Fig. 17. Additional examples for conditional samples from this process, as in Fig. 7, can be found in Fig. 15 and Fig. 16.

A.4 Limitations and Societal Impacts

Training deep generative models consumes a significant amount of energy (see also Sec. A.1 regarding the used hardware; the ImageNet model for example was trained for 19 days). With regard to the environment, it is important that we reduce the energy consumption as much as possible. To take a step in this direction, we followed previous works and relied on a strongly compressed, learned representation of images. Because we can fine-tune the corresponding encoder and decoder models from pre-trained ones, the costs for this step are largely amortized and subsequent levels of our hierarchy benefit from a drastically reduced sequence length. Nonetheless, it should be noted that such a strong compression scheme for images does not result in perfect reconstructions. For applications which require very high fidelity, such a level of compression might be unsuitable due to artifacts in the reconstructed images. Additionally, the use of adversarial learning in this stage can potentiate biases of datasets by its mode-seeking behavior. Both of these issues can be lessened with larger sequence lengths at the cost of higher energy requirements.

The transformer architecture which is used to model the transitions in our hierarchy is generally considered to be less biased compared to convolutional architectures. However, it also cannot benefit from useful inductive biases and therefore requires a large amount of data and resources to learn all required relationships. In early experiments, we noticed that on small datasets, such as CIFAR-10 , the transformer models overfit before they reach good performance. Thus, in its current implementation our approach requires datasets of sufficient size. Future works should evaluate different architectures, regularizations or augmentations to enable its use on small datasets. On the other extreme, we find that with large datasets, the main bottleneck is the computational resources that can be spent on the training of the transformer models. On the largest datasets, Conceptual Captions and ImageNet, we find that performance still improves after two weeks of training. Thus, consistent with other works on scaling up generative models, we expect that performance of our model will keep increasing with the available resources.

To ensure comparability with other approaches, we use standard benchmark datasets for deep generative models, even if some of them are known to contain offensive content .

A.5 Sampling Speed

Here, we discuss the effects of varying the number of encoder vs. decoder layers in ImageBART on sampling speed as presented in Sec. 4.4. On each diffusion scale, the encoder layers only have to run once whereas the decoder layers have to run ndata_dimn_{\text{data\_dim}} times. This results in an approximate complexity of order nscalesC(nencoder_layers+ndata_dimndecoder_layers)n_{\text{scales}}C(n_{\text{encoder\_layers}}+n_{\text{data\_dim}}n_{\text{decoder\_layers}}), where CC is the complexity of a single transformer layer. The speedup from such an encoder-decoder transformer over a decoder only transformer with nencoder_layers+ndecoder_layersn_{\text{encoder\_layers}}+n_{\text{decoder\_layers}} layers is therefore

A.6 Additional Samples & Nearest Neighbors

We provide additional samples from our models in Fig. 18-25. Additionally, we also provide nearest neighbors (measured in VGG feature space) for our FFHQ and LSUN-Churches models in Fig. 26 and Fig. 27, respectively.

References