Autoregressive Image Generation using Residual Quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, Wook-Shin Han
Introduction
Vector quantization (VQ) becomes a fundamental technique for autoregerssive (AR) models to generate high-resolution images . Specifically, an image is represented as a sequence of discrete codes, after the feature map of the image is quantized by VQ and rearranged by an ordering such as raster scan . After the quantization, AR model is trained to sequentially predict the codes in the sequence. That is, AR models can generate high-resolution images without predicting whole pixels in an image.
We postulate that reducing the sequence length of codes is important for AR modeling of images. A short sequence of codes can significantly reduce the computational costs of an AR model, since an AR uses the codes in previous positions to predict the next code. However, previous studies have a limitation to reducing the sequence length of images in terms of the rate-distortion trade-off . Namely, VQ-VAE requires an exponentially increasing size of codebook to reduce the resolution of the quantized feature map, while conserving the quality of reconstructed images. However, a huge codebook leads to the increase of model parameters and the codebook collapse problem , which makes the training of VQ-VAE unstable.
In this study, we propose a Residual-Quantized VAE (RQ-VAE), which uses a residual quantization (RQ) to precisely approximate the feature map and reduce its spatial resolution. Instead of increasing the codebook size, RQ uses a fixed size of codebook to recursively quantize the feature map in a coarse-to-fine manner. After iterations of RQ, the feature map is represented as a stacked map of discrete codes. Since RQ can compose as many vectors as the codebook size to the power of , RQ-VAE can precisely approximate a feature map, while conserving the information of the encoded image without a huge codebook. Thanks to the precise approximation, RQ-VAE can further reduce the spatial resolution of the quantized feature map than previous studies . For example, our RQ-VAE can use 88 resolution of feature maps for AR modeling of 256256 images.
In addition, We propose RQ-Transformer to predict the codes extracted by RQ-VAE. For the input of RQ-Transformer, the quantized feature map in RQ-VAE is converted into a sequence of feature vectors. Then, RQ-Transformer predicts the next codes to estimate the feature vector at the next position. Thanks to the reduced resolution of feature maps by RQ-VAE, RQ-Transformer can significantly reduce the computational costs and easily learn the long-range interactions of inputs. We also propose two training techniques for RQ-Transformer, soft labeling and stochastic sampling for the codes of RQ-VAE. They further improve the performance of RQ-Transformer by resolving the exposure bias in the training of AR models. Consequently, as shown in Figure 1, our model can generate high-quality images.
Our main contributions are summarized as follows. 1) We propose RQ-VAE, which represents an image as a stacked map of discrete codes, while producing high-fidelity reconstructed images. 2) We propose RQ-Transformer to effectively predict the codes of RQ-VAE and its training techniques to resolve the exposure bias. 3) We show that our approach outperforms previous AR models and significantly improves the quality of generated images, computational costs, and sampling speed.
Related Work
AR models have shown promising results of image generation as well as text and audio generation. AR modeling of raw pixels is possible , but it is infeasible for high-resolution images due to the slow speed and low quality of generated images. Thus, previous studies incorporate VQ-VAE , which uses VQ to represent an image as discrete codes, and uses an AR model to predict the codes of VQ-VAE. VQ-GAN , improves the perceptual quality of reconstructed images using adversarial and perceptual loss . However, when the resolution of the feature map is further reduced, VQ-GAN cannot precisely approximate the feature map of an image due to the limited size of codebook.
Composite quantizations have been used in other applications to represent a vector as a composition of codes for the precise approximation under a given codebook size . For the nearest neighbor search, product quantization (PQ) approximates a vector as the sum of linearly independent vectors in the codebook. As a generalized version of PQ, additive quantization (AQ) uses the dependent vectors in the codebook, but finding the codes is an NP-hard task . Residual quantization (RQ, also known as stacked quantization) iteratively quantizes a vector and its residuals and represents the vector as a stack of codes, which has been used for neural network compression . For AR modeling of images, our RQ-VAE adopts RQ to discretize the feature map of an image. However, different from previous studies, RQ-VAE uses a single shared codebook for all quantization steps.
Methods
We propose the two-stage framework with RQ-VAE and RQ-Transformer for AR modeling of images (see Figure 2). RQ-VAE uses a codebook to represent an image as a stacked map of discrete codes. Then, our RQ-Transformer autoregressively predicts the next codes at the next spatial position. We also introduce how our RQ-Transformer resolves the exposure bias in the training of AR models.
In this section, we first introduce the formulation of VQ and VQVAE. Then, we propose RQ-VAE, which can precisely approximate a feature map without increasing the codebook size, and explain how RQ-VAE represents an image as a stacked map of discrete codes.
We remark that reducing the spatial resolution of , , is important for AR modeling, since the computational cost of an AR model increases with . However, since VQ-VAE conducts a lossy compression of images, there is a trade-off between reducing and conserving the information of . Specifically, VQ-VAE with the codebook size uses bits to represent an image as the codes. Note that the best achievable reconstruction error depends on the number of bits in terms of the rate-distortion theory . Thus, to further reduce to but preserve the reconstruction quality, VQ-VAE requires the codebook of size . However, VQ-VAE with a large codebook is inefficient due to the codebook collapse problem with unstable training.
1.2 Residual Quantization
Instead of increasing the codebook size, we adopt a residual quantization (RQ) to discretize a vector . Given a quantization depth , RQ represents as an ordered codes
where is the codebook of size , and is the code of at depth . Starting with -th residual , RQ recursively computes , which is the code of the residual , and the next residual as
for . In addition, we define as the partial sum of up to code embeddings, and is the quantized vector of .
The recursive quantization of RQ approximates the vector in a coarse-to-fine manner. Note that is the closest code embedding in the codebook to . Then, the remaining codes are subsequently chosen to reduce the quantization error at each depth. Hence, the partial sum up to , , provides a finer approximation as increases.
Although we can separately construct a codebook for each depth , a single shared codebook is used for every quantization depth. The shared codebook has two advantages for RQ to approximate a vector . First, using separate codebooks requires an extensive hyperparameter search to determine the codebook size at each depth, but the shared codebook only requires to determine the total codebook size . Second, the shared codebook makes all code embeddings available for every quantization depth. Thus, a code can be used at every depth to maximize its utility.
1.3 RQ-VAE
For brevity, the quantized feature map at depth is also denoted by . Finally, the decoder reconstructs the input image from as .
Our RQ-VAE can make AR models to effectively generate high-resolution images with low computational costs. For a fixed downsampling factor , RQ-VAE can produce more realistic reconstructions than VQ-VAE, since RQ-VAE can precisely approximate a feature map using a given codebook size. Note that the fidelity of reconstructed images is critical for the maximum quality of generated images. In addition, the precise approximation by RQ-VAE allows more increase of and decrease of than VQ-VAE, while preserving the reconstruction quality. Consequently, RQ-VAE enables an AR model to reduce its computational costs, increase the speed of image generation, and learn the long-range interactions of codes better.
To train the encoder and the decoder of RQ-VAE, we use the gradient descent with respect to the loss with a multiplicative factor , The reconstruction loss and the commitment loss are defined as
RQ-VAE is also trained with adversarial learning to improve the perceptual quality of reconstructed images. The patch-based adversarial loss and the perceptual loss are used together as described in the previous study . We include the details in the supplementary material.
2 Stage 2: RQ-Transformer
In this section, we propose RQ-Transformer in Figure 2 to autoregressively predict a code stack of RQ-VAE. After we formulate the AR modeling of codes extracted by RQ-VAE, we introduce how our RQ-Transformer efficiently learns the stacked map of discrete codes. We also propose the training techniques for RQ-Transformer to prevent the exposure bias in the training of AR models.
After RQ-VAE extracts a code map , the raster scan order rearranges the spatial indices of to a 2D array of codes where . That is, , which is a -th row of , contains codes as
Regarding as discrete latent variables of an image, AR models learn which is autoregressively factorized as
2.2 RQ-Transformer Architecture
A naïve approach can unfold into a sequence of length using the raster-scan order and feed it to the conventional transformer . However, it neither leverages the reduced length of by RQ-VAE and nor reduces the computational costs. Thus, we propose RQ-Transformer to efficiently learn the codes extracted by RQ-VAE with depth . As shown in Figure 2, RQ-Transformer consists of spatial transformer and depth transformer.
The spatial transformer is a stack of masked self-attention blocks to extract a context vector that summarizes the information in previous positions. For the input of the spatial transformer, we reuse the learned codebook of RQ-VAE with depth . Specifically, we define the input of the spatial transformer as
Given the context vector , the depth transformer autoregressively predicts codes at position . At position and depth , the input of the depth transformer is defined as the sum of the code embeddings of up to depth such that
RQ-Transformer is trained to minimize , which is the negative log-likelihood (NLL) loss:
Our RQ-Transformer can efficiently learn and predict the code maps of RQ-VAE, since RQ-Transformer has lower computational complexity than the naïve approach, which uses the unfolded 1D sequence of codes. When computing length of sequences, a transformer with layers has of computational complexity . On the other hand, let us consider a RQ-Transformer with total layers, where the number of layers in the spatial transformer and depth transformer is and , respectively. Then, the spatial transformer requires and the depth transformer requires , since the maximum sequence lengths for the spatial transformer and depth transformer are and , respectively. Hence, the computational complexity of RQ-Transformer is , which is much less than . In Section 4.3, we show that our RQ-Transformer has a faster speed of image generation than previous AR models.
2.3 Soft Labeling and Stochastic Sampling
The exposure bias is known to deteriorate the performance of an AR model due to the error accumulation from the discrepancy of predictions in training and inference. During an inference of RQ-Transformer, the prediction errors can also accumulate along with the depth , since finer estimation of the feature vector becomes harder as increases.
As approaches zero, is sharpened and converges to the one-hot distribution .
Based on the distance between code embeddings, soft labeling is used to improve the training of RQ-Transformer by explicit supervision on the geometric relationship between the codes in RQ-VAE. For a position and a depth , let be a feature vector of an image and be a residual vector at depth in Eq. 4. Then, the NLL loss in Eq. 14 uses the one-hot label as the supervision of . Instead of the one-hot labels, we use the softened distribution as the supervision.
Along with the soft labeling above, we propose stochastic sampling of the code map from RQ-VAE to reduce the discrepancy in training and inference. Instead of the deterministic code selection of RQ in Eq. 4, we select the code by sampling from . Note that our stochastic sampling is equivalent to the original code selection of SQ in the limit of . The stochastic sampling provides different compositions of codes for a given feature map of an image.
Experiments
In this section, we empirically validate our model for high-quality image generation. We evaluate our model on unconditional image generation benchmarks in Section 4.1 and conditional image generation in Section 4.2. The computational efficiency of RQ-Transformer is shown in Section 4.3. We also conduct an ablation study to understand the effectiveness of RQ-VAE in Section 4.4.
For a fair comparison, we adopt the model architecture of VQ-GAN . However, since RQ-VAEs convert 2562563 RGB images into 884 codes, we add an encoder and decode block to RQ-VAE and further decreases the resolution of the feature map by half. All RQ-Transformers have and except for the model of 1.4B parameters that has and . We include all details of implementation in the supplementary material.
The quality of unconditional image generation is evaluated on the LSUN- and FFHQ datasets. The codebook size is 2048 for FFHQ and 16384 for LSUN. For the FFHQ dataset, RQ-VAE is trained from scratch for 100 epochs. We also use early stopping for RQ-Transformer when the validation loss is minimized, since the small size of FFHQ leads to overfitting of AR models. For the LSUN datasets, we use a pretrained RQ-VAE on ImageNet and finetune the model for one epoch on each dataset. Considering the dataset size, we use RQ-Transformer of 612M parameters for LSUN- and 370M parameters for LSUN-church and FFHQ. For the evaluation measure, we use Frechet Inception Distance (FID) between 50K generated samples and all training samples. Following the previous studies , we also use top- and top- sampling to report the best performance.
Table 1 shows that our model outperforms the other AR models on unconditional image generation. For small-scale datasets such as LSUN-church and FFHQ, our model outperforms DCT and VQ-GAN with marginal improvements. However, for a larger scale of datasets such as LSUN-, our model significantly outperforms other AR models and diffusion-based models . We conjecture that the performance improvement comes from the shorter sequence length by RQ-VAE, since SQ-Transformer can easily learn the long-range interactions between codes in the short length of the sequence. In the first two rows of Figure 3, we show that RQ-Transformer can unconditionally generate high-quality images.
2 Conditional Image Generation
We use ImageNet and CC-3M for a class- and text-conditioned image generation, respectively. We train RQ-VAE with 16,384 on ImageNet training data for 10 epochs and reuse the trained RQ-VAE for CC-3M. For ImageNet, we also use RQ-VAE trained for 50 epochs to examine the effect of improved reconstruction quality on image generation quality of RQ-Transformer in Table 2. For conditioning, we append the embeddings of class and text conditions to the start of input for spatial transformer. The texts of CC-3M are represented as a sequence of at most 32 tokens using a byte pair encoding .
Table 2 shows that our model significantly outperforms previous models on ImageNet. Our RQ-Transformer of 480M parameters is competitive with the previous AR models including VQ-VAE2 , DCT , and VQ-GAN without rejection sampling, although our model has 3 less parameters than VQ-GAN. In addition, RQ-Transformer of 821M parameters outperforms the previous AR models without rejection sampling. Our stochastic sampling is also effective for performance improvement, while RQ-Transformer without it still outperforms other AR models. RQ-Transformer of 1.4B parameters achieves 11.56 of FID score without rejection sampling. When we increase the training epoch of RQ-VAE from 10 into 50 and improve the reconstruction quality, RQ-Transformer of 1.4B parameters further improves the performance and achieves 8.71 of FID. Moreover, when we further increase the number of parameters to 3.8B, RQ-Transformer achieves 7.55 of FID score without rejection sampling and is competitive with BigGAN . When ResNet-101 is used for rejection sampling with 5% and 12.5% of acceptance rates for 1.4B and 3.8B parameters, respectively, our model outperforms ADM and achieves the state-of-the-art score of FID. Figure 3 also shows that our model can generate high-quality images.
RQ-Transformer can also generate high-quality images based on various text conditions of CC-3M. RQ-Transformer shows significantly higher performance than VQ-GAN with a similar number of parameters. In addition, although RQ-Transformer has 23% of parameters, our model significantly outperforms ImageBART on both FID and CLIP score (with ViT-B/32 ). The results imply that RQ-Transformer can easily learn the relationship between a text and an image when the reduced sequence length is used for the image. Figure 3 shows that RQ-Transformer trained on CC-3M can generate high-quality images using various text conditions. In addition, the text conditions in Figure 1 are novel compositions of visual concepts, which are unseen in training.
3 Computational Efficiency of RQ-Transformer
In Figure 4, we evaluate the sampling speed of RQ-Transformer and make a comparison with VQ-GAN. Both the models have 1.4B parameters. The shape of the input code map for VQ-GAN and RQ-Transformer are set to be 16161 and 884, respectively. We use a single NVIDIA A100 GPU for each model to generate 5000 samples with 100, 200, and 500 of batch size. The reported speeds in Figure 4 do not include the decoding time of the stage 1 model to focus on the effect of RQ-Transformer architecture. The decoding time of VQ-GAN and RQ-VAE is about 0.008 sec/image.
For the batch size of 100 and 200, RQ-Transformer shows 4.1 and 5.6 speed-up compared with VQ-GAN. Moreover, thanks to the memory saving from the short sequence length of RQ-VAE, RQ-Transformer can increase the batch size up to 500, which is not allowed for VQ-GAN. Thus, RQ-Transformer can further accelerate the sampling speed, which is seconds per image, and be 7.3 faster than VQ-GAN with batch size 200. Thus, RQ-Transformer is more computationally efficient than previous AR models, while achieving state-of-the-art results on high-resolution image generation benchmarks.
4 Ablation Study on RQ-VAE
We conduct an ablation study to understand the effect of RQ with respect to the codebook size () and the shape of the code map (). We measure the rFID, which is FID between original images and reconstructed images, on ImageNet validation data. Table 4 shows that increasing the quantization depth is more effective to improve the reconstruction quality than increasing the codebook size . Here, we remark that RQ-VAE with is equivalent to VQ-GAN. For a fixed codebook size 16,384, the rFID significantly deteriorates as the spatial resolution is reduced from 1616 to 88. Even when the codebook size is increased to 131,072, the rFID cannot recover the rFID with 1616 feature maps, since the restoration of rFID requires the codebook of size 16,3844 in terms of the rate-distortion trade-off. Contrastively, note that the rFIDs are significantly improved when we increase the quantization depth with a codebook of fixed size 16,384. Thus, our RQ-VAE can further reduce the spatial resolution than VQ-GAN, while conserving the reconstruction quality. Although RQ-VAE with can further improve the reconstruction quality, we use RQ-VAE with 884 code map for AR modeling of images, considering the computational costs of RQ-Transformer. In addition, the longer training of RQ-VAE can further improve the reconstruction quality, but we train RQ-VAE for 10 epochs as the default due to its increased training time.
Figure 5 and 6 substantiate our claim that RQ-VAE conducts the coarse-to-fine estimation of feature maps. For example, Figure 5 shows the reconstructed images of a quantized feature map at depth in Eq. 4. When we only use the codes at , the reconstructed image is blurry and only contains coarse information of the original image. However, as increases and the information of remaining codes is sequentially added, the reconstructed image includes more clear and fine-grained details of the image. We visualize the distribution of the code usage at each depth over the norm of code embeddings in Figure 6. Since RQ conducts the coarse-to-fine approximation of a feature map, a smaller norm of code embeddings are used as increases. Moreover, the overlaps between the code usage distributions show that many codes are shared in different levels of depth . Thus, the shared codebook of RQ-VAE can maximize the utility of its codes.
Conclusion
Discrete representation of visual images is important for an AR model to generate high-resolution images. In this work, we have proposed RQ-VAE and RQ-Transformer for high-quality image generation. Under a fixed codebook size, RQ-VAE can precisely approximate a feature map of an image to represent the image as a short sequence of codes. Thus, RQ-Transformer effectively learns to predict the codes to generate high-quality images with low computational costs. Consequently, our approach outperforms the previous AR models on various image generation benchmarks such as LSUNs, FFHQ, ImageNet, and CC-3M.
Our study has three main limitations. First, our model does not outperform StyleGAN2 on unconditional image generation, especially with a small-scale dataset such as FFHQ, due to overfitting of AR models. Thus, regularizing AR models is worth exploration for high-resolution image generation on a small dataset. Second, our study does not enlarge the model and training data for text-to-image generation. As a previous study shows that a huge transformer can effectively learn the zero-shot text-to-image generation, increasing the number of parameters is an interesting future work. Third, AR models can only capture unidirectional contexts to generate images compared to other generative models. Thus, modeling of bidirectional contexts can further improve the quality of image generation and enable AR models to be used for image manipulation such as image inpainting and outpainting .
Although our study significantly reduces the computational costs for AR modeling of images, training of large-scale AR models is still expensive, consumes high amounts of electrical energy, and can leave a huge carbon footprint, as the scale of model and training dataset becomes large. Thus, efficient training of large-scale AR models is still worth exploration to avoid environmental pollution.
Acknowledgements
This work was supported by Institute of Information & communications Technology Planning & Evaluation(IITP) grant funded by the Korea government(MSIT) (No.2018-0-01398: Development of a Conversational, Self-tuning DBMS; No.2021-0-00537: Visual Common Sense).
References
Appendix A Implementation Details
For the architecture of RQ-VAE, we follow the architecture of VQ-GAN for a fair comparison. However, we add two residual blocks with 512 channels each followed by a down-/up-sampling block to extract feature maps of resolution 88.
A.2 Architecture of RQ-Transformer
The RQ-Transformer, which consists of the spatial transformer and the depth transformer, adopts a stack of self-attention blocks for each compartment. In Table 5, we include the detailed information of hyperparameters to implement our RQ-Transformers. All RQ-Transformers in Table 5 uses RQ-VAE with 884 shape of codes. For CC-3M, the length of text conditions is 32, and the last token in text conditions predicts the code at the first position of images. Thus, the total sequence length () of RQ-Transformer is 95.
A.3 Training Details
For ImageNet, RQ-VAE is trained for 10 epochs with batch size 128. We use the Adam optimizer with and , and learning rate is set 0.00004. The learning rate is linearly warmed up during the first 0.5 epoch. We do not use learning rate decay, weight decaying, nor dropout. For the adversarial and perceptual loss, we follow the experimental setting of VQ-GAN . In particular, the weight for the adversarial loss is set 0.75 and the weight for the perceptual loss is set 1.0. To increase the codebook usage of RQ-VAE, we use random restart of unused codes proposed in JukeBox . For LSUN-, we use the pretrained RQ-VAE on ImageNet and finetune it for one epoch with 0.000004 of learning rate. For FFHQ, we train RQ-VAE for 150 epochs of training data with 0.00004 of learning rate and five epochs of warm-up. For CC-3M, we use the pretrained RQ-VAE on ImageNet without finetuning.
All RQ-Transformers are trained using the AdamW optimizer with and . We use the cosine learning rate schedule with 0.0005 of the initial learning rate. The RQ-Transformer is trained for 90, 200, 300 epochs for LSUN-bedroom, -cat, and -church respectively. The weight decay is set 0.0001, and the batch size is 16 for FFHQ and 2048 for other datasets. In all experiments, the dropout rate of each self-attention block is set 0.1 except 0.3 for 3.8B parameters of RQ-Transformer. We use eight NVIDIA A100 GPUs to train RQ-Transformer of 1.4B parameters, and four GPUs to train RQ-Transformers of other sizes. The training time is 9 days for LSUN-cat, LSUN-bedroom, 4.5 days for ImageNet, and CC-3M, and 1 day for LSUN-church and FFHQ. We use the early stopping at 39 epoch for the FFHQ dataset, considering the overfitting of the RQ-Transformer due to the small scale of the dataset.
Appendix B Additional Results of Generated Images by RQ-Transformer
We show the additional examples of unconditional image generation by RQ-VAEs trained on LSUN- and FFHQ. Figure 7, 8, 9, and 10 show the results of LSUN-cat, LSUN-bedroom LSUN-church, and FFHQ, respectively. For the top- (top-) sampling, 512 (0.9), 8192 (0.85), 1400 (1.0), and 2048 (0.95) are used respectively.
B.2 Nearest Neighbor Search of Generated Images for FFHQ
For the training of FFHQ, we use early stopping for RQ-Transformer when the validation loss is minimized, since RQ-Transformer can memorize all training samples due to the small scale of FFHQ. Despite the use of early stopping, we further examine whether our model memorizes the training samples or generates new images. To visualize the nearest neighbors in the training images of FFHQ to generated images, we use a KD-tree , which is constructed by the VGG-16 features of training images. Figure 11 shows that our model does not memorize the training data, but generates new face images for unconditional sample generation of FFHQ.
B.3 Ablation Study on Soft Labeling and Stochastic Sampling
For 821M parameters of RQ-Transformer trained on ImageNet, RQ-Transformer achieves 14.06 of FID score when neither stochastic sampling nor soft labeling is used. When stochastic sampling is applied to the training of RQ-Transformer, 13.24 of FID score is achieved. When only soft labeling is used without stochastic sampling, RQ-Transformer achieves 14.87 of FID score, and the performance worsens. However, when both stochastic sampling and soft labeling are used together, RQ-Transformer achieves 13.11 of FID score, which is improved performance than baseline.
B.4 Additional Examples of Class-Conditioned Image Generation for ImageNet
We visualize the additional examples of class-conditional image generation by RQ-Transformer trained on ImageNet. Figure 12, 13, and 14 show the generated samples by RQ-Transformer with 1.4B parameters conditioned on a few selected classes. Those images are sampled with top- 512 and top- 0.95.In addition, Figure 16 shows the generated samples using the rejection sampling with ResNet-101 is applied with various acceptance rates. The images in Figure 16 are sampled with fixed for top- sampling and different acceptance rates of the rejection sampling and top- values. The (top-, acceptance rate)s are (512, 0,5), (1024, 0.25), and (2048, 0.05), and their corresponding FID scores are 7.08, 5.62, and 4.45. Figure 15 shows the generated samples of RQ-Transformer with 3.8B parameters using rejection sampling with (4098, 0.125).
B.5 Additional Examples of Text-Conditioned Image Generation for CC-3M
We visualize the additional examples of text-conditioned image generation by RQ-Transformer trained on CC-3M. Figure 17 shows the generated samples conditioned by various texts, which are unseen during training. Specifically, we manually choose four pairs of sentences, which share visual content with different contexts and styles, to validate the compositional generalization of our model. All images are sampled with top- 1024 and top- 0.9. Additionally, Figure 18 shows the samples conditioned on randomly chosen texts from the validation set of CC-3M. For the images in Figure 18, we use the re-ranking with CLIP similarity score as in and and select the image with the highest CLIP score among 16 generated images.
B.6 The Effects of Top-k𝑘k & Top-p𝑝p Sampling on FID Scores
In this section, we show the FID scores of the RQ-Transformer trained on ImageNet according to the choice of and for top- and top- sampling, respectively. Figure 19 and 20 shows the FID scores of 821M and 1400M parameters of RQ-Transformer according to different s and s. Although we report the global minimum FID score, the minimum FID score at each is not significantly deviating from the global minimum. For instance, the minimum FID attained by RQ-Transformer with 1.4B parameters is 11.58 while the minimum for each is at most 11.87. When the rejection sampling of generated images is used to select high-quality images, Figure 21 shows that higher top- values are effective as the acceptance rate decreases, since various and high-quality samples can be generated with higher top- values. Finally, for the CC-3M dataset, Figure 22 shows the FID scores and CLIP similarity scores according to different top- and top- values.
Appendix C Additional Results of Reconstruction Images by RQ-VAE
In this section, we further explain that RQ-VAE with depth conducts the coarse-to-fine approximation of a feature map. Table 6 shows the reconstruction error , the commitment loss , and the perceptual loss , when RQ-VAE uses the partial sum of up to code embeddings for the quantized feature map of an image. All three losses, which are the reconstruction and perceptual loss of a reconstructed image, and the commitment loss of the feature map, monotonically decrease as increases. The results imply that RQ-VAE can precisely approximate the feature map of an image, when RQ-VAE iteratively quantizes the feature map and its residuals. Figure 23 also shows that the reconstructed images contain more fine-grained information of the original images as increases. Thus, the experimental results validate that our RQ-VAE conducts the coarse-to-fine approximation, and RQ-Transformer can learn to generate the feature vector at the next position in a coarse-to-fine manner.
C.2 The Effects of Adversarial and Perceptual Losses on Training of RQ-VAE
In Figure 24, we visualize the reconstructed images by RQ-VAEs, which are trained without and with adversarial and perceptual losses. When the adversarial and perceptual losses are not used (the second and third columns), the reconstructed images are blurry, since the codebook is insufficient to include all information of local details in the original images. However, despite the blurriness, note that RQ-VAE with (the third column) much improves the quality of reconstructed images than VQ-VAE (or RQ-VAE with , the second column).
Although the adversarial and perceptual losses are used to improve the quality of image reconstruction, RQ is still important to generate high-quality reconstructed images with low distortion. When the adversarial and perceptual losses are used in the training of RQ-VAEs (the fourth and fifth columns), the reconstructed images are much clear and include fine-grained details of the original images. However, the reconstructed images by VQ-VAE (or RQ-VAE with , the fourth column) include the unrealistic artifacts and the high distortion of the original images. Contrastively, when RQ with is used to encode the information of the original images, the reconstructed images by RQ-VAE (the fifth column) are significantly realistic and do not distort the visual information in the original images.
C.3 Using D𝐷D Non-Shared Codebooks of Size D/K𝐷𝐾D/K for RQ-VAE
As mentioned in Section 3.1.2, a single codebook of size is shared for every quantization depth instead of non-shared codebooks of size . When we replace the shared codebook of size 16,384 with four non-shared codebooks of size 4,096, rFID of RQ-VAE increases from 4.73 to 5.73, since the non-shared codebooks can approximate at most clusters only. In fact, a shared codebook with =4,096 has 5.94 of rFID, which is similar to 5.73 above. Thus, the shared codebook is more effective to increase the quality of image reconstruction with limited codebook size than non-shared codebooks.