Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
Yuchao Gu, Xintao Wang, Yixiao Ge, Ying Shan, Xiaohu Qie, Mike Zheng Shou
Introduction
In recent years, remarkable progress has been made in image synthesis using likelihood-based generative methods, such as diffusion models , autoregressive (AR) , and non-autoregressive (NAR) transformers. These models offer stable training and better diversity compared to Generative Adversarial Networks (GANs) . However, unlike GANs, which can generate high-resolution (e.g., 2562 and 5122) images at one forward pass, likelihood-based methods usually require multiple forward passes by sequential decoding or iterative refinement . Consequently, early works , which maximize likelihood on pixel space, are limited in their ability to synthesize high-resolution images due to the high computational cost and slow decoding speed.
Instead of directly modeling the underlying distribution in the pixel space, recent vector-quantized (VQ-based) generative models construct a discrete latent space for generative transformers. There are two basic components in VQ-based generative models, i.e., VQ tokenizers and generative transformers. VQ tokenizers learn to quantize images into discrete codes, and then decode the codes to recover the input images, which process is termed as reconstruction. Then, a generative transformer is trained to learn the underlying distribution in the discrete latent space constructed by the VQ tokenizer. Once trained, the generative transformer can be used to sample images from the underlying distribution, and this process is termed as generation. Thanks to the discrete latent space, VQ-based generative models can easily scale up to synthesize high-resolution images without prohibitive computation cost.
The VQ tokenizer has received much attention as the core component in VQ-based generative models. Various techniques, such as factorized codes and smaller compression ratio in VIT-VQGAN , recursive quantization in Residual Quantization , and multichannel quantization with spatial modulated decoder in MoVQ , have been used to compress more fine-grained details into VQ tokenizers, leading to steadily improving reconstruction fidelity. However, none of the previous works have carefully examined a fundamental question, how the improved reconstruction of VQ tokenizers affects the generation. Lacking such analysis is due to two main reasons: 1) the underlying assumption that “better reconstruction, better generation”, and 2) the absence of a visualization pipeline to intuitively compare generation results of various VQ tokenizers.
In this paper, we introduce a visualization pipeline for examining how different VQ tokenizers influence generative transformers. Unlike previous works that compare randomly-sampled generation results, our approach models specific images and facilitates a straightforward comparison of the generative transformer’s ability using different VQ tokenizers. The key idea is to reduce the flexibility of the sampling process by providing ground-truth context to generative transformers, which can be easily implemented with an autoregressive (AR) transformer with causal attention.
Our proposed visualization pipeline leads us to two important observations. 1) Improving the reconstruction fidelity of VQ tokenizers does not necessarily improve the generation. 2) Learning to compress semantic features within VQ tokenizers significantly improves generative transformers’ ability to capture textures and structures. As shown in Fig. 1, increasing the semantic ratio (=1) improves the AR transformer’s ability to capture texture and structure, while decreasing it (=0) results in the transformer modeling rough colors instead. These observations arise due to the competing objectives of reconstruction and generation optimization. Reconstruction aims to retain variation in the dataset by favoring latent spaces with larger variance (i.e., weaker separability), whereas generation optimization favors latent spaces with smaller variance (i.e., better separability) to optimize a classification objective.
Our observations reveal that there are two competing objectives for VQ tokenizers: semantic compression and details preservation, but recent VQ tokenizers have primarily focused on the latter. To balance the two objectives for better generation, we propose Semantic-Quantized GAN (SeQ-GAN), which consists of two learning phases. The first phase utilizes a semantic-enhanced perceptual loss to achieve semantic compression, while the second phase finetunes the decoder to restore fine-grained details while preserving structures and textures. Compared to previous VQ tokenizers, SeQ-GAN compresses semantic features rather than fine-grained details (e.g., high-frequency details, colors) into codebook and finetunes the decoder to restore those details, which does not affect transformer learning but improves local details generation.
Our main contributions are summarized as follows. (1) We rethink the common assumption ”better reconstruction, better generation” in recent VQ tokenizers, and propose a visualization pipeline to explore the impact of different VQ tokenizers on generative transformers. (2) We identify two competing objectives in optimizing VQ tokenizers: semantic compression and details preservation, and introduce SeQ-GAN as a solution that balances these objectives to achieve better generation quality. (3) Our SeQ-GAN achieves significant improvements over prior VQ tokenizers in both conditional and unconditional image generation, as demonstrated through experiments with both AR and NAR transformers. (Generation results are shown in Fig. 2).
Related Work
VQ-based Generative Models. The VQ-based generative model is first introduced by VQ-VAE , which constructs a discrete latent space by VQ tokenizers and learns the underlying latent distribution by prior models . VQGAN improves upon this by utilizing perceptual loss and adversarial learning in training VQ tokenizers, and using autoregressive transformers as the prior model, leading to significant improvements in generation quality. VQ-based generative models have been applied in various generation tasks, such as image generation , video generation , text-to-image generation , and face restoration .
Building on the success of VQGAN , recent works have focused on improving the two fundamental components of VQ-based generative models: VQ tokenizers and generative transformers. To enhance VQ tokenizers, VIT-VQGAN proposes quantizing image features into factorized and L2-normed codes with a larger codebook and small compression ratio, achieving finer reconstruction results. Residual Quantization recursively quantizes feature maps using a shared codebook to precisely approximate image features. MoVQ enhances the VQ tokenizer’s decoder with modulation and proposes multi-channel quantization with a shared codebook, resulting in state-of-the-art reconstruction results. Different from previous works, we argue that improving reconstruction fidelity does not necessarily lead to better generation quality.
Another line orthogonal to our work is improving generative transformers. Early works adopt autoregressive (AR) transformers . However, AR transformers suffer from low sampling speed and ignore bidirectional contexts. To overcome these limitations, non-autoregressive (NAR) transformers are introduced based on different theories, like mask image modeling (i.e., MaskGIT ) and discrete diffusion (i.e., VQ-diffusion ). In this paper, we demonstrate that integrating our SeQ-GAN as the VQ tokenizer consistently enhances the generation quality of both AR and NAR transformers.
Visual Tokenizers for Generative Pretraining. Recent works in large-scale generative visual pretraining also explore the potential of the visual tokenizer. Instead of directly performing mask image modeling on pixels , the pioneer BEiT reconstructs masked patches quantized by a discrete VAE . Follow-up works further strengthen the semantics of the visual tokenizer, such as PeCo , which adopts contrastive perceptual loss during tokenizer training, and mc-BEiT , which softens and re-weights the masked prediction target during visual pretraining. To further reduce the low-level representation in the visual tokenizer, iBOT abandons reconstructing pixels, but updates the tokenizer online during the pretraining. BEiT-v2 formulates the training objective of the visual tokenizer by reconstructing semantic features extracted by CLIP . Unlike prior attempts to remove low-level representation interference in visual pretraining, we highlight the importance of semantic compression and details preservation in training VQ tokenizers for image synthesis.
Methodology
In this section, we first review how VQ tokenizers affect generation in VQ-based generative models in Sec. 3.1. Then, we present a visualization pipeline in Sec. 3.2 to examine the impact of different VQ tokenizers on generative transformers. Based on this pipeline, we make two critical observations in Sec. 3.3, highlighting the competing objectives in designing VQ tokenizers. Finally, we propose SeQ-GAN in Sec. 3.4 as a solution that balances these objectives to improve generation quality.
In this section, we cover the fundamental process of VQ-based generative models and highlight the potential impact of VQ tokenizers on generation results.
The decoder is responsible for decoding the quantized features back to the image space, i.e., .
The training objective of the VQ tokenizer is to minimize the reconstruction error with respect to the input image. Following VQGAN to use adversarial loss () and perceptual loss () , the reconstruction objective can be formulated as
In Eq. 2, means stop-gradient and is known as the commitment loss , where the commitment weight is set to 0.25 following .
Training generative transformers. As shown in Fig. 3, the encoder and codebook of a trained VQ tokenizer define a discrete latent space that quantizes an image into a sequence of discrete indices for generative transformer training. This sequence serves as input and label in training the generative transformer with token classification loss. In this paper, we use the autoregressive (AR) transformer in VQGAN and the non-autoregressive (NAR) transformer in MaskGIT . Therefore, the quality of the discrete latent space defined by the encoder and codebook of VQ tokenizers will influence the generative transformer training.
Generation: sampling from generative transformers. After training a generative transformer, we can sample discrete index sequences from it through either autoregressive decoding or iterative refinement . To map the discrete indices back to visual details, we retrieve the corresponding feature from the codebook and decode it into image space using the VQ tokenizer’s decoder. Therefore, the decoder will affect the generation quality by influencing the index-to-visual-details mapping.
2 Pipeline for Visualizing VQ Generative Models
Recent VQ-based generative models examine their designs by looking into the random sampled generation results, where different sampling techniques are adopted (e.g., top- top- sampling , classifier-free guidance , or rejection sampling ). However, instead of examining random samples, we are more curious about how generative transformers model specific images, enabling us to check the influence of different VQ tokenizers on generative transformers side by side. To achieve that goal, we propose to reduce the flexibility of the sampling process by providing ground-truth (GT) contexts for predicting each index, which can be easily implemented by AR transformers.
The pipeline is shown in Fig. 4. First, we train a VQ tokenizer along with its corresponding AR transformer. To analyze a specific image, we obtain the GT index sequence from the VQ tokenizer and feed it to the trained AR transformer, similar to the teacher forcing strategy used in training AR transformers. Because the AR transformer adopts casual attention , it does not directly access the GT indices, but can accesses all GT context indices for predicting each index. Given the same context (i.e., preceding GT indices), the next index prediction task is well-controlled and thus we can get the top-1 predicted index sequence within one forward pass. Finally, we decode the GT sequence and the AR predicted sequence back to the image space by the decoder of the VQ tokenizer. Following this approach, we are able to visualize both the reconstruction of VQ tokenizers and the upper limit prediction of AR transformers for specific images.
3 Rethinking the Objectives of VQ Tokenizers
Motivation. Recent advancements in VQ tokenizers have led to improved reconstruction results, with MoVQ in particular enhancing their decoder with modulation to add variation to quantized code and achieve the highest reconstruction fidelity. However, few studies investigate whether improvements of reconstruction fidelity of VQ tokenizers benefit generation quality. To address this gap, we conduct the following experiment to answer this question.
Experimental Settings. In Sec. 3.1, we identify two key factors in VQ tokenizers that affect generation: 1) the quality of the discrete latent space defined by the encoder/codebook, and 2) the index-to-visual-details mapping defined by the decoder. Inspired by MoVQ , we keep the configuration of encoder/codebook the same, and enhance the decoder to strengthen the index-to-visual-detail mapping. Our baseline is a convolution-only VQGAN , and we add two extra convolution blocks or two interleaved regional and dilated attention blocks at each resolution level to enhance the decoder. Based on each tokenizer, we train the generative transformer with different configurations, including different parameter sizes (AR and AR-Large), different types (AR and NAR), and different training iterations (AR-Large and AR-Large-2). Additional experimental settings can be found in the Sec. 6.1.
Results. The results presented in Table. 1 show that enhancing the decoder improves reconstruction fidelity, but it does not necessarily lead to better generation quality. Surprisingly, the baseline tokenizer achieves the best generation quality. Assuming that the quality of the discrete latent space (defined by encoder/codebook) remains unchanged, enhancing the decoder should improve generation quality by improving the index-visual-details mapping. However, in reality, enhancing the decoder leads to a degradation in generation quality. This suggests that jointly learning the encoder/codebook with an enhanced decoder actually degrades the quality of the discrete latent space.
Using the proposed visualization pipeline in Sec. 3.2, we visualize the reconstruction and AR prediction results of the baseline tokenizer and its attention-enhanced variant in Fig. 5. Although the attention-enhanced variant leads to a more consistent reconstruction, the AR transformer faces challenges in capturing the details and can only predict a rough color for the main object, even given ground-truth contexts. This highlights the generative transformers’ difficulties in modeling the discrete latent space.
The discrepancy between reconstruction and generation is due to the conflicting optimization objectives. In reconstruction training, VQ tokenizers prefer a latent space with larger variance (i.e., weaker separability) to retain the variation of datasets, while generative transformer training prefers smaller variance (i.e., better separability), because it optimizes the classification objective (i.e., cross-entropy). Therefore, in Table. 1 and Fig. 5, a powerful decoder promotes encoding more variation in the codebook, which hinders the separability of the discrete latent space and thus results in suboptimal generation performance. Through the result, we arrive at the following observation.
Observation 1. Improving the reconstruction fidelity of VQ tokenizers does not necessarily improve the generation.
3.2 Details Preservation vs. Semantic Compression
Motivation. Observation 1 suggests that compressing more fine-grained details within the tokenizer in reconstruction does not always improve generation. Therefore, we shift our focus towards exploring the role of semantics in VQ tokenizers for better generation quality.
Semantic-Enhanced Perceptual Loss. Unlike generative pretraining that uses fully semantic tokenizers, image synthesis requires consideration of low-level details. To balance the trade-off between low-level details and semantics in VQ tokenizers, we introduce a semantic-enhanced perceptual loss that controls the details/semantic ratio.
Specifically, given an input and a reference image, we extract their activation features and from a pre-trained VGG network. For each layer , the feature is of shape . Then, the perceptual loss can be calculated as To preserve details, perceptual loss used in previous VQ tokenizers adopts the features from both the shallow and high layers, which we denote as in this paper. To better compress semantic information during reconstruction, we propose a semantic-enhanced perceptual loss , which removes the features from shallow layers and further includes the logit feature (i.e., feature before softmax classifier). The layers to extract features can be summarized as
: {relu-{1_2, 2_2, 3_3, 4_3, 5_3}},
: {relu5_3, logit}.
We re-weight the two perceptual losses to control the proportion between the details and semantic information by
Results. Using the proposed semantic-enhanced perceptual loss, we examine the impact of different semantic ratios () during VQ tokenizer training on generation quality. In particular, we find that increasing initially improves the reconstruction FID (rFID), with the best rFID achieved at =0.4 before decreasing. However, increasing consistently improves the generation FID. Our visualizations in Fig. 6(a) and Fig. 1(b) demonstrate that the semantic-enhanced VQ tokenizer (=1) enables the AR transformer to capture more overall structures and textures than the baseline tokenizer (=0). We provide additional visualizations in the Sec. 6.2. These results help us arrive at the following observation.
Observation 2. Semantic compression within VQ tokenizers benefits the generative transformer.
3.3 Discussion
Tokenizers in Natural Language Processing (NLP) are naturally discrete and semantically meaningful, and in large-scale generative visual pretraining , fully semantic visual tokenizers that abandon low-level information are preferred. However, VQ tokenizers in VQ-based generative models should consider low-level details. Previous works prioritize preserving details to achieve better reconstruction fidelity, but we find solely compressing fine-grained details within VQ tokenizers will degrade the discrete latent space and hinder transformer training. We argue that both semantic compression and detail preservation should be considered when designing VQ tokenizers for image synthesis.
4 Our Solution: SeQ-GAN
To achieve better generation quality, we propose the Semantic-Quantized GAN (SeQ-GAN) as the VQ tokenizer in VQ-based generative models, balancing the objectives of semantic compression and details preservation.
Fig. 7 illustrates the two-phase approach of SeQ-GAN for tokenizer learning. In the first phase, we prioritize semantic compression by applying the proposed semantic-enhanced perceptual loss in Eq. 3. However, semantic compression with VQ tokenizers may cause some loss of color fidelity and high-frequency details. To address this, we enhance the decoder in the second phase using interleaved block regional and dilated attention . We fix the encoder and codebook of the tokenizer and finetune the enhanced decoder with to achieve better detail preservation. Note that in the second phase of tokenizer learning, we fix the discrete latent space by fixing the encoder and codebook. Therefore, our decoder-only finetuning enhances the generation quality of local details without affecting the transformer learning of structures and textures.
Experiments
We train the SeQ-GAN on ImageNet , FFHQ and LSUN , separately. In the first phase, we train the SeQ-GAN on ImageNet and FFHQ using the Adam optimizer with a learning rate of 1e-4 for 500,000 iterations. For LSUN-{cat, bedroom, church}, we follow RQ-VAE to use the pretrained SeQ-GAN on ImageNet and finetune for one epoch on each dataset. In the second phase, we finetune the enhanced decoder of SeQ-GAN on three datasets for 200,000 iterations with a learning rate of 5e-5. Detailed settings are provided in the Sec. 6.1.
The results are summarized in Table. 2. Since VIT-VQGAN , RQ-VAE and MoVQ prioritize the reconstruction fidelity by compressing more fine-grained details within the tokenizer, they usually require a larger latent size. Our SeQ-GAN does not pursue the reconstruction fidelity, but optimizes for better generation quality. Therefore, SeQ-GAN does not achieve the best reconstruction fidelity. However, compared to VQGAN , with the same latent size and codebook size, SeQ-GAN still has a large improvement in rFID and codebook usage.
2 Unconditional Image Generation
We train AR and NAR transformers on top of SeQ-GAN for unconditional image generation on FFHQ and LSUN datasets. All models are trained for 500,000 iterations with the Adam optimizer, using a learning rate of 1e-4. Detailed hyperparameters are in the Sec. 6.1.
From the results in Table. 3, previous state-of-the-art results are achieved by continuous diffusion model ADM and StyleGAN2 , while VQ-based generative models typically lag behind. Using SeQ-GAN as the VQ tokenizer enables our AR/NAR transformers with 171M parameters to surpass VIT-VQGAN , RQ-VAE , and MoVQ , despite having fewer parameters. Our method achieves comparable performance to ADM and StyleGAN2 on both FFHQ and LSUN datasets.
3 Conditional Image Generation
We train AR and NAR transformers with our SeQ-GAN tokenizer on 256256 ImageNet generation. The model is trained with a learning rate of 1e-4 for 300 epochs to enable direct comparison with VIT-VQGAN and MaskGIT . Further training settings can be found in the Sec. 6.1.
Results are summarized in Table. 4. Our SeQ-GAN+AR (364M, 256 sample steps) achieves FID of 6.25 and IS of 140.9, a remarkable improvement over VIT-VQGAN (714M, 256 sample steps), which obtains 11.2 FID and 97.2 IS. Compared to MaskGIT , which obtains 6.18 FID, our SeQ-GAN+NAR achieves a better 4.99 FID with a similar sampling step and model size. Compared to MoVQ+NAR (389M, 12 sample steps), obtaining 7.22 FID and 130.1 IS, our SeQ-GAN+NAR-L (364M, 12 sample steps) achieves much better performance of 4.55 FID and 200.4 IS.
4 Ablation
Codebook regularization. We ablate the strategy for increasing codebook usage in the baseline setting (one-phase training with ). As shown in Table. 5, although the factorized and L2-normed codebook in VIT-VQGAN can largely enhance the reconstruction fidelity, its large codebook size results in a suboptimal performance on the AR transformer. Moreover, optimizing the NAR transformer on a large codebook size is unstable. Compared to the offline K-means clustering used in previous codebook learning , the entropy regularization used in our paper achieves a better reconstruction and generation performance.
Design of semantic-enhanced perceptual loss. Our baseline, , utilizes all five layers to compute perceptual loss. As shown in Table. 6, adding the logit feature improves both rFID and generation FID. Variant-D achieves the best rFID, while achieves the best generation FID. This demonstrates that reconstruction fidelity does not necessarily correlate with generation performance. Removing more shallow layers consistently improves generation quality, highlighting the importance of semantics when optimizing VQ tokenizers for generation quality. However, adopting the logit feature (variant-E) without the spatial feature results in significantly worse performance. While adjusting the balance between details and semantics by removing different perceptual layers is possible, it usually requires extensive parameter tuning to match the loss scale. Instead, we fix and and simply tune the semantic ratio in Eq. 3 to achieve our goal.
Influence of the second phase tokenizer learning. SeQ-GAN is trained with semantic-enhanced perceptual loss in the first phase, which can result in some loss of color fidelity and high-frequency details. However, by finetuning the enhanced decoder in the second phase, those details can be preserved for the generation. As shown in Fig. 8(a), the second phase learning can restore color distortion (e.g., windows). Furthermore, Fig. 8(b) shows that second phase learning consistently improves generation FID. It’s worth noting that joint learning the encoder/codebook and an enhanced decoder degrades the generation performance in our observation 1 (see Sec. 3.3.1). Therefore, the decoder-only finetuning is an effective way to promote details preservation without degrading discrete latent space.
Conclusion
This work examines a fundamental question in VQ-based generative models, “how the improved reconstruction of VQ tokenizers affects the generation”. To answer this question, we introduce a visualization pipeline to examine the influence of different tokenizers on AR transformers. Based on this pipeline, we find both semantic compression and details preservation should be considered in optimizing VQ tokenizers, in which previous works prioritize the latter. Based on this finding, we propose a simple solution SeQ-GAN, which achieves remarkable improvement over existing VQ-based generative models on image synthesis.
References
Appendix
In this section, we first present detailed experimental settings in Sec. 6.1. Next, in Sec. 6.2, we offer additional visualization and analysis to further our understanding of the observations. We then present more qualitative results of our method in Sec. 6.3. Finally, we discuss the limitations of our approach and potential directions for future work in Sec. 6.4.
Tokenizer learning. As shown in Table. 7, SeQ-GAN’s architecture is based on VQGAN . However, we modified the architecture in the first learning phase by removing the attention and constructing a convolution-only VQGAN. In the second learning phase, we enhanced the decoder with block regional and dilated attention (BD Attn) to make attention suitable for high-resolution feature maps. SeQ-GAN has a total of 54.5M and 57.9M parameters for the first and second learning phases, respectively. We use the style-based discriminator for training SeQ-GAN, as suggested in VIT-VQGAN . The hyperparameters used for training SeQ-GAN are summarized in Table. 9.
Generative transformer training. The autoregressive (AR) and non-autoregressive (NAR) transformers share the same architecture, except that the AR transformer adopts causal attention. As detailed in Table. 8, the AR/NAR transformers and their large variant have 172M and 305M parameters, respectively. We train both types of generative transformer using the hyperparameters listed in Table. 10. During sampling, we adopt the basic sampling techniques from VQGAN (i.e., top- sampling ) and MaskGIT (i.e., adjusting sample temperature), while excluding the classifier-free guidance and rejection sampling for simplicity.
Detailed settings for observation and ablation experiments. Our observation (see Sec. 3.3 in the manuscript) and ablation experiments (see Sec. 4.4 in the manuscript) are conducted on the ImageNet dataset. We mostly follow the same configurations as the benchmark experiments listed in Table. 9, with the exception that we use a batch size of 64 for SeQ-GAN learning. Based on each VQ tokenizer, we train the generative transformer on ImageNet with a batch size of 64 for 500,000 iterations, while keeping other settings the same as in Table. 10.
For Observation 1 (see Sec. 3.3.1), we evaluate different VQ tokenizers on various transformer configurations: 1) Different parameter sizes, including AR with 172M parameters and AR-Large with 305M parameters. 2) Different types of transformers, including both autoregressive and non-autoregressive transformers with 172M parameters. 3) Different training iterations, including AR-Large and AR-Large-2, where we add an extra 500,000 iterations to the AR-Large model to investigate whether longer training eliminates the difference in VQ tokenizer.
2 More Visualization and Analysis of the Observations
In this section, we present additional visualizations to support our observations and proposed solutions.
First, we train SeQ-GAN with varying semantic ratios and plot the validation loss curve for each corresponding generative transformer training in Fig. 10. Our results show that a larger semantic ratio results in lower validation loss, indicating that generative transformers are better able to model the discrete space constructed by VQ tokenizers when more semantics are incorporated.
Next, we employ our proposed visualization pipeline to examine the reconstruction and AR prediction using SeQ-GAN with two different semantic ratios (). Fig. 11 demonstrates that the generative transformer trained on SeQ-GAN (=1) is better able to model each instance (e.g., Row (a-c)), the semantic features (e.g., the cat’s face in Row (d) and the eagle’s beak in Row (e)), and the structure (e.g., the peaked roof in Row (f)). Note that compared to the SeQ-GAN (=0) in Fig. 11, the reconstruction of the SeQ-GAN (=1) loses some color fidelity and high-frequency details, leading to similar problems of lost details and spatial distortion in the generation results (see Fig. 9 (1st phase)).
Finally, to address the issue of lost details and spatial distortion resulting from removing shallow layers in during the first phase of tokenizer training, we use a two-phase tokenizer learning approach in our SeQ-GAN. In the second phase, we finetune an enhanced decoder to restore the lost details. To demonstrate the effectiveness of our two-phase tokenizer learning on generation quality, we decode the transformer-sampled indices to image space using the decoder from both SeQ-GAN (1st phase) and SeQ-GAN (2nd phase), and present the generation results in Fig. 9. Our visualization clearly shows that SeQ-GAN (2nd phase) preserves more details and enhances generation quality compared to SeQ-GAN (1st phase).
3 More Qualitative Results
We provide qualitative comparisons to BigGAN , VQGAN and MaskGIT in Fig. 12, Fig. 13 and Fig. 14. For MaskGIT and BigGAN, the samples are extracted from the paper and for VQGAN, we use their pre-generated samples in the official codebasehttps://github.com/CompVis/taming-transformers. Our SeQ-GAN+NAR produces results with better quality and diversity than previous methods. From the uncurated results in Fig. 15, Fig. 16, Fig. 17 and Fig. 18, our SeQ-GAN+NAR can generate images with high quality and diversity on unconditional image generation.
4 Limitation and Future Work
Our observation 1 indicates that the quality of the discrete latent space in VQ-based generative models cannot be directly assessed by reconstruction fidelity, as the reconstruction and generation have different optimization goals. Thus, future work could design more intuitive methods to evaluate the quality of the discrete latent space.
In addition, observation 2 highlights the importance of semantics in the discrete latent space for visual synthesis. We have kept our approach simple by controlling the semantic ratio through the modification of the perceptual loss. Future works on VQ tokenizers can explore more effective ways to balance semantic compression and details preservation. For example, contrastive learning may improve the semantics compression of VQ tokenizers.
4.2 Limitation
Our method has the limitation inherited from likelihood-based generative models. While techniques such as rejection sampling and classifier-free guidance can be used to filter out samples with bad shapes and improve sample quality in conditional image generation, there are few sampling techniques available for improving unconditional image generation. Classifier-based metrics such as FID tend to focus more on textures than overall shapes, and thus may not be consistent with human perception, as pointed out in . To address this, StyleGAN2 introduces the perceptual path length (PPL) metric , which is more related to the shape quality of samples. StyleGAN2 regularizes the GAN training to favor lower PPL. Although generative transformers with the SeQ-GAN tokenizer can achieve a better FID than StyleGAN2 in unconditional image generation, some samples still have poor overall shapes (as seen in the uncurated samples in Fig. 16). Therefore, an interesting area for future research is to investigate the design of sampling techniques for unconditional image generation in likelihood-based generative models to improve overall shape quality.