CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP
Zihao Wang, Wei Liu, Qian He, Xinglong Wu, Zili Yi
Introduction
Text guided image generation in the general domain has been a challenging and frontier task in recent years. Early approaches (e.g., DMGAN , AttnGAN , DF-GAN , Obj-GAN , OPGAN , SD-GAN , CPGAN , XMC-GAN ) that directly generate pixels from the given text embeddings with a convolutional generator have shown promising revolutions to generate images in limited domains. However, when designated to generate images in the general domain, these methods see poor results in terms of image quality and text-image matching.
Recently, transformer-based text-to-image generators such as DALL-E and CogView have achieved great progress. Such progress is owed to two factors. First, the discretized representations of images achieved by vector quantized models such as VQ-VAE and VQ-GAN enable an image to be represented in the same way as natural language, thus enabling a transformer to be trained upon the cross-modality text-image data in a unified framework. Second, the progress in terms of large model (consisting of tens or hundreds of billions of parameters) training significantly leverages the model capacities of modeling cross-modality data in general domains. So far, these large-transformer-based methods achieve the best performance in terms of image quality, text-image relevance and range of domains. However, a limitation is that they require hundreds of millions of high-quality paired text-image data for the training, which are not publicly available and typically too expensive to acquire.
On the other hand, inspired by the recent progress in cross-modality language-vision pre-training and the unveiling of CLIP model , various optimization-based methods attempt to search in the image space based on a query text by optimizing the text-image matching score of a pre-trained CLIP model. The image search domain used by these methods could be the latent codes of a pre-trained GAN model (e.g. BigGAN , StyleGAN , SWAGAN ), the codebook of a VQGAN model , Diffusion Denoising Models , structured representations such as a set of strokes (ClipDraw ) or triangles , or a SIREN network that maps spatial coordinates to pixels. These methods relieve the demands of huge paired datasets and computing resources. However, the images generated by methods of this stream are either limited to a specific domain (e.g., faces) or suffer low-quality (e.g., unnatural, structure-distorted or physically meaningless).
An analysis of methods of the two streams motivates us to seek a balance between the two. With the language-vision priors learned by the CLIP model, we shall be able to train a text-to-image generator without the use of any paired data. Considering the joint language-vision embedding space of a pre-trained CLIP is shared by both modalities, the image embedding extracted with the image encoder of the CLIP upon an image is also a good representation of textual semantics. If we train a transformer that maps an image embedding to the image itself, then the inference pipeline with text inputs is automatically bridged: starts from a text, goes though the united embedding space, and finally generates an image.
Specifically, we first extract the cross-modality embedding of the image in the joint language-vision embedding space with a pre-trained CLIP model . Next, we convert the image into a sequence of discrete tokens in the VQGAN codebook space which can be trained with the unlabeled image dataset in hand. Finally, an autoregressive transformer that predicts the image tokens based on its joint language-vision embedding is trained. During inference, the transformer can generate coherent image tokens based on the text embedding extracted upon an input text with the text encoder of CLIP, and the generated image tokens can be further reconstructed into an image with the VQGAN decoder: see Figure 2. Such a scheme relies on the assumption that the distribution of images embeddings used for training is well aligned with that of text embeddings during test, which shall hold if we intentionally ensure the statistical coherency of the semantic distribution of the training image data and text data used for test.
We propose a scheme to train a reliable and general text-to-image generator without any paired text-image data, but with a set of unlabeled images and a pre-trained CLIP model as prior. Our approach provides a promising new direction for high-fidelity text-to-image generation with accessible resources.
Qualitative and quantitative evaluations verify that our method outperforms optimization-based text-to-image methods (e.g., VQGAN+CLIP , BigGAN-CLIP ) and CNN-based methods (DF-GAN ) in terms of image quality while not compromising the text-image matching. Our model can even achieve comparable performance as the flagship supervised model like CogView which is trained with huge amounts of paired data.
Related Work
Oord et al. first present an approach called Vector Quantized Variational Autoencoder (VQVAE) to learn discrete representations of images and model their distribution autoregressively with a convolutional architecture. extends this approach to use a hierarchy of learned representations to represent images of higher-resolution. introduces self-attention layers to the bottleneck of the convolutional architecture with the expect to capture long-range interactions in high-resolution images, and adversarial training losses to enforce the learning of perceptually rich codebooks. In our approach, we use the perceptually rich discrete representations of images that preserve more photorealistic details and natural textures.
Vision-Language Modelling Pre-training methods that have recently moved from raw text to multi-modal data (e.g., image-text) have revolutionized numerous multi-modal tasks (e.g., image-text matching , image captioning , Visual Question Answering ). Cross-modality tasks require the understanding of both modalities, and the alignment and relationships between the two modalities. The pre-training enables the encoder to produce representations with fused cross-modality information, thus benefiting downstream tasks.
The unveiling of CLIP model is a big step for multi-modal pre-training. In the CLIP architecture, the image modality and the language modality are mapped with the image encoder and the text encoder respectively to the shared multi-modal embedding space. The text-image similarity score computed in the shared multi-modal embedding space of CLIP can serve as a metric of text-image alignment or an objective for text-guided image generation . However, in our strategy, we use the shared multi-modal embedding space in a creative way, i.e., training a reverse model that maps the shared embedding to the image modality.
Text-to-Image Generation While there have been several attempts to improve the controllability of image generation by conditioning image synthesis on explainable priors (categories, attributes, label maps, edge maps, key points, depth-map), they often require users to follow some fixed control patterns. However, text-to-image generation that enables free-style user controls is a good choice, as natural language is easy to express and rich in information.
Recent years see great progresses in this field and many methods have been proposed. The key differences of existing text-to-image approaches rely on what are used to represent texts (e.g., word embeddings) and images (e.g., GAN, VAE, VQVAE or raw pixels) respectively, and what model is used to bridge the two modalities. Early methods for text-to-image attempt to train a convolutional generator that predicts pixels directly from the given text embeddings. Recently, transformer-based generators that map the textual embeddings to the discretized representations of images (VQGAN or VQVAE ) have achieved significantly better results than traditional CNN-based methods. Other streams rely on a pre-trained GAN model (e.g., StyleGAN ) and attempt manipulate the style space based on textual inputs , or rely on a pre-trained text-image matching model (e.g., CLIP ) and attempt to optimize the image representations to satisfy the textual guidance . Our method takes the advantages of both transformer-based and CLIP-based methods. We make use of the knowledge priors learnt with CLIP and train a powerful transformer without any paired data.
Approach
In this section, we introduce the details about our text-to-image generation framework and the training strategy.
As shown in Fig. 3, our model is made up of three components: a pre-trained language-image matching model (CLIP), an image tokenizer (VQ-GAN) and a conditional autoregressive transformer that takes the image embedding of an image extracted from CLIP as the condition , then generates the discrete image tokens of the same image.
Contrastive Language–Image Pre-training (CLIP) has achieved great success in mapping the language-image inputs to a common embedding space. Given an image or a sentence as the input (denoted as ), the CLIP model can embed them into a common representation space:
The official pre-trained CLIP model provided by is trained on 400-millions of text-image pairs with InfoNCE Loss:
which learns robust representations of hetero-modality data to ensure the semantically relevant data to be close to each other in the common embedding space.
Whereas it is too expensive to obtain large-scale and high-quality text-image pairs within our domain-of-interest, the CLIP model pre-trained on 400-million noisy pairs collected from Internet has shown enough capability to model language-vision data in general domains. Specifically, we use the ViT-B/32 variation of CLIP in our experiments.
2 Learning an Efficient Image Tokenizer
The recent VQ-VAE and VQ-GAN models have shown promising results to compress image patches into discrete image tokens. Such mechanisms enable images to be represented in the same way as natural language and easier to process with transformers.
where the indices sequence of is denoted as .
The VQ-decoder is used to reconstruct an image from the token sequence , i.e., .
The VQGAN model can be optimized with an objective consisting of the reconstruction loss:
3 Conditional Autoregressive Transformer
Once the complete set of tokens are restored with respect to the image embedding, the pre-trained decoder could reconstruct the tokens back to an image, .
4 Training Strategy
We employ the two-stage training strategy.
First Stage We first train a VQ-GAN model with the image dataset in a self-supervised manner. As mentioned in Sec. 3.2, all parameters of the encoder , decoder , codebook and discriminator will be optimized during training. The training objective is:
Second Stage The conditional autoregressive transformer is trained at this stage. Since we have paired input-output data (embeddingimage), our objective is a sum of the embedding reconstruction loss and a loss to maximize the likelihood of the corresponding image token.
The maximum-likelihood of the token sequence is enforce with
To ensure the generated image can be mapped back to its embedding with the CLIP image encoder, we employ the embedding reconstruction loss:
The training objective is the weighted combination of the two losses above:
where we set in our implementation.
Experiment
In this section, we describe how we evaluate our method and compare with previous approaches. We first introduce the datasets used for training and validation and the implementation details of our approach on these datasets. Then we make comprehensive comparisons between our method and previous text-to-image methods both quantitatively and qualitatively.
We train and evaluate our methods on two datasets: MS-COCO and ImageNet .
MS-COCO is a widely used dataset for language-vision benchmarks. It contains 80k images for training and 40k test set images. Each image has 5 short sentence descriptions, which is not used in our method but used by competing methods (e.g., DM-GAN , DF-GAN , and AttnGAN ). We use the 2014 split of MS-COCO dataset in our experimental setting. We use the complete set of images to train the vQGAN and the text-to-image generator, while we only use the textual descriptions form the validation split for the quantitative evaluation and visual demonstration.
ImageNet has long been used to evaluate conditional generation tasks. It contains more than 14 million images, and a little more than 21 thousand groups or classes. We use the complete set of images to train the VQGAN and our text-to-image generator. For evaluations, we construct the input textual descriptions either by fitting the template of “a photo of a [class name]” (class name is an ImageNet category) or manually composing a caption like “a photo of some [descriptive] objects with some [features]” (descriptive can be some constraints of the color, size or other properties of the object, and features could be some accessories of the object).
2 Implementation Details.
For both datasets, we train a VQGAN with and . We use GPT2 as the architecture of our conditional transformer. We trained a 24-layers GPT2-medium for MS-COCO and a 48-layers GPT2-XL for ImageNet. The details of params are shown in Tab. 1. The CLIP we used is the pre-trained ViT-B/32 model released by OpenAI .
3 Comparisons
We compare our results with four existing approaches that are representative methods of different research streams (e.g., CNN-based methods, optimization-based methods and transformer-based methods). These methods are:
DF-GAN , DM-GAN and AttnGAN represent the traditional approaches which use CNN generator to directly generate images with a textual condition. These methods provide pre-trained models on the MS-COCO dataset, so we can directly use those models for comparisons.
CogView is a flagship transformer-based method and serves as a good representation of fully-supervised and large transformer-based models , which are trained on tens of millions of high-quality text-image pairs. It achieves the best text-image relevance and FID metrics (see more details in ), but the results by CogView suffer lack of perceptual details as images are encoded and decoded with VQ-VAE . Its pre-trained model is a large model with 4-Billion parameters trained on 50 million text-image pairs in the general domain and should cover the distribution of both ImageNet and MS-COCO images very well.
VQGAN-CLIP represents for zero-shot opimization-based approach. It utilizes CLIP scores to guidance the optimization direction of latent codes of a pre-trained VQGAN model without any extra training. The CLIP model used in our comparisons is the pre-trained ViT-B/32 model released by OpenAI . The VQGAN models for the two datasets respectively are the same as used in our method.
BigGAN-CLIP Since DF-GAN cannot be trained on ImageNet (without text labels), we use the BigGAN-CLIP as a substitute when conducting visual and quantitative comparisons on ImageNet dataset. Here the BigGAN model is the one pre-trained on ImageNet dataset and provided by .
4 Quantitative Results
Evaluation Metrics To evaluate the quality of generated images, we use the stantard metrics as in : Inception Score (IS) , Fr´echet Inception Distance (FID) and CapS . IS calculates KL-divergence between conditional distribution and marginal distribution given an image classifier. FID computes the Fr´echet distance between the distribution of Inception features of synthetic images and real-world images. CapS measures the semantic similarities between the input text and the generated image. The quantitative results are computed over 30,000 images generated based on diverse validation captions and 30,000 ground-truth images related to the captions.
MS-COCO Tab. 2 shows the comparison between our method and previous methods on text-guided scene image synthesis. Our method achieves the best FID-0 and FID-1 due to the perceptually rich results generated by VQGAN and coherent image structures. The CapS score is lower than CogView by 4% but significantly better than other competing methods.
ImageNet. Tab. 3 shows the class-conditional generation results on ImageNet. We compared our method with similar methods based on discrete image tokens. Since the models require text inputs, we use the textual prompting (A photo of a {ImageNet label}.) as the input. Our method achieves the best FID metrics, implying that our method achieves better image quality and more coherent semantic distribution.
5 Qualitative Results
As shown in Fig. 4 and Fig. 5, compared to the competing methods, our method could generate high-fidelity images with more details. Generally speaking, the results of VQGAN-CLIP are non-realistic and suffer severe image distortion. CogView can generate good image structures but fails to produce realistic textures as they use the VQVAE to discretize images. The BigGAN-CLIP method sees more natural details than VQGAN-CLIP but suffer distorted image structures either. The DF-GAN can generate acceptable image structures with perceptually rich details but is prone to producing local regional artifacts. The visual evaluation matches the quantitative results well.
As shown in Fig. 4, our model can successfully capture the semantic concepts such as “reflection in the water” ( column), but fails to match the numeric concepts like “three plush bears” ( column). These concepts are captured by CogView model very well.
To examine the generalization ability of our method, we attempt to generate images under out-of-distribution language descriptions. As shown in Fig. 5, some of the descriptions (e.g. “a dog with a cigarrete”, “a lemon with hair and face”) do not even have a corresponding real-world image. Our model that is trained on the realistic images can surprisingly generate images well-aligned with these out-of-distribution texts. However, CogView that is trained upon amounts of image with textual labels fails to match those decorative words (e.g., “flying”, “with a cigarrete”, “with big beak”).
We also explore the generalization ability of our method in terms of stylized synthesis. We attempt to generate images under special style descriptions (e.g., “sketch”, “oil painting” or even “style of Edvard Munch”). As shown in Fig. 6, our model can successfully synthesize stylized pictures even without seeing many stylized training samples as no style augmentation is applied during training.
Conclusion
In this paper, we propose the CLIP-GEN strategy to train a reliable and general text-to-image generator without using any paired text-image data, but with the image-language priors of CLIP and a set of unlabeled image data. Such strategy enables us to make use of the available huge text-free image dataset (e.g., ImageNet) to train a text-to-image generator as powerful as flagship models like CogView that is trained with huge amounts of paired data. The proposed strategy should see greater breakthroughs in the future if we use larger numbers of unlabeled images that are available on the Internet.