Scaling Laws of Synthetic Images for Model Training ... for Now

Lijie Fan, Kaifeng Chen, Dilip Krishnan, Dina Katabi, Phillip Isola, Yonglong Tian

Introduction

The quality and quantity of data play a crucial role in training vision models. Historically, the emphasis has been on creating large, meticulously curated image datasets with categorical labels at the image level for training supervised models . Prominent examples include CIFAR and ImageNet . While creating these datasets is effective on a smaller scale, their expansion to hundreds of millions of samples presents significant challenges. These challenges include the intensive labor required for curation at scale, as well as the increasing potential for noise and quality issues as the datasets scale up.

Recently, there has been an increasing interest in training vision models using language supervision . This shift is exemplified by models like CLIP , which move beyond the fixed, predefined categories typical of datasets like ImageNet. Training these models requires extensive image-text pair datasets. Developments ranging from the creation of the Conceptual Captions dataset , which comprises millions of image-text pairs, to the LAION dataset , encompassing billions of pairs, are examples of this growing trend. However, this approach is not without its challenges. The massive scale of data sourcing, often through web scraping, introduces significant noise. Scalability issues also persist. Moreover, the immense size of these datasets presents practical difficulties in terms of storage and data transfer. For instance, LAION-2B requires tens of terabytes of disk space and could take days, if not weeks, to download.

Fortunately, recent breakthroughs in text-to-image models have introduced exciting new possibilities in the realm of synthetic data generation. These models, capable of producing high-quality images from textual descriptions, offer several significant advantages. Firstly, they allow precise control over image content through input texts, which could provide categorical labels or paired text supervision for free. Secondly, they are bandwidth-efficient, as only the model needs to be transferred, not the entire dataset. For instance, models like Stable Diffusion occupy merely 5 GB of disk space, which is 2000×\times more efficient compared to the massive LAION-2B dataset. Thirdly, they facilitate easier scalability with markedly reduced human labor for dataset curation. These benefits naturally lead to the question of whether it’s feasible to scale up vision datasets with synthetic images for training supervised models.

However, the use of synthetic images is also not without its drawbacks. When scaled to tens or hundreds of millions of images, these models may produce images of lower quality or with misaligned concepts, and might also struggle with maintaining diversity. In this paper, we tackle a pivotal question: How effective is the scaling of synthetic images, specifically generated for training supervised vision models? We examine the scaling behavior of synthetic images created by cutting-edge text-to-image models, comparing their efficacy to real images in two key scenarios: the training of supervised classifiers and the training of vision models with language supervision, such as CLIP. Additionally, we explore a range of factors that markedly impact the scaling efficiency of synthetic images. These include the choice of text-to-image model, the classifier-free guidance scale employed, and the nature of text prompts used for generating training images. A summarized comparison of the scaling ability between real and synthetic images is shown in Figure 1.

An empirical study on the scaling behavior of images synthesized by three major text-to-image models (Stable Diffusion , Imagen , and Muse ) shows that model performance exhibits power law scaling as a function of the number of synthetic images they are trained on. This trend holds until computation budget and model size become limiting factors .

We identify several factors that can significantly alter the scaling ability of synthetic data, including prompt design, classifier free guidance, and the choice of models.

In supervised settings, synthetic data does not scale as effectively as real data. However, there are exceptions where synthetic data demonstrates better scaling: (1) with classes that text-to-image models are particularly adept at generating, and (2) when the test data deviates significantly from the training data, e.g., out of distribution data.

In CLIP training, the disparity in scaling performance between synthetic and real data is less pronounced. Incorporating synthetic data with real data leads to enhanced zero-shot performance in most scenarios.

Related Work

Text to image models. Recent breakthroughs in text-to-image models, primarily driven by advances in diffusion models , have enabled the generation of high-quality, photo-realistic images using neural networks. Key examples of such models include Imagen , which performs diffusion in pixel space, and Stable Diffusion , which operates in the latent space of an autoencoder. DALL-E 3 also exemplifies this category. An alternative family of models, based on visual tokens, utilizes VQGAN and Transformers . Prominent examples within this category include Parti and Muse . Additionally, recent advancements have been exploring the scaling Generative Adversarial Networks (GANs) for text-to-image generation, as demonstrated in works such as .

Learning from synthetic data. Synthetic data has proven to be effective in improving performance across various domains . Synthetic images, in particular, have been extensively utilized in a range of different computer vision tasks, including object detection , semantic segmentation , autonomous driving , and robotics . More recently, there has been evidence that combining synthetic images generated by text-to-image models with real images can improve the performance on supervised learning tasks . Particularly, has fine-tuned the text-to-image model using the target dataset, e.g. ImageNet, while this paper studies the capabilities of off-the-shelf text-to-image models. Additionally, there are efforts developing methods for learning transferable representations from synthetic images .

Neural scaling laws. Scaling up model size, data amount, and training budget has unlocked new capabilities of deep models . Recent studies suggest the testing loss behaves as a power low with respect to each of these three resources when the other two are proper, in large language models (LLMs), machine translation , auto-regressive generative models , and transfer learning . Similar behavior is observed in multi-modal models . Chinchilla suggests scaling up data proportionally to model size, to obtain compute-optimal LLMs. propose to fit scaling laws by extrapolating training curves. Of particular interest, theoretically shows one can break the power law with respect to data size with an ideal data pruning strategy. In this paper, we focus on the scaling behavior of synthetic data for training models.

Preliminaries

We first study the scaling behavior of synthetic images generated with state-of-the-art text-to-image models under the ImageNet supervised training setting.

There are three primary factors influencing the generated images used for supervised training: (1) choice of text-to-image model, (2) the classifier-free guidance scale, and (3) the class-specific prompt used for the text input. We will now provide a detailed description of each of these factors:

Text-to-Image Models. We conducted the study on three state-of-the-art text-to-image models of different types:

Stable Diffusion , a model that drives the diffusion process in the latent space of a pre-trained autoencoder.

Imagen , a model that drives the diffusion process directly in the raw pixel space.

Muse , a visual token-based generation model trained with masked generative modeling, that performs discrete diffusion in the latent space of an autoencoder.

These models have distinct architectural designs, but are all capable of generating photo-realistic images. Since Imagen and Muse are not publicly available, we base our work on a version trained on internal data sources.

Guidance Scale. All modern text-to-image models primarily rely on the classifier-free guidance (CFG) technique to generate images based on textual input . Increasing the CFG scale typically improves the alignment between the generated images and the input text, resulting in higher-quality output images. However, this also tends to reduce the diversity of content in the generated images. Through empirical analysis, we determined that when generating images for training supervised classifiers, it is advisable to use a relatively lower CFG scale compared to the default value used in generation. This ensures that the generated images exhibit a higher degree of diversity, particularly when generating images from texts describing the same class. We conducted a detailed analysis and determined the optimal CFG scale ranges for different models: [1.5, 10.0] for Stable Diffusion, [1.0, 2.0] for Imagen, and [0.1, 1.0] for Muse.

Class-specific Prompts. To generate images for each class in ImageNet, we employed different techniques to create corresponding text prompts. This allows us to generate images conditioned on the specific ImageNet class via the prompts. Take the class ‘Tench’ as example, we can have prompts as:

Classnames: Directly use the ImageNet class name. (‘Tench’)

Classnames + Description: Combine class name with its WordNet description. (‘tench, freshwater dace-like game fish of Europe and western Asia …’)

Classnames + Hypernyms: Combine ImageNet class name with its Wordnet hypernyms. (‘Tench, Tinca tinca, cyprinid, cyprinid fish’)

Word2Sen: Use a pre-trained T5 model as used in to convert the ImageNet class name into a sentence. We generate 100 sentences for each class. (‘a tench with fish in the distance.’)

CLIP templates: Generate either 7 or 80 sentences with the text templates CLIP used for zero-shot classification task. (‘a photo of the large tench’)

IN-Captions: Combine the class name with captions from ImageNet(IN) training images. Captions are generated by BLIP2 . (‘Tench, a man holding a fish’)

2 Metrics: Recognizability and Diversity

The above factors give us a number of configurations to generate synthetic data. We now proceed to define metrics to analyze the resulting images, and then analyze the scaling behavior exhibited by the images generated under this configuration. The generated images should possess two crucial attributes: (1) Recognizabilty: Synthetic images should exhibit high precision, meaning they correctly represent the intended class, and high recall, implying that images for other classes should not mistakenly contain elements of this class. (2) Diversity: It is essential that the generated images are diverse from each other to improve generalization.

We define two measures to quantify the recognizability and diversity of images generated under a specific configuration. We generate 50 images for each ImageNet class, resulting in a synthetic test set comprising 50,000 images. Subsequently, we define the two metrics as follows:

Recognizabiliy: Use a pre-trained ImageNet classifier (a ViT-B with 86.2%\% accuracy from ) to classify the generated images and compute the F1 score for each class. The final metric is given by averaging F1 score across all classes.

DiversityWe also tried replacing the diversity metric with FID or LPIPS , please refer to Appendix D for details.: Following , we extract features from the same pre-trained model and compute the standard deviation on the feature space for images from every class, and then compute the average score across all classes.

3 Scaling Law for Synthetic Data

Prior works on scaling laws, such as , have observed that, for sufficiently large models, the test loss LL and dataset scale DD, approximately follow a power-law relationship:

where kk is a constant. Thus LDL_{D} exhibits linear dependence on DD in log space. Let DID_{I} be 1.3 million, roughly the size of the ImageNet training set with real data. We re-write Equation 1 as:

The slope −k-k and y-intercept −log⁡LDI-\log L_{D_{I}} would determine a unique scaling curve in log space. With this, we provide quantitative definitions for two key metrics for scaling:

Scaling Ability: Quantifies the scaling effectiveness of synthetic images generated by a particular text-to-image configuration. Stepper curves means loss scales better with data, therefore we represent scaling ability by the negative of the slope: kk.

Performance at 1.3M: Measures the classification performance of models (as negative log loss) when trained on a dataset with a scale equivalent to 1.3M, the size of the ImageNet training set. It is represented by the y-intercept −log⁡LDI-\log L_{D_{I}}.

Scaling on Supervised Training

We train supervised classification models exclusively using the images generated by text-to-image models and evaluate their performance by computing cross-entropy loss and top-1 accuracy on the ImageNet validation set, which contains real images. Training iterations are scheduled linearly based on the training data size in logarithmic space. All generated images are resized to a resolution of 256x256 pixels. Unless stated otherwise, we employ the ViT-B model with a patch size of 16 as our backbone architecture. Training hyperparameters details are provided in Appendix A.

2 Performance at 1.3M

We commenced by generating synthetic ImageNet datasets, each containing 1.3 million synthetic images, using various configurations of text-to-image models, CFG scales, and prompts as outlined in Section 3. In total, we created synthetic ImageNets in 54 distinct configurations, with detailed information provided in Appendix C. Figure 2 displays the validation loss on the real ImageNet validation set, represented by −log⁡LDI-\log L_{D_{I}} as defined in Equation 2. A higher value correlates with increased classification accuracy, signaling better performance. Comparisons focusing on classification accuracy are also included in Appendix C.

Within this study, we investigate the impact of different prompt sets (Section 3) within the Stable Diffusion model. Squares (◼), Circles (⚫), and Diamonds (◆) represent prompt configurations involving IN-captions, Classnames, and all other prompt setups, respectively. For Muse and Imagen configurations, we maintain the prompt set as IN-Captions and vary the CFG scale within the ranges and [0.1, 1], respectively. Triangles (▲) represent the performance of images generated with Muse, while Stars (★) represent the performance of images generated with Imagen. Several key findings emerge from the results:

Diversity and Recognizability trade-off: Across different configurations, we observe a trade-off between diversity and recognizability. The top-right corner of the figure represents the best performance, indicating configurations that can generate both accurate and diverse images. Configurations perform poorly when either recognizability or diversity falls below a certain threshold. Also see Figure A4 bottom left for scattering colorized by accuracy.

Effect of Prompt Sets: Choosing different prompt sets can impact performance. Using a more diverse prompt set shifts the configuration towards the bottom-right of the figure. Transitioning from Classname to IN-captions for text-to-image prompts may contribute to this shift, likely due to the increased diversity on the text side, which inherently leads to more diverse generated images.

Impact of CFG Scale: When prompts are fixed, controlling the CFG scale also affects the performance of the classification model. Increasing the CFG scale shifts the configuration towards the upper-left part of the figure, where recognizability is increased, but diversity decreases. This initially leads to improved performance, followed by a decrease.

Text-to-Image Model Performance: In terms of text-to-image models, when all configured to use IN-Captions as prompts, Stable Diffusion, Imagen, and Muse demonstrate a quite similar trend in balancing recognizability and diversity. This similarity in their trade-off is reflected in their close proximity to each other in the plot.

3 Scaling Ability

We next proceed to analyze the scaling behavior of different models, as well as the difference between training supervised models on synthetic images and on the real ImageNet training set. Figure 3 illustrates the scaling behavior across various configurations. Specifically, for Stable Diffusion, we depict the scaling behavior of different configurations with various prompts and CFG scales. We select the optimal configuration for Muse and Imagen from Section 4.2, using IN-Caption as prompts and the corresponding optimal CFG scale for each model. From the figure, several observations can be made:

Power-law Relationship: Training on synthetic images follows a power-law relationship from 0.125 million to 4 million training images. Validation loss and training data size exhibit a linear correlation when analyzed in log space.

Scaling Disparity: Training on synthetic images does not scale as effectively as training on the real ImageNet training set images, and typically has a smaller scaling ability. This difference can be attributed to the curation of ImageNet training images and performing validation under an in-domain setting.

Impact of Prompts and CFG Scale: Using default prompt sets and CFG scale for image generation results in poor scaling ability, i.e. a very flat slope and smaller kk value. However, by tuning the prompts and CFG scale properly, the generated images become much more diverse, leading to an increased scaling behavior for synthetic images, bringing it closer to the scaling ability observed with real images. Nevertheless, the best scaling configuration is still significantly worse than scaling with real data.

4 Scaling beyond 4M

We naturally wonder about the scaling behavior when the dataset size exceeds 4 million images and whether the validation loss will continue to decrease. In Figure 3, we also illustrate the scaling curve for Stable Diffusion up to 64 million images, and for Muse and Imagen up to 8 million training images (in gray background). The results indicate that the relationship changes when the dataset scale exceeds around 4 million images.

We hypothesize that this could be due to the loss being constrained by insufficient model capacity. According to , the power-law relationship between validation loss and training dataset size requires the model to have sufficient capacity to fit the dataset and converge. Therefore, when the dataset size exceeds 4 million images, and if we continue to use ViT-B as the backbone architecture, the validation loss in log space no longer exhibits a linear trend. To address this, we retrain the supervised models with ViT-L as the backbone architecture for the best Stable Diffusion and Imagen configuration, as shown in the dotted lines. This improvement in model capacity could achieve a lower validation loss and maintain a roughly linear ratio up to the 8 million scale and slightly postpones the inflection point.

5 Out-Of-Distribution Scaling

We also investigate the scaling behavior on out-of-distribution (OOD) validation sets to determine whether it differs from the in-domain setup on the ImageNet validation set. We employ the supervised ImageNet classifiers and test them on four OOD validation sets, which include ImageNet-A , ImageNet-R , ImageNet-Sketch , and Imagenet-V2 . The scaling curves for validation loss and top-1 validation accuracy are presented in Figure 4.

Our empirical results indicate that in scenarios where the domain gap is relatively small, such as testing on ImageNet-v2, the scaling behavior mirrors the observation in in-domain setups, with real images showing superior scaling performance. However, a much more intriguing observation emerges when the domain shift is more pronounced, as seen in tests on ImageNet-R and ImageNet-Sketch. In these instances, the disparity in scaling capabilities between synthetic and real images narrows. Consequently, scaling up synthetic images becomes particularly beneficial and useful. Remarkably, in situations with sufficiently large dataset scales, synthetic images can even outperform real images from ImageNet training set (e.g. for ImageNet-R and ImageNet-Sketch with Muse), highlighting the potential of synthetic images in bridging significant domain gaps. Interestingly, when images are generated with 80 CLIP templates as text prompt (the light blue line in the plot) instead of IN-Captions, the improvements over real images on ImageNet-R and ImageNet-Sketch are more significant, although the scaling ability on the ImageNet validation set is worse (as shown in Figure 3). This suggests that carefully crafting text prompts can unlock further potential in increasing the efficacy of synthetic images, particularly for OOD scenarios.

Zoom-in: Per Class Analysis

In addition to our general analysis of scaling behavior and its impact on overall performance across all classes, we take a more detailed approach to gain a better understanding of the key factors affecting class-specific scaling behavior. Our analysis involves assessing the scaling ability of each specific class in the 1,000 categories in ImageNet. We aim to establish connections between the scaling ability of each class and the characteristics of the images generated by text-to-image models, aiming to figure out potential reasons why synthetic data does not scale as good as real images. For this analysis, we focus on images generated by Stable Diffusion, using the optimal CFG of 2.0 and IN-Caption prompts.

When we fix the generation configuration of specific text-to-image model, CFG scales and prompts as described in Section 3.1, text-to-image models could still exhibit varying degrees of recognizability and diversity when generating images for different object classes. To explore how these factors influence the scaling ability of each class, we conducted an analysis focusing on the correlation between scaling ability and both recognizability and diversity, all computed for each class individually. These correlations, and their implications for scaling efficiency, are depicted in Figure 5. Our analysis underscores the potential positive role of recognizability in determining the scaling ability for the synthetic images, for each specific class, generated by text-to-image models. We identified a positive correlation between recognizability and scaling ability, indicating that the precision in generating the intended class significantly enhances the scaling effectiveness of synthetic images. In contrast, the influence of diversity within each class seems to be more limited. Our findings reveal only a negligible correlation between diversity and scaling ability. This might be attributed to the increased noise introduced when computing diversity for specific categories, as opposed to the overall dataset.

2 What makes a ‘Poor’ class

To gain a deeper insight into how scaling ability is distributed across different classes, and identify classes that do not scale well, we created a scatter plot with scaling ability on the X-axis and 1.3M performance on the Y-axis, as shown in Figure 8. Each class is represented as a dot in this plot, with top-right positions indicating better performance when scaled up to 4 million images. Points are colored based on either diversity or recognizability, or their final performance at 4M dataset scale.

Based on their positioning in the scatter plot, classes can be categorized into three distinct groups. Points in the bottom-left section represent ‘Poor’ classes, which are marked by both limited scaling ability and poor overall performance. In contrast, classes located in the upper-right section are deemed ‘Easy’ characterized by strong initial performance as well as robust scaling ability. Lastly, classes situated in the mid-right section can be described as ‘Scaling’. These classes may exhibit poor initial performance but demonstrate considerable improvement as the dataset size increases.

In Figure 8, we showcase two classes from each of the ‘Scaling’, ‘Easy’, and ‘Poor’ categories to illustrate and compare their scaling behaviors against real images within the same class. This analysis highlights an intriguing finding: certain ‘Scaling’ classes demonstrate a scaling ability that surpasses that of real images, thereby emphasizing the potential utility of synthetic images in these scenarios. We present additional results for ‘Scaling’ classes in Appendix G.2.

Additionally, we present visualizations of the generated images from these categories in Figure 8. Our findings show that text-to-image models adeptly generate images for ‘Scaling’ and ‘Easy’ classes with commendable accuracy and diversity. However, these models face challenges in accurately rendering the correct concepts for ‘Poor’ classes.

Scaling on CLIP

We investigated the scaling behavior of synthetic data in CLIP training using the extensive LAION-400M dataset. The synthetic images were generated using Stable Diffusion. We compare across different CFG scales and choose the optimal one (1.51.5) for CLIP training. For evaluation, we followed the prompt templates from and conduct zero-shot classification on ImageNet and 15 different fine-grained classification datasets, including Food-101 , Stanford Cars , Oxford Pets etc. The training scale begins with 1 million image-text pairs, progressively scaling up to encompass the full dataset of 371371The LAION-400M dataset we used contains slightly less samples compared to the orignal one because of link rot. million samples. All models use ViT-B as the backbone with a patch size of 1616, and are trained for 3232 epochs across all dataset scales. Detailed training hyper-parameters are in Appendix B. Comparisons on different CFG scales are also available in Appendix H.

2 Scaling Analysis

We evaluated the scaling behavior across three different data setups: (1) using only synthetic images, (2) using only real images, and (3) using a combination of both synthetic and real images. Dataset scale here refers to the number of captions. When combining synthetic and real images for training, we maintained a consistent text scale throughout. During each training iteration, we randomly selected one image, either real or synthetic, for use. The comparative analysis of these setups, evaluated on zero-shot classification loss and accuracy on ImageNet validation set, is depicted in Figure 9.

The analysis revealed that for all three scenarios, zero-shot classification loss adheres to the power-law relationship when the data amount is under around 64 million, compared to the 4 million scale in supervised training. In this range, the loss and data scale maintain a linear relationship in logarithmic space. Additionally, while the scaling efficiency (reflected in the slope of the curve) of synthetic data is somewhat lower than that of real data, this discrepancy is less pronounced than in the supervised classifier settings. However, a noticeable performance gap persists between synthetic and real images, which is likely attributable to concept mismatches between generated images and corresponding texts in certain classes, as discussed in Section 5.2.

Moreover, our results indicate that combining synthetic and real images during CLIP training can significantly enhance zero-shot performance, particularly when the dataset is limited. For instance, in training scenarios with fewer than 10 million image-text pairs, integrating synthetic images with real data can boost performance by up to 5%5\%.

3 Scaling on downstream datasets

We followed the same setup and extended our comparison to include the scaling behavior of synthetic versus real images on 15 fine-grained classification datasets, detailed in Table 1. This analysis indicates a scaling behavior in these datasets that is consistent with our findings from the ImageNet evaluations. Notably, a combination of synthetic and real images demonstrated superior performance in most scenarios, particularly when the total dataset size was under 100 million samples. In cases with extremely limited data availability, such as with just 1 million samples, training on synthetic images occasionally yielded better performance than with real images, for some certain tasks, such as Pets and SUN397 .

Discussion

In this paper, we investigate the scaling laws of synthetic data in model training and identify three key factors that significantly influence scaling behavior: the choice of models, the classifier-free guidance scale, and the selection of prompts. After optimizing these elements and increasing the scale of training data, we find that, as expected, synthetic data still does not scale as effectively as real data, particularly for supervised classification on ImageNet. This limitation largely stems from the inability of standard text-to-image models to accurately generate certain concepts. However, our study also highlights several scenarios where synthetic data proves advantageous: (1) In certain classes, synthetic data demonstrates better scaling behavior compared to real data; (2) Synthetic data is particularly effective when real data is scarce, for instance, in CLIP training with limited datasets; (3) Models trained on synthetic data may exhibit superior generalization to out-of-distribution data. We hope our findings will pave the way for further research in this field.

Acknowledgements. The authors would like to thank Shobhita Sundaram, Julia Chae, Sara Beery, and the VisCam team for fruitful discussions, Yuanzhen Li for helping with computation resources, and Jason Baldridge and Sergey Ioffe for guidance and check on publication policy.

References

Appendix A Details on Supervised Training

Supervised training was conducted on both the real ImageNet training set and various synthetic ImageNet generated by text-to-iamge models at different dataset scales. To ensure fair comparisons across different setups, identical training hyper-parameters were used for both real and synthetic images. Our training approach aligns with the setup described in , utilizing binary cross-entropy loss. The number of total training iterations and warm-up iterations were adjusted in proportion to the dataset scale in logarithmic space. For instance, models at the 1 million scale were trained for 95k iterations with a 10k iteration warm-up period. At the 2 million scale, training was extended to 190k iterations with a 20k iteration warm-up, and for the 4 million scale, the training and warm-up periods were increased to 285k and 30k iterations, respectively. More detailed descriptions of the training hyper-parameters are provided in Table A1.

A.2 Details on Text Prompts

In this section, we provide more details on the different configurations of text prompts used for Class-specific Prompts, as outlined in Section 3.1. For the Classnames + Hypernyms configuration, we utilized all hypernyms associated with each specific ImageNet category, separated by commas. Regarding CLIP templates, we employed two sets of prompt templates with different number of sentences. The first one includes the 80 distinct sentence originally used in the CLIP paper and its inference codehttps://github.com/openai/CLIP/tree/main/notebooks. The second set includes a subset of 7 templates, as recommended in . Additionally, we incorporated two more prompt configurations for comparison, following the approach in : (1) Classnames + Description + Places, which combines ImageNet class names with their WordNet descriptions, followed by a background category sampled from the Places dataset . (2) Classnames + Hypernyms + Places, which is similar to the previous configuration but replaces the descriptions with WordNet hypernyms, also incorporating a background category from Places.

Together with the configurations described in Section 3.1 of the main paper, these methods result in a total of 88 different configurations for text prompts when generating images for ImageNet categories. Additional visualizations of images produced by each these prompt configurations are included in Appendix F.

A.3 Evaluation on Downstream Datasets

In addition to evaluating the trained supervised classifiers directly, we also conducted linear probing on 15 different fine-grained classification datasets. Detailed descriptions of these datasets can be found in Appendix B.2. To perform linear probing on these datasets, we first removed the linear classification head from the classifier trained on ImageNet. Then, we extracted features from both the training and testing sets of each dataset, without applying any data augmentation. Subsequently, logistic regression was employed on these extracted features. The logistic regression layer was optimized using L-BFGS, with a maximum number of iterations equals 500500. For a detailed comparison of these results, please refer to Appendix E.

Appendix B Details on CLIP Training

Previous studies along with our empirical analysis, indicate the necessity of using different training parameters according to dataset scale. Specifically, for smaller-scale datasets, a larger learning rate and weight decay are recommended to mitigate overfitting. Conversely, for larger datasets, both the learning rate and weight decay should be reduced. Accordingly, we have followed two distinct sets of hyper-parameters within the CLIP training pipeline, one tailored for datasets with fewer than 100 million captions following the parameter in , and another for those exceeding this threshold following the parameter in . The specific parameters for both configurations are outlined in Table A3. Models are trained for 32 epochs across all data scales. The number of warmup steps was set to 600 for the 1 million scale, 1200 for the 2 million scale, and 2000 for scales of 4 million or greater. It is important to note that we maintained consistent training hyper-parameters across all three different types of data sources (synthetic, real, synthetic+real) at the same data scale to ensure fair comparisons. For an in-depth comparison of the effects of different hyper-parameters in CLIP training at different scales, please refer to the experimental details provided in Appendix H.1.

B.2 Downstream Dataset

For all the pre-trained CLIP models, we conducted zero-shot evaluations on ImageNet and 15 other widely used downstream classification datasets. These datasets include Food-101 , Stanford Cars , SUN397 , Oxford Pets , among others. Detailed information about these evaluation datasets can be found in Table A2. It’s important to note that for zero-shot evaluations, only the test images from these datasets are used.

B.3 Zero-shot Evaluation Details

We employed the same text prompt templates as referenced in , following a similar text ensembling strategy. For each category, text features were computed for every single template, and the mean average of these features across all templates was used to represent the final text feature for that specific category. Given that CLIP training involves a trainable temperature parameter, τ\tau, it is necessary to incorporate this parameter during zero-shot evaluation to accurately compute the zero-shot classification loss. Let zimgz_{img} be the image feature from the visual encoder, ztxtz_{txt} denote the aggregated text feature. Assuming a total of CC classes, with ztxtcz_{txt_{c}} as the text feature for c−thc-th category, the zero-shot classification loss is calculated as follows:

Here sim(zimg,ztxtc)\text{sim}(z_{img},z_{txt_{c}}) calculates the dot product, measuring the similarity between image feature and text features for each category.

Appendix C Details for Performance at 1.3M

As detailed in Section 4.2 of the main paper, we adopted 54 different configurations encompassing distinct text-to-image models, CFG scales, and text prompts to generate 1.3 million synthetic images for each configuration. Following the generation process, a supervised model was trained on the images geenrated by each configuration. To facilitate a clearer comparison in our tables, we have categorized these 54 configurations into three distinct groups, with each group focusing on specific comparative factors:

The first group exclusively uses Stable Diffusion as the text-to-image model. The primary comparison focus here is on the impact of varying text prompt configurations.

The second group standardizes the text prompt configuration to IN-Caption. This group’s aim is to assess the effects of using different text-to-image models and to understand the behavior of CFG scales within each specific model, and to find the optimal CFG scale for each of them.

The third group also exclusively uses Stable Diffusion as the text-to-image model. Here, the comparison emphasis is on the impact of different CFG scales under different text prompt configurations.

By grouping the configurations into these three different groups, we aim to provide a more structured and comprehensible analysis. In each of the three groups, we present the detailed validation loss (the negative log loss here is used to plot Figure 2 in the main paper) and top-1 accuracy on ImageNet validation set for models trained with different configurations, all under the scale of 1.3 million images.

Table A4 presents the analysis for the first group. It shows that using IN-Caption as the text prompt yields the best performance across both CFG scales of 22 and 7.57.5. This superior performance is largely attributed to its ability to guide the text-to-image model to generate diverse images while maintaining high recognizability, thereby justifying our choice of IN-Caption for most of our experiments. In the second group, detailed in Table A5, we observe that different text-to-image models exhibit varying optimal CFG scales for image generation to train supervised models. Specifically, Stable Diffusion, Imagen, and Muse reach their optimal performance at CFG scales of 2, 1.5, and 0.3, respectively. These findings validate our decision to employ these specific CFG scales in our later study of scaling behavior for each model. Table A6 covers the third group’s comparisons, focusing on finding the optimal CFG scales for Stable Diffusion under different text prompts. The results indicate that for text prompts with less diversity, such as using Classnames or Classnames+Hypernym, a smaller optimal CFG scale (1.5) is better since it will lead to more diverse images during the generation process. In contrast, for more diverse text prompts, like IN-Captions and CLIP templates with 80 sentences, since there is more diversity on the text side relativly, a larger optimal CFG scale (2) is more effective.

In addition, we have included a recognizability versus diversity plot for each of the three comparison groups in Figure A2. Each point in these plots represents a specific configuration, and is color-coded based on either the top-1 classification accuracy or the negative log loss on ImageNet validation set. The figures illustrate a trade-off between diversity and recognizability. Optimal performance is typically observed when there is a relatively better and more balanced trade-off between these two factors. Configurations characterized by either low diversity or low recognizability tend to result in suboptimal performance, indicating the necessity of maintaining a balance between these two factors.

Appendix D Evaluation under FID and LPIPS

In addition to diversity, we computed two other key metrics: FID (Frechet Inception Distance) and LPIPS (Learned Perceptual Image Patch Similarity). Both of them are standard evaluation metrics for the text-to-image generation models. Our study examines the performance variations in relation to these two metrics. As detailed in Section 3.2 of the main paper, these metric scores are also calculated using the synthetic test sets, which comprises 50,00050,000 images for each configuration:

The FID scores are derived by measuring the Frechet Inception Distance between the synthetic test set, containing these 50,000 generated images, and the real ImageNet validation set.

For LPIPS, we perform the calculation on a per-class basis. We randomly select and compute the similarity between 250 pairs of synthetic images for each class, and the final LPIPS metric is computed as the average across all classes.

The comparison of FID and LPIPS scores across each group is presented in Tables A4, A5, and A6. Additionally, in Figures A2 and A4, we plot a detailed comparison of the performance across different image generation configuration groups, substituting diversity with either FID or LPIPS. Considering that lower scores for FID indicate better distribution match and for LPIPS implies larger intra-class diversity, we take the negative of these values for plotting purposes. This adjustment ensures consistency with the diversity plot on the X-axis, positioning better results towards the right.

Furthermore, we incorporate comparisons using diversity, FID, or LPIPS as the X-axis for all 54 text-to-image generation configurations in Figure A4. Our findings reveal that while there is a moderate correlation between the FID score or LPIPS of generated images and the classification performance of models trained on them, the relationship is not definitive. In some cases, configurations with the same level of recognizability but lower FID scores or LPIPS show inferior classification performance. This suggests that while FID and LPIPS are effective metrics for evaluating the quality of the images generated by text-to-image models, their correlation with the performance of supervised classifiers trained on synthetic images is not as strong as expected. This observation underscores the need for a more specific metric tailored to evaluate the performance of supervised classifiers trained on such synthetic images.

Appendix E Detailed Scaling Behavior Comparison

In Tables A7 and A8, we present a comparison of the scaling behavior of supervised models trained under various configurations. This comparison is based on linear probing performed on 15 fine-grained classification datasets, as detailed in Appendix A.3. Our findings indicate that, in general, the scaling behavior observed in linear probing on these downstream datasets aligns with the trends seen in the ImageNet validation set. However, there are instances where training on synthetic images surpasses the performance of training on real images, in the Food-101 dataset for example.

Additionally, we have also included the detailed comparison on the out-of-distribution (OOD) validation sets, including ImageNet-A , ImageNet-R , ImageNet-Sketch , and Imagenet-V2 . The results from these comparisons demonstrate that training on synthetic images can yield improved performance on OOD test sets, exemplified by the results on ImageNet-R.

Appendix F Visualization on Generated Images

To better understand the impact of various text prompts used in generating training images, we provide additional visualizations of images created using different text prompt configurations for specific ImageNet categories. These visualizations were generated using Stable Diffusion, with the CFG scale set to 2.

In Figure A6, we present a detailed visualization of the images generated with different text prompt configurations for three different ImageNet categories: Goldfish, Golden Retriever, and shopping carts. The visualizations illustrate that incorporating more detailed information into the prompt tends to encourage the text-to-image model to generate more diverse images. However, this increased diversity may potentially compromise the accuracy of the category of interest in the generated images.

Appendix G More per-class analysis

To delve deeper into how recognizability and diversity are distributed across the 1000 ImageNet classes and their influence on scaling ability (kk), we categorized all classes into 10 different groups based on their scaling ability, ranging from lowest to highest. For each group, we calculated the average recognizability and diversity of the classes within it. The result of this analysis is illustrated in a detailed bar plot in Figure A5.

The analysis reveals the following trend: as the scaling ability of a group increases, the average diversity initially rises and then begins to decrease. This trend suggests that at the initial stages, enhanced diversity contributes to the generation of more varied synthetic images, helping the supervised classifier to learn more robust features during training. However, beyond a certain point, further increases in diversity can be harmful, potentially compromising the accuracy of the generated images or leading to the omission of key objects or the generation of wrong concepts.

In contrast, the average recognizability consistently increases as scaling ability increases, indicating a stronger correlation between the scaling ability and recognizability for each class. This consistent improvement shows the significance of recognizability as a more relevant metric for class-based analysis.

G.2 More Results on ‘Scaling’ Classes

In Figure A8, we provide a detailed comparison of the scaling behavior for models trained on either real or synthetic images from Stable Diffusion, specifically focusing on the ‘scaling’ classes as described in Section 5.2. Additionally, Figure A8 presents visualizations of synthetic images generated for these classes, using the same setup as described in Section 5.

We further explore ’Scaling’ classes for supervised classifiers trained on images generated by the Imagen and Muse models. The scaling behaviors of these classes, in comparison to models trained with real images, along with their visualizations, are presented in Figure A10, Figure A10, Figure A12 and Figure A12. Our analysis reveals that certain classes, such as sweatshirts, exhibit consistently good scaling across different text-to-image models. Meanwhile, there are classes that show particularly strong scaling performance with specific text-to-image models.

For the ten ‘scaling’ classes selected in Stable Diffusion, we observed that models trained on synthetic images exhibit scaling abilities that are comparable to, and in some cases even superior to, those trained on real images. A notable example can be seen in the ‘bighorn sheep’ and ‘spotlight’ categories, where models trained on synthetic images already outperform those trained on real images at dataset scales below 1 million, and this advantage continues to grow as the scale increases, since there are only 1.3M real images.

This finding suggests that for certain concepts, text-to-image models are indeed capable of generating images that are more conducive to train supervised classifiers effectively. As text-to-image models continue to improve, we anticipate that such instances will become more frequent. Eventually, it’s plausible that models trained on synthetic images could surpass the performance of those trained on real images across the entire ImageNet validation set.

G.3 More ‘Poor’ Classes

In Table A9, we identify and list ‘poor’ classes where supervised models, trained on synthetic images, face challenges in accurate classification. For each of the three text-to-image models — Stable Diffusion, Imagen, and Muse — we highlight 40 distinct categories that pose difficulties. Notably, certain categories, such as tiger cats and vine snakes, are common challenges across different text-to-image models. Future research in the development of text-to-image models could benefit from focusing on these categories. Improving the accuracy in generating images of these ‘poor’ classes is crucial, as their current limitations are a key factor hindering the ability of synthetic images to have better scaling ability and performance than real images, in the supervised learning contexts.

G.4 Per-class FID and LPIPS

Following the same setup outlined in Section 5.1 of the main paper, we also computed the correlations between scaling ability (kk in Equation 2 of the main paper) and both FID and LPIPS scores. Unlike the previous analysis focusing on recognizability and diversity, this evaluation specifically studies the relationship of scaling ability with these two metrics.

To calculate the per-class LPIPS scores, we used the same method as detailed previously. However, for per-class FID computation, the existing synthetic test sets, containing only 50 images per class, were deemed insufficient, since FID score is sensitive to the number of images. Therefore we sample 13001300 images from the synthetic training images and compute the FID with images from real ImageNet training set for each class. Similar to our previous approach, we took the negative of the per-class FID and LPIPS scores for consistency, as lower scores indicate better performance. The correlations obtained are depicted in Figure A13.

The results from this figure indicate a lack of strong correlation between scaling ability and either FID or LPIPS scores. This finding highlights the necessity for a more tailored metric that is specifically designed to assess the scaling ability of supervised classifiers trained on synthetic images.

Appendix H More Results on CLIP Scaling

In Table A3, we detail the use of two distinct sets of hyper-parameters for CLIP training, tailored to different dataset scales. Config (a) in the table, labeled as ‘S’ here, is designed for smaller dataset scales with fewer than 100 million images. Conversely, config (b), represented as ‘L’ here, is intended for larger dataset scales equal to or exceeding 100 million images. To validate the necessity of these configurations, we present an empirical study in Table A10.

Here we train CLIP models on subsets of the LAION-400M dataset with 10M, 50M, and 100M samples, exclusively utilizing real images and applying the two different hyper-parameter sets. Our findings indicate that for scales of 10M and 50M, the ‘S’ hyper-parameter configuration yields superior results, with the performance difference being reduced at the 50M scale. In contrast, at the 100M scale, the ‘L’ configuration demonstrates enhanced performance. Therefore, based on these empirical results, we opted to utilize the ‘S’ hyper-parameter set for smaller data scales and the ‘L’ set for larger scales.

H.2 Comparison on different CFGs

To identify the most effective CFG scale for generating synthetic images to train CLIP models, we utilized the Stable Diffusion to create synthetic images for the CC12M dataset at four different CFG scales: 1.25, 1.5, 1.75, and 2.5. Following the generation of these images, CLIP models were trained using the synthetic images and their corresponding texts. The efficacy of these trained models was then evaluated through zero-shot classification on ImageNet and various downstream classification datasets.

The detailed comparison of these different CFG scales are presented in Table A11. Based on these results, we determined that a CFG scale of 1.5 delivers the best zero-shot classification performance on ImageNet. Consequently, we chose CFG=1.5=1.5 for the majority of our CLIP experiments.

H.3 Detailed experiment results for all scales

Table A12 provides detailed scaling behavior for CLIP models trained utilizing either synthetic, real, or a combination of synthetic and real images. We also present the scaling behavior comparison in detailed plots for each specific downstream dataset in Figure A15. Figure A14 shows the average scaling behavior over all 15 downstream datasets. The models were trained on subsets of the LAION-400M dataset, beginning with 1 million samples and scaling up exponentially to the entire set of 371M. Our findings indicate synthetic images does not scale as good as real onees, yet integrating synthetic images with real images in the training of CLIP models can be advantageous, particularly in scenarios where the dataset size is relatively limited.