DataComp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander Ratner, Shuran Song, Hannaneh Hajishirzi, Ali Farhadi, Romain Beaumont, Sewoong Oh, Alex Dimakis, Jenia Jitsev, Yair Carmon, Vaishaal Shankar, Ludwig Schmidt
Introduction
Recent advances in multimodal learning such as CLIP , DALL-E , Stable Diffusion , Flamingo , and GPT-4 offer unprecedented generalization capabilities in zero-shot classification, image generation, and in-context learning. While these advances use different algorithmic techniques, e.g., contrastive learning, diffusion, or auto-regressive modeling, they all rest on a common foundation: large datasets containing paired image-text examples. For instance, CLIP’s training set contains 400 million image-text pairs, and Stable Diffusion was trained on the two billion examples from LAION-2B . This new generation of image-text datasets is 1,000 times larger than previous datasets such as ImageNet, which contains 1.2M images .
Despite the central role of image-text datasets, little is known about them. Many state-of-the-art datasets are proprietary, and even for public datasets such as LAION-2B , it is unclear how design choices such as the data source or filtering techniques affect the resulting models. While there are thousands of ablation studies for algorithmic design choices (loss function, model architecture, etc.), datasets are often treated as monolithic artifacts without detailed investigation. Moreover, datasets currently lack the benchmark-driven development process that has enabled a steady stream of improvements on the model side and isolates data enhancements from changes to the model. These issues impede further progress in multimodal learning, as evidenced by recent work showing that public datasets currently do not match the scaling behavior of proprietary alternatives .
In this paper, we take a step towards a more rigorous dataset development process. Our first and central contribution is DataComp, a new benchmark for multimodal dataset design. DataComp flips the traditional benchmarking paradigm in machine learning where the dataset is fixed and researchers propose new training algorithms. Instead, we hold the entire training code and computational budget constant so that participants innovate by proposing new training sets. To evaluate the quality of a training set, we score the resulting model with a testbed of 38 classification and retrieval tasks such as ImageNet , ImageNetV2 , DTD , EuroSAT , SUN-397 , and MSCOCO .
DataComp focuses on two key challenges that arise when assembling large training datasets: what data sources to train on, and how to filter a given data source. Each challenge corresponds to one track in our benchmark. To facilitate the filtering track, our second contribution is CommonPool, a dataset of 12.8B image-text pairs collected from Common Crawl and currently the largest public image-text dataset. We release CommonPool as an index of image url-text pairs under a CC-BY-4.0 license, and apply content checks in its construction to remove unsafe or unwanted content. In the filtering track, the goal of participants is to find the best subset of CommonPool to train on. In the second track, Bring Your Own Data (BYOD), participants may leverage any data source, as long as it does not overlap with our evaluation testbed.
Our third contribution is an investigation of scaling trends for dataset design. In particular, DataComp contains four scales, where we vary the training budget and the candidate pool size from 12.8M to 12.8B samples (see Table 2). Expressed in GPU hours, the cost of a single training run ranges from 4 to 40,000 GPU hours on the A100 cluster we used for development. The different scales enable researchers with different resources to participate in our benchmark. Moreover, our results show that the ranking of filtering approaches is largely consistent across scale.
Our fourth contribution is over three hundred baseline experiments, including techniques such as querying captions for relevant keywords, filtering based on image embeddings, and applying a threshold on CLIP scores. A key result from our baselines experiments is that smaller, more stringently filtered datasets can lead to models that generalize better than larger datasets coming from the same pool. At the 12.8B scale, our best filtering baseline increases ImageNet zero-shot accuracy by 6.9 percentage points (pp) relative to the unfiltered pool (see Table 3). For the BYOD track, our initial experiments show that 109M additional data points (less than 1% of the 12.8B pool) improve the CLIP-filtered subsets of CommonPool by up to 1.2 pp ImageNet accuracy (see Table 18).
Finally, our fifth contribution is DataComp-1B, a new state-of-the-art multimodal dataset. We obtain DataComp-1B by combining our two most promising filtering baselines. DataComp-1B enables training a CLIP ViT-L/14 model to an ImageNet zero-shot accuracy of 79.2% (see Table 1), corresponding to a computational cost reduction when compared to a larger CLIP ViT-g/14 model trained on LAION-2B for about longer. Moreover, our model outperforms OpenAI’s original CLIP ViT-L/14 by 3.7 percentage points, while using the same compute budget.
To make DataComp a shared environment for controlled dataset experiments, we publicly release our candidate pool url index, our tooling for assembling these pools, our filtering baselines, and our code for training and evaluating models at www.datacomp.ai. We believe that our infrastructure will help put research on dataset design on rigorous empirical foundations, draw attention to this understudied research area, and lead to the next generation of multimodal datasets.
Related Work
We review the most closely related work and include additional related work in Appendix C.
Classical work considers dataset cleaning and outlier removal to discard samples that may lead to undesirable model bias. A related line of work develops coreset selection algorithms , which aim to select data subsets that lead to the same performance as training on the entire dataset. These techniques appear to scale poorly to larger data regimes . More recent efforts in subset selection often operate on already curated datasets (e.g., CIFAR-10, ImageNet) or on smaller data regimes (e.g., YFCC-15M ). These settings often do not reflect newer training paradigms that involve (1) noisy image-text pairs instead of category labeled images and (2) large scale datasets (e.g., billions of samples). While data-centric investigations have led to community competitions like dcbench and DataPerf , existing benchmarks have likewise operated at small data scales compared to datasets like LAION-2B , which contains over two billion images. DataComp bridges this gap by aligning data-centric investigation with large scale image-text training.
There has also been renewed interest in dataset pruning and deduplication. Sorscher et al. show that data pruning can improve traditional scaling trends on ImageNet, but do not consider image-text training or larger datasets. Raffel et al. remove sentence redundancies when creating the C4 corpus. Subsequent work further demonstrated the benefits of deduplication for better language modeling . Radenovic et al. introduce CAT filtering for image-text datasets—a rule-based system to retain high quality samples. Abbas et al. propose SemDeDup, which starts with the CAT-filtered LAION-440M subset, further employing clustering to remove semantic duplicates. DataComp facilitates data-centric investigation at an even larger scale (i.e., 12.8B sample scale) and provides a common experimental setting for fair comparison amongst dataset creation algorithms.
Datasets have been instrumental to building multimodal models like CLIP , Flamingo , Stable Diffusion , DALL-E and GPT-4 . These methods succeeded by training on large, heterogeneous datasets rather than solely through advanced modelling techniques. For example, OpenAI’s CLIP trains on 400M image-text pairs from the web, roughly the size of ImageNet . Prior work on scaling image-text datasets also provides promising trends with respect to zero-shot model performance . Additional large scale datasets like FILIP-300M , FLD-900M , and PaLI-10B were constructed to train multimodal models. However, many datasets used to train such models (including the dataset for OpenAI’s CLIP) are proprietary, making it hard to conduct data-centric investigations.
Even for public image-text datasets like SBU , Flickr30k , MS-COCO , TaiSu , Conceptual Captions , CC12M , RedCaps , WIT , Shutterstock , YFCC-100M , COYO-700M , LAION-400M , or LAION-2B little is known about what constitutes a good image-text dataset. Preliminary analysis suggests that different image-text data sources lead to CLIP models with different properties . However, previous work is limited to smaller scale data (10-15M examples). Birhane et al. examine LAION-400M and find NSFW imagery and racial slurs, centering the dangers in web-scale multimodal datasets. To combat toxicity, we preprocess our pool to remove NSFW content and blur human faces detected in images. For more details on our safety preprocessing see Section 3.2, Appendices E and G.
The DataComp benchmark
DataComp is meant to facilitate data-centric experimentation. While traditional benchmarks emphasize model design, DataComp is centered around dataset development, where the resulting datasets can be used to train high accuracy models. We focus on large image-text datasets and quantify a dataset submission by training a CLIP model on it from scratch and evaluating on 38 downstream image classification and retrieval tasks. We additionally have three secret test sets, which will be released after a year, to guard against overfitting. To facilitate such investigations, we provide a candidate pool of uncurated image-text pairs sourced from the public internet. Our benchmark offers two tracks: one where participants must filter samples from the pools we provide, and another where participants can use external data. Moreover, DataComp is structured to accommodate participants with diverse levels of computational resources: each track is broken down into four scales with varying compute requirements. We now discuss high-level design decisions, construction of a 12.8B image-text data pool to facilitate the competition, benchmark tracks, model training, and evaluation.
In many areas of machine learning, larger datasets lead to better performing models . Hence comparing only datasets with the same size is a natural starting point. However, this approach is flawed as controlling the dataset size ignores critical curation constraints: candidate pool size (i.e., number of image-text pairs to harvest) and training compute. For instance, assembling a dataset like LAION-2B consists of identifying data sources (e.g., Common Crawl or Reddit) and filtering the data source. Notably, the final dataset size is a design choice and is only upper-bounded by the data sources. Hence, the true data constraint is the size of the reservoir of samples: candidate pool to be filtered. To make DataComp a realistic benchmark, we therefore fix the candidate pool in the filtering track, but give participants control over the training set size.
Compute cost is another relevant constraint. To put datasets of different size on equal footing, we specify the total number of training samples seen. Consider the 12.8B compute scale and filtered datasets and , with 6.4B and 3.2B image-text pairs respectively. At this scale, we train by making two passes over , while making four passes over . A key result from our experiments is that smaller, more stringently filtered datasets can lead to models that generalize better.
Two key procedures in assembling a training dataset are filtering a data source and aggregating data sources . To reflect this structure, DataComp has two tracks: filtering, where participants select a subset of the samples from CommonPool, and Bring Your Own Data (BYOD), where participants can use any source of data. Key decisions for each tracks are described in Sections 3.2 and 3.3, respectively. For full competition track rules see Appendix A.
To facilitate study of scaling trends and accommodate participants with various computational resources, we structure DataComp using four scales of compute: small, medium, large and xlarge. Each new scale increases the number of samples seen during training by 10 (from 12.8M to 12.8B samples seen), and the pool we provide by the same factor (from 12.8M samples to 12.8B samples). Table 2 gives the experimental configuration used for each scale. For the small scale, our runs took 4 hours on an A100 GPU, and for the xlarge scale 81 hours on 512 GPUs.
2 CommonPool generation, for the filtering track
We construct a large-scale pool of image-text pairs, CommonPool, from Common Crawl . CommonPool is distributed as an image url-text pair index under a CC-BY-4.0 license. Our pool construction pipeline has four steps: url extraction and data download, NSFW detection, evaluation set deduplication, and face blurring. We additionally provide per sample metadata (e.g., CLIP features). Starting from the xlarge CommonPool, we take successive random subsets to create large, medium, and small CommonPool (e.g., medium is a subset of large).
We first use cc2dataset , which utilizes Apache Spark , to extract pairs of image urls and nonempty alt-text from all Common Crawl snapshots from 2014 to 2022. We then deduplicate the url-text pairs and randomly shuffle. This step results in 88B possible samples. Not all samples are downloadable; other samples are not suitable due to NSFW content or overlap with our evaluation sets. We attempt to download 40B samples using img2dataset resulting in 16.8B image-text pairs. For more details, see Appendix D.
Since Common Crawl is a snapshot of the internet, we require strict preprocessing to remove unsafe content. We use Detoxify to prune samples that contain unsafe text (e.g., obscene, sexually explicit, or threatening language). We also discard samples with explicit visual content. To do so, we train a classifier on CLIP ViT-L/14 features, using the NSFW dataset used in LAION-5B . We validate our classifier against the Google commercial image safety API. See Appendix E for details. Around 19% of image-text pairs are considered NSFW, taking the pool of 16.8B downloads to 13.6B samples.
To prevent accidental overfitting to certain test sets in our evaluation suite, we perform a thorough near-duplicate removal between the candidate pool and our evaluation sets, using a state-of-the-art image deduplication model . Appendix F contains additional details. The model flags 3% of the 16.8B images as near-duplicates, reducing the 13.6B pool to 13.1B samples. From here we select a random subset to get the xlarge pool of 12.8B samples.
To protect the privacy of individuals, we detect and blur faces from images in our pool using a face detector . As observed by Yang et al. , obfuscating faces has little impact on model performance, as we also observe in our experiments (Appendix G).
To bootstrap participants we distribute metadata for each sample in CommonPool (e.g., image url, alt-text, original image resolution, CLIP features, and CLIP similarity scores). Following Carlini et al. , we release SHA256 hashes for each image to guard against data poisoning in subsequent CommonPool downloads. For additional details see Appendix H. We open-source our metadata processing pipeline as dataset2metadata .
3 The bring your own data (BYOD) track
While CommonPool can be used to study different filtering techniques, state-of-the-art models often train on data from different sources. For instance, the Flamingo model uses both multimodal massive web (M3W) and ALIGN datasets . To facilitate non-proprietary research on curating data from many sources, we instantiate a separate DataComp track to allow participants to combine multiple data streams. For example, participants could construct a training set from CC12M , YFCC100M , and data sources they label themselves. In Section 4.2 and Appendix P.2 we describe our exploration using existing public, image-text datasets. These datasets are acquired from their respective sources and are not re-release as part of DataComp.
4 Training
5 Evaluation
We evaluate on a suite of 38 image classification and retrieval tasks. We also study two additional fairness tasks, detailed in Section 5 and Appendix Q. As discussed in Section 3.2, we remove test set images from DataComp to avoid contamination. Image classification datasets range from satellite imagery recognition to classifying metastatic tissues. In total we have (with some overlap): 22 of the datasets evaluated in Radford et al. , 6 ImageNet distribution shifts (i.e., ImageNet-Sketch , ImageNet-V2 , ImageNet-A , ImageNet-O , ImageNet-R , and ObjectNet ), 13 datasets from VTAB , and 3 datasets from WILDS . Retrieval datasets include Flickr30k , MSCOCO , and the WinoGAViL commonsense association task . To aggregate results over all evaluation tasks, we average the preferred metric for each task.
DataComp adopts a zero-shot evaluation protocol: models are tested without training on the evaluation tasks. This approach is computationally efficient and measures a model’s ability to perform well without any additional training. We find a strong rank correlation () between performance in linear probe zero-shot settings (Appendix Figure 16). Additional details are in Appendix O.
Baselines
We study six simple filtering methods for the filtering track; see Section P.1 for further details.
We simply use the entire pool as the subset, without any filtering. Since each pool size is equal to the sample budget, training consists of one pass over the data.
To isolate the effects of increasing the compute budget from increasing the dataset size, we form subsets consisting of 1%, 10%, 25%, 50% and 75% of the pool chosen at random.
We consider many simple filtering operations inspired by Schuhmann et al. and Byeon et al. : filtering by language (English captions, using either fasttext or cld3 ); filtering by caption length (over two words and five characters); and filtering by image size (smaller dimension above 200 pixels and aspect ratio below three). We also experiment with combining language and caption length filtering and combining language, caption length, image size fitering. Unless otherwise specified, “basic” refers fasttext English, caption length, and image size filtering.
We experiment with CLIP score filtering (also employed by LAION), where we take only examples having cosine similarity scores between CLIP image and text embeddings that exceed a pre-defined threshold. We investigate a range of thresholds and two OpenAI CLIP models for computing the scores: the ViT-B/32 model (as in LAION) and the larger ViT-L/14. We also combine CLIP score thresholds and cld3 English filtering to reproduce the LAION-2B filtering scheme. Table 16 in Section P.1 summarizes the different CLIP score configurations.
We select examples that contain text overlapping with ImageNet class names, which serve as a proxy for relevance to downstream tasks. Specifically, we select English captions (according to fasttext) that contain words from ImageNet-21K or ImageNet-1K class synsets.
We select a subset of examples whose visual content overlaps with ImageNet classes. After applying English language (fasttext) and caption length filtering, we cluster the image embeddings extracted by the OpenAI ViT-L/14 model for each image into 100K groups using Faiss . We then find the nearest neighbor group for every ImageNet training example, and keep examples belonging to these groups. We apply this procedure using either ImageNet-21K (14M images) or ImageNet-1K (1.2M images), forming two subsets.
2 BYOD baselines
We experiment with multiple external data sources, including four moderately sized datasets (10 to 58M samples) studied by Nguyen et al. —CC12M , YFCC15M , RedCaps and Shutterstock —and the larger LAION-2B . Additional experiments, along with more details about the data sources are provided in Appendix P.2. We consider these data sources as they are and do not perform additional preprocessing. We also present experiments combining some of the data sources (using only the external datasets, or in addition to data from our pool).
Results and discussion
Our key results are in Table 3. Most notably, the intersection between image-based filtering and CLIP score filtering excels on most tasks. The exception is at the small scale and for retrieval datasets.Cherti et al. also observe that models rank differently on classification and retrieval tasks. Furthermore, other filtering strategies like basic, CLIP score, image-based, text-based filtering show better downstream performance when compared to no filtering. A much larger suite of experiment results can be found in Appendix R.
We hope DataComp catalyzes the search for the next generation of multimodal datasets. We contribute DataComp-1B, which is the output of the Image-based CLIP score (L/14 30%) baseline filter at the xlarge scale of the filtering track. Our dataset is comprised of 1.4B samples, which not only is smaller than the LAION-2B dataset with 2.3B samples, but also comes from a smaller pool. Nevertheless, a CLIP L/14 trained on DataComp-1B outperforms the LAION-2B competitor by 6.1 percentage points on ImageNet (see Table 1). Moreover, training on DataComp-1B improves ImageNet accuracy by 3.7 percentage points over OpenAI’s ViT-L/14 trained with the same compute budget. Additionally, even if we restrict ourselves to 400M samples, we can still find a subset of DataComp-1B that outperforms OpenAI’s ViT-L/14, as seen in Table 24. These results demonstrate the impact that DataComp can make and provide a foundation upon which participants can build.
Appendix P.2 Table 18 shows results for several baselines in the BYOD track. We find several instances where adding external data sources improves performance over using just data from CommonPool. For example, at the large scale, combining CLIP-filtered data from CommonPool with external data from CC12M , YFCC15M , RedCaps and Shutterstock boosts ImageNet accuracy by 4.3 percentage points. See Appendix P.2 for more experiments and details.
In Figure 2, we see that randomly selecting subsets of the pool has little effect and degrades performance substantially when only small fractions are used. When filtering with CLIP scores, the optimal training set comes from selecting 30% of the pool with the highest scores. The difference in performance trends between random subsets and CLIP score filtering highlights the importance of filtering strategies for selecting samples.
2 DataComp design analyses
To validate our pool construction, we show that we can build datasets comparable to LAION-2B by employing their filtering technique on our pool. LAION-2B selects all samples where the caption is in English and the cosine similarity score from a trained ViT-B/32 CLIP model is above 0.28. We compare this filtering approach on our pool using the same number samples, 130M samples at the large scale. We find that the different data sources perform comparably: 55.3% vs 55.7% accuracy on ImageNet, and 0.501 vs 0.489 average performance over our evaluation sets using our pool and LAION-2B, respectively.
We find that the ranking between filtering strategies is typically consistent across different scales. This is illustrated in Figure 3, which shows that the baselines at small and medium scales are positively correlated. Moreover, as shown in Appendix Table 22, the rank correlations of performance is high, between 0.71 and 0.90 for different scale pairs.
DataComp fixes the training procedure, so a natural question is whether better datasets from DataComp are better outside of DataComp. While DataComp-1B is trained at the xlarge scale, we show in Appendix Table 23 that even when substituting the ViT-L/14 for a ViT-B/16 or ViT-B/32, training on DataComp-1B outperforms training on OpenAI’s WIT and LAION-2B. Additionally, we found that modifying hyperparameters such as training steps and batch size minimally affects the relative ordering of different data curation methods on downstream performance. Details on hyperparameter ablations are in Appendix L.
3 Evaluation trends
Similarly to Kornblith et al. , in Appendix Figure 25 we find that ImageNet performance is highly correlated with the average performance across all datasets we study, with an overall correlation of 0.99. Note that unlike Kornblith et al. we evaluate zero-shot performance rather than transfer learning. However, ImageNet performance is not representative of all evaluation tasks, as the correlation between ImageNet accuracy and accuracy on other individual datasets varies substantially, in some cases even exhibiting a negative correlation, as discussed in Appendix R.
While typical models trained on a target task suffer large performance drops under data distribution shift, zero-shot CLIP models are known to exhibit strong performance across many distributions . In Appendix Figure 26, we show that CLIP models trained with data from our pool are more robust to distribution shift than ImageNet-trained models from Taori et al. ’s testbed. Examining geographic diversity, we find that our models are better than ImageNet-trained models, but fall short of models fine-tuned on diverse curated datasets (see Appendix Figure 21). We also perform a face classification analysis and identify demographic biases in our models: notably, the BYOD datasets we consider can increase the risk of misclassification. See Appendix Q for more fairness and diversity analyses.
Limitations and conclusion
In terms of societal risks, creating an index of image-text pairs from the public internet can be problematic. The internet contains unsafe, toxic, and sensitive content, which ideally should not percolate into machine learning datasets. Though we take steps to remove NSFW content and blur human faces to protect privacy, we hope future work will further explore the biases and risks from CommonPool and DataComp-1B. We see several additional directions for future work, including 1) Curating more data sources. 2) Improved data filtering algorithms. 3) Further supervision signals (e.g., image captions coming from captioning models). 4) Additional input modalities (e.g., video, 3D objects). 5) Broader evaluations for vision-and-language and robotics tasks.
Overall, we see DataComp as a first step towards improving training datasets, and hope our new benchmark will foster further research. By providing a controlled experimental setting, DataComp enables researchers to iterate on dataset design on rigorous empirical foundations. We open-source all of our code, data, and infrastructure, and hope these resources will help the community build the next generation of multimodal datasets.
Acknowledgements
SYG and JH are supported by NSF Graduate Research Fellowships. GS is supported by the Onassis Foundation - Scholarship ID: F ZS 056-1/2022-2023. GD has been supported by the Onassis Fellowship (Scholarship ID: F ZS 012-1/2022-2023), the Bodossaki Fellowship and the Leventis Fellowship. This research has been supported by NSF Grants AF 1901292, CNS 2148141, DMS 2134012, TRIPODS II-DMS 2023166, Tripods CCF 1934932, IFML CCF 2019844 and research gifts by Western Digital, WNCG IAP, UT Austin Machine Learning Lab (MLL), Cisco, the Len Blavatnik and the Blavatnik Family Foundation, the Stanly P. Finch Centennial Professorship in Engineering, Open Philanthropy, Google, Microsoft, and the Allen Institute for AI.
We would like to thank Amro Abbas, Danny Bickson, Alper Canberk, Jessie Chapman, Brian Cheung, Tim Dettmers, Joshua Gardner, Nancy Garland, Sachin Goyal, Huy Ha, Zaid Harchaoui, Ari Holtzman, Andrew Hundt, Andy Jones, Adam Klivans, Ronak Mehta, Sachit Menon, Ari Morcos, Raviteja Mullapudi, Jonathon Shlens, Brandon McKinzie, Alexander Toshev, David Grangier, Navdeep Jaitly, Kentrell Owens, Marco Tulio Ribeiro, Shiori Sagawa, Christoph Schuhmann, Matthew Wallingford, and Ross Wightman for helpful feedback at various stages of the project. We are particularly grateful to Daniel Levy and Alec Radford for early encouragement to pursue this project and feedback on the experimental design.
We thank Stability AI and the Gauss Centre for Supercomputing e.V.https://gauss-centre.eu for providing us with compute resources to train models. We are thankful for the compute time provided through the John von Neumann Institute for Computing (NIC) on the GCS Supercomputer JUWELS Booster at Jülich Supercomputing Centre (JSC), and for storage resources on JUST granted and operated by JSC, as well as computing and storage resources from the Helmholtz Data Federation (HDF).
References
Appendix
We provide concrete rules below for the two competition tracks that comprise DataComp: filtering and BYOD. Additionally, we provide a checklist, which encourages participants to specify design decisions, which allows for more granular comparison between submissions.
Participants can enter submissions for one or many different scales: small, medium, large or xlarge, which represent the raw number of image-text pairs in CommonPool that should be filtered.
After choosing a scale, participants generate a list of uids, where each uid refers to a CommonPool sample. The list of uids is used to recover image-text pairs from the pool, which is used for downstream CLIP training.
Participants are not allowed to modify the training procedure. Hence, changing hyperparameters, model architecture, optimizer, compute budget, or number of training steps is not allowed. Changing any other training details is also not allowed.
Participants are strongly encouraged to submit and open-source both the list of uids and the code used to generate this list; however, this is not required.
To avoid overfitting, we do not permit running any code or algorithmic dependence on the test images of the evaluation tasks. However, use of other images associated with these tasks (e.g., supervised training sets) is permitted.
Participants can use templates or class labels from the downstream tasks in their filtering algorithms.
For clarity, we include some examples of permitted and forbidden uses:
We permit using the ImageNet class label “triceratops” in a filtering algorithm.
We forbid examining individual or aggregate predictions on the test sets of the evaluation tasks.
A.2 Bring your own data track: amendments
To facilitate more open-ended exploration, we provide amendments to the Track 1 competition to allow for more diverse submissions in Track 2.
Participants are allowed to augment CommonPool data with existing datasets, so long as these data sources do not contain test images from the evaluation tasks. Participants can use data from any CommonPool; however, they are not required to do so.
Assembling one’s own dataset is allowed; however, test images from the evaluation tasks can neither be contained nor otherwise used to construct said dataset. We encourage releasing the image urls or the images themselves in addition to the text for each image. We also encourage rigorous documentation of face-blurring and other data safety checks (see Section 3.2 for more details). We reserve the right to run our own safety code on participant provided data and disqualify entries that do not meet adequate safety standards.
The following checklist provides the basis for more fine-grained comparison between submissions.
Images from the evaluation tasks are included in my submission. If yes, please specify which datasets.
I used an existing datasets (e.g., YFCC100M ) in my submission. If yes, please specify which datasets. (Note: applies to BYOD only)
I curated my own data. If yes, please provide (1) image data or urls, (2) text for each image, (3) list of safety steps taken including but not limited to face blurring, explicit content image and text filtering. (Note: applies to BYOD only)
Appendix B Contributions
For this section, contributors are ordered alphabetically.
Giannis Daras, Alex Fang (content filtering lead), Samir Yitzhak Gadre (metadata lead), Ryan Marten (deduplication lead), Vivek Ramanujan, Vaishaal Shankar, George Smyrnis (face blurring lead)
B.2 Participant tooling
Romain Beaumont, Yair Carmon, Alex Fang, Jonathan Hayase (lead), Gabriel Ilharco, Vivek Ramanujan, Vaishaal Shankar, Georgios Smyrnis
Mehdi Cherti, Gabriel Ilharco, Jenia Jitsev, Vivek Ramanujan, Georgios Smyrnis, Mitchell Wortsman (lead)
Romain Beaumont, Yonatan Bitton, Mehdi Cherti, Dhruba Ghosh (lead), Gabriel Ilharco
B.3 Baselines
Yair Carmon, Rahim Enterazi, Alex Fang, Samir Yitzhak Gadre, Gabriel Ilharco, Kalyani Marathe, Thao Nguyen, Eyal Orgad (co-lead), Georgios Smyrnis, Mitchell Wortsman, Jieyu Zhang (co-lead)
Alex Fang, Gabriel Ilharco, Samir Yitzhak Gadre
B.4 Leadership and Advising
Romain Beaumont, Yair Carmon, Alexandros G. Dimakis, Ali Farhadi, Hannaneh Hajishirzi, Jenia Jitsev, Pang Wei Koh, Ranjay Krishna, Stephen Mussmann, Sewoong Oh, Alexander Ratner, Olga Saukh, Ludwig Schmidt, Vaishaal Shankar, Shuran Song, Richard Vencu
Yair Carmon, Alexandros G. Dimakis, Jenia Jitsev, Sewoong Oh, Ludwig Schmidt, Vaishaal Shankar
Appendix C Additional related work
Here we expand on the related work described in Section 2.
Image dataset safety is an active area of research, especially in the context of large-scale dataset construction. In addition to Birhane et al. , who study problematic content in LAION-400M, Yang et al. study the ImageNet dataset and reveal limitations associated with the ImageNet curation strategy—with negative implications for downstream model fairness. Prabhu & Birhane also study the ImageNet dataset and find pornographic content. Both Birhane et al. and Prabhu & Birhane survey ethical conundrums and harms that are borne out of improper dataset curation. In an effort to combat dataset toxicity, we conduct NSFW preprocessing (Section 3.2, Appendix E) and blur detected faces (Section 3.2, Appendix G) during pool construction. We also conduct preliminary fairness evaluations (Section 5.3, Appendix Q) for models trained on our data. We hope CommonPool will serve as a research artifact for future work examining dataset safety.
Beyond data selection, Chan et al. investigate the effects of dataset distribution on emergent properties of transformers, while Fang et al. look at the relationship between data and model robustness to distribution shifts. We hope our extensive evaluation suite comprised of 38 diverse tasks will facilitate similar studies when training multimodal models at large scale.
Others study how to reduce the burdens of training data annotation in the curation process. Classic approaches include distant supervision , crowd-sourced labels , heuristic rules and feature annotation , among others. A recent line of work known as data programming or programmatic weak supervision attempts to reduce annotation cost and is found in many industry applications . In data programming, developers write programmatic labeling functions to automatically label a large amount of unlabeled data. The labeling functions could produce noisy and conflicting labels, so researchers have developed methods to aggregate noisy votes to produce the final training labels .
Previous literature also studies methods for training data attribution, which seek to link a model’s behavior (e.g., its accuracy on a particular task or subset of data) to particular subsets of its training data. Such methods include influence functions, a classic technique from robust statistics that uses a second-order Taylor expansion to approximate the effect of removing a training point on the learned model parameters , as well as methods that fit attribution functions directly to the dynamics of repeated training runs . Training data attribution methods assume that we have already trained a model, though they can be subsequently used to refine the training data (e.g., by identifying potentially mislabeled training points ). Our focus in this paper is instead on data curation methods—that is, methods for selecting a subset of the training data to train a model in the first place.
In the context of natural language processing, Swayamdipta et al. proposes a tool for characterizing samples in a dataset based on training dynamics, labelling instances as ambiguous, easy to learn or hard to learn. Previous literature such as work by Le Bras et al. , Li & Vasconcelos , Gururangan et al. advocate for removing easy instances from the training data. Ethayarajh et al. propose a measure of how difficult a dataset is to learn, -usable information. Such techniques could be promising directions of further exploration in the context of our benchmark.
Finally, another related line of work is studying scaling trends. In addition to Sorscher et al. , researchers have investigated how model performance changes as a function of compute budget, model size, and number of training samples . However, this line of work does not consider how dataset design may affects scaling trends. Beyond dataset size, we measure the effects of different dataset sources and filtering strategies. While scaling trends are central to our investigations, the purpose of our benchmark is to search for the next generation of large multimodal datasets to facilitate more accurate and reliable models.
Appendix D Parsing Common Crawl
Common Crawl releases metadata files for the websites that they index (i.e., WAT files). They release these files approximately once a month. We consider all files available from 2014 through November of 2022. We first parse these files, utilizing Apache Spark to extract image urls and corresponding alt-text. We map each url, text pair to a uid hash and remove duplicates. This results in 88 billion url, text pairs, which are randomized via a distributed shuffle. Note, we do not consider image content when running uid deduplication at this step. Hence, two identical images with different urls and the same caption would both be retained.
Appendix E Not safe for work (NSFW) filtering
Our data is sourced from Common Crawl, which contains snapshots of the web. Therefore, we apply multiple layers of NSFW content filtering to remove problematic images and captions from CommonPool.
First, we filter our captions with Detoxify , a language model for toxic comment classification. Specifically, we use the multilingual XLM-RoBERTa variant. The model outputs scores between zero and one for the following categories: toxicity, severe toxicity, obscene, identity attack, insult, threat, and sexually explicit. As we had no ground truth for our data, we manually spot check a 1 million random subset of CommonPool at varying thresholds. We found that a threshold of 0.1 provided good coverage of filtering out NSFW text. If any of the detoxify category scores exceeds the threshold, the sample is discarded. Qualitatively, we found that the model struggled with multilingual content, acronyms, and innuendo. Even at 0.1, we noticed there are some captions that are NSFW. However, lowering the threshold further heavily affected false positives. We therefore use a 0.1 threshold for all NSFW categories, which on a random subset of one million captions achieves positive rates shown in Table 4.
Second, on the vision side, we use a modified version of LAION-5B’s CLIP-based binary classification NSFW model, which takes CLIP ViT-L/14 visual embeddings as input. We remove the initial multi-category encoder from the model, and retrain on the same data with an initial normalization layer followed by a 4-layer multilayer perceptron. Our retrained model matches the performance of the original model on their manually annotated testset. Specifically, we achieve 97.4% classification accuracy on a held out test set compared to 96.1% for the original LAION NSFW image filtering model. Additional details about the training data can be found in Appendix C.5 of the LAION-5B paper. In brief, the training data contains 682K images that is roughly balanced with images from safe for work and NSFW categories.
To evaluate our model and determine a threshold, we used Google Vision API’s SafeSearch explicit content detector to generate labels for an 40,000 random subset of our candidate pool. Specifically, an image is NSFW if SafeSearch classifies it as likely or very likely adult (i.e., sexually explicit). As shown in Table 5, we found that by thresholding at 0.1 we achieve high recall relative to SafeSearch and very few true positives after manual review. We also manually reviewed images classified by SafeSearch as likely or very likely racy and found that the images were either benign, subjectively suggestive but not explicit, or already found in the set of images labeled as adult.
Appendix F Deduplication against evaluation sets
To prevent data leakage, we filter CommonPool by removing duplicate and near-duplicate matches of evaluation set images. See Figure 4 for example query images from Common Crawl and corresponding near-duplicates in our evaluations sets. We consider images as duplicates when the cosine similarity between a query (Common Crawl image) feature and a reference (evaluation image) feature is higher than a fixed threshold. We employ the deduplication model proposed by Yokoo , which earned 1st place in the Facebook AI Image Similarity Challenge (ISC) . We choose a cosine similarity threshold of 0.604169 to maximize the true duplicates detected, without removing too many false duplicates from the pool. We compare against OpenAI’s CLIP ViT-B/32 as a baseline on ISC. We find that for our threshold, the ISC model achieves precision 0.9 and recall 0.8. At a threshold of 0.96, CLIP achieves the same precision 0.9, but a significantly worse recall of 0.02. Approximately 2.8% of downloaded samples are flagged as evaluation set near-duplicates.
To verify the performance of our de-duplication models with greater granularity, we modify the evaluation procedure in Douze et al. to include transformations which are representative of naturally-occurring duplications on the Internet. Specifically, we study: 1) jpeg compression (encoding), 2) image flips, 3) image rotations, 4) aspect ratio modifications, and 5) grayscaling. To do this, we sample 20% of the images from each of our evaluation datasets uniformly at random to serve as a reference set of about 140,000 images. Next we sample 560,000 images uniformly at random from LAION-2B to serve as distractors, for a 4-to-1 distractor to reference ratio. Finally, we apply each of the augmentations above and use threshold filtering to determine duplicates. Figure 5 shows the results from the deduplication model compared with OpenAI’s CLIP ViT-L/14. At high recall values, we see that CLIP filtering results in removing over 2 the data as that of the deduplication model from Yokoo .
Appendix G Face blurring
As an extra step to safeguard against issues of privacy that may arise from the use of data scraped from the web, we include face blurring as part of our pool creation. To create face metadata, we use the SCRFD face detector to extract bounding boxes for the faces in our images. These bounding boxes are included as part of the image metadata in our pool. We make use of the pretrained SCRFD-10G model. We use the same preprocessing as the one described in the official repository of the paper, with the exception of providing input images (by padding each image to square and then resizing) to limit computation costs. Invoking this model provides us with bounding boxes along with an associated score, which we then compare against a threshold of to keep or discard this bounding box. This threshold is the default one used in the repository of SCRFD for the visualization of bounding boxes, and we found it to perform well on our data as discussed next.
In Table 6 we can see the result of face detection on a set of 3293 images from CommonPool. We evaluate the detection on whether the image has visible faces or not (where images such as cartoon drawings of non-real human faces are not considered as positives), and whether the detector has detected these visible faces. We considered an image as a true positive if all the clearly visible faces in the image were detected, based on the above thresholding process. We did not do extensive box labeling. True positives are instead determined by human inspection. We compare the quality of these detections with the Amazon Rekognition system, which is the one upon which the face detections on ImageNet were based . Note that in this scenario, the recall of the detectors is more important than precision (as detecting a few more bounding boxes across our pool does not affect privacy).
To utilize these bounding boxes on our data, we apply a standard blurring pipeline, as proposed by Yang et al. . The result of this process is an image where the faces is blurred and there is a smooth transition from blurred to clean parts of the image. In Figure 6 we see the distribution of faces for the small CommonPool. Note that the majority of images do not contain faces.
As part of our competition pipeline, images are by default blurred during the download process. In Table 7 we can see the results of training on a set of images with the size of our medium scale after filtering with each method, with and without the application of face blurring as provided by our detector. We can see that the difference in performance is small, which suggests that the application of face blurring does not significantly affect the performance on our downstream tasks. However, we note that this design decision may be more detrimental in generative settings, especially when a generative model needs to output faces. Our competition is primarily focused on discriminative tasks, and as such when designing our dataset, we wished to prioritize the safety and privacy of individuals through blurring faces in our download tooling by default.
Finally, we evaluated the detector we used for potential biases. More specifically, we used the detector on the validation set of the FairFace dataset . We found that the central face of the image was detected in all the images of the validation set, regardless of subgroup annotate in the dataset.
Appendix H DataComp CommonPool creation pipeline
Creating CommonPool was a multistep process, which involved (1) parsing image urls and alt-text from Common Crawl dumps and downloading these images, (2) tagging images with metadata and (3) conducting safety content filtering and evaluation set duplication. In this section we provide an overview of the data pipeline used to create CommonPool. For an overview of our “data funnel” see Figure 7.
For the first step, we use parse Common Crawl metadata files to harvest image-text pairs (Section D). We use img2dataset to obtain 16.8B downloaded samples. This is the first, unfiltered version of CommonPool, and contains only basic information for our images (i.e., the original image height, width, and alt-text caption). During this step we also resize images such that their largest dimension does not exceed 512 pixels. This eases storage requirements for large images, but is still larger than the 224 pixel resolution used for later training stages.
For the second step, we process our unfiltered pool and create richer metadata for each image-text pair. We generate the following for each sample:
CLIP ViT-B/32 and CLIP ViT-L/14 image and text features, with their associated similarities.
NSFW scores for the image and the text, using the analysis described in Appendix E.
Deduplication score for the image, as described in Appendix F.
Bounding boxes for faces detected in the image, using the method described in Appendix G.
For the third and final step, we filter our image-text pairs based on the metadata generated during the second stage. We filter out image-text pairs where the NSFW and deduplication scores exceed the respective thresholds (Section E). From the images that pass through this filtering, we keep only the desired amount (e.g., 12.8B images from the xlarge CommonPool). Smaller pools are telescoping subsets of larger pools. We package the metadata and image urls, which is made publicly available to the participants. Note, we do not release raw image data but rather image urls pointing to images.
A summary of the metadata for each sample is found in Table 8. To validate our pipeline for duplication and CLIP feature correctness, we also take ImageNet train though metadata generation as a unit test. Using the deduplication features, we detect that 100% of the images are in fact duplicates. Additionally using the CLIP ViT-B/32 and CLIP ViT-L/14 image features and corresponding text features from OpenAI’s 80-prompt ensemble, we achieve 63.36% and 75.54% top-1 accuracies, which match the performance reported in the CLIP paper .
When creating pools of different scale (i.e., number of samples), we ensure that smaller pools are subsets of larger pools. For instance, the small CommonPool is a subset of the xlarge CommonPool.
After CommonPool is created, the participants can then download the final image-text pairs using the provided files via img2dataset. To further ease the computational burden on participants, we additionally provide metadata for each sample in CommonPool. Note that when downloading, our img2dataset configuration automatically blurs faces. Hence this is an automatic step on not something participants must do ad hoc.
Appendix I CommonPool statistics
To provide more information about the kinds of samples in our CommonPool, we conduct additional analysis on the small pool, which is an i.i.d. sample of downloaded data and a subset of the larger pools.
In Figure 8 we show CLIP similarity similarity scores between images and their corresponding text. We notice a flatter distribution of CLIP ViT-L/14 scores than corresponding B/32 scores.
Turning our attention to images in CommonPool, in Figure 9, we visualize the aspect ratios and sizes of original images (i.e., before they are downloaded and resized). In Figure 10, we display a distribution of image height and width after download resizing. Notice that the majority of images are around pixels, which is the final resized resolution used for training.
Analysing the textual component of each sample, we visualize frequency of the number of CLIP BPE tokens in the captions (Figure 11) and most common languages (Figure 12). Token counts follow a long-tailed distribution with much more mass in the short sequence range, while English is the predominant language in CommonPool according to fasttext and cld3.
We also look at url statistics. In Figure 13 we see common domain names in CommonPool (e.g., wordpress domains) and common suffixes (e.g., .com or .net).
Appendix J Efficient training on data subsets
When training at large scale, it is important to use efficient access patterns to load training data. This typically means that data must be loaded using large sequential reads instead of random reads in order to maximize throughput. In DataComp, this is facilitated by the WebDatasethttps://github.com/webdataset/webdataset format which stores the training examples in tar files (called “shards”) and WebDataLoader which makes it easy to load data stored in this format.
Given an arbitrary subset of a pool, we would like to efficiently train on that subset. Because WebDataset format does not permit efficient random access (a feature inherited from tar), we must read through the entire pool to select the required images. There are two ways to implement this filtering:
Filter during training: we apply a predicate during training data loading that discards data not present in the subset.
Filter before training: we iterate over the pool, selecting the images in the subset, and write them to a new WebDataset.
After some profiling, we concluded that option 1 had too much overhead in the case where the subset is much smaller than the pool. To see why, note that if the subset is an -fraction of the pool size, then we would end up reading a factor more data than needed for training. Instead, we give an implementation of option 2, which performs at most twice as many reads as needed for training.Since in DataComp, the number of examples seen is equal to the pool size.
Our tool, called the resharder, reads a set of uids in NumPy array format, scans through the pool, selecting those examples, and writes them to a new WebDataset. The resharder uses multiprocessing to make good use of hardware and can be distributed over many computers to further increase throughput. The resharder also supports streaming data to and from cloud storage such as Amazon S3. The resharder is provided to participants as part of the competition tooling.
Appendix K Effect of duplicates in the training data
Given that CommonPool was constructed by scraping the web for image and text pairs, there is a likelihood that some of our images are duplicates of each other, even if they originated from different web sources and have different captions. Here we examine the effect of removing such duplicates. We used the technique proposed by Webster et al. , where CLIP image features are first compressed and then used to do an approximate nearest neighbor search. After this process, two images and are considered duplicates if , where is some threshold and is the distance of a vector with its quantized version used for approximate nearest neighbor search. For each image, we search duplicates across its nearest neighbors, and keep it if it’s the one with the highest CLIP ViT-L/14 similarity score across its duplicates. Results can be seen in Table 9, both when this technique is used by itself and in conjunction with ViT-B/32 filtering. We can see that results are similar to when only using CLIP filtering.
Appendix L Hyperparameter ablations
Recall that in DataComp, we freeze the training procedure and hyperparameters to focus the competition on dataset curation. However, this leads to the natural question: do “better” datasets (i.e., datasets that lead to higher accuracy models on zero-shot downstream tasks) remain consistent when training is modified. Hence we ablate key experimental choices: batch size, model architecture, and number of training steps.
We ablate over the batch size hyperparameter, doubling the batch size at the medium scale, but holding all other hyperparameters constant. As see in Table 10, we find that the delta rankings are largely consistent, for both ImageNet and Average performance, with rankings changing by at most plus or minus one position. More specifically, rank correlation before and after doubling batch size is 0.96 for ImageNet and 0.98 for the Average over 38 datasets metric.
L.2 Model architecture
We choose to use the ViT architecture because of favorable CLIP scaling trends over vanilla ResNets as reported by Radford et al. . However, we still hope that better datasets for downstream ViT performance will lead to better datasets to train convolutional architectures. We look at the medium scale, swapping the ViT-B/32 architecture with a ConvNeXt model with matched giga multiplier–accumulate operations (GMACs). Looking at Table 11, we see that ranking of different filtering methods is again relatively consistent (i.e., 1.0 rank correlation for ImageNet and 0.87 rank correlation for the average metric). We conclude that improvements in dataset filtering have potential to improve more than just CLIP ViT model performance.
L.3 Number of training steps
Recall that one of our major design decisions for DataComp is to fix the hyperparameters associated with model training, following closely hyperparameters from prior work . We choose to fix hyperparameters to place emphasis on data curation and remove confounders arising from hyperparameter differences between participants. Here we ablate our hyperparameter configuration by training small baselines for 10 more steps. In Figure 14 we see positive correlation for ImageNet accuracy for the ablated and original hyperparameter configurations. We see similar correlation for average performance. See Table 12 for specific values.
Appendix M Detector-based baselines
While controlling for factors such as class balance is common in the supervised settings, experimenting with analogous strategies in the context of multimodal datasets and CLIP training is a pertinent direction. Towards this end, we use the Detic detector to annotate the medium pool (128M samples) by extracting bounding boxes and class labels for the 1203 LVIS objects categories. Following the original Detic paper, we retain predictions whose confidence score exceeds 0.5. Based on these annotations, we construct the following five strategies:
Object exists: Subset for which there exists at least one detection from the 1203 LVIS categories.
Object centered: Subset for which there exists at least one detection from the 1203 LVIS categories with a bounding box center falling in the center grid cell of a 3x3 grid superimposed on the image.
Balancing by class: We define 1204 buckets—1203 buckets corresponding to the LVIS classes and an additional bucket for images that do not have any detections. For each image in the medium pool, we assign the image to the bucket(s) corresponding to the detected classes. We then construct a dataset such that there are an equal number of samples from each bucket and the total number of samples specified by a particular scale (e.g., 128M samples for medium scale). Note, for rare classes there can be many repeated samples and for common classes only a subset of the total samples will be in the dataset.
Balancing by position: We define 26 buckets—0, 1, …, 24 corresponding to 5x5 grid locations in an image. An image is added to bucket(s) when it contains a bounding box whose center falls in the bucket’s grid cell. The 25th bucket contains images for which there are no detections. We again construct a dataset such that there are an equal number of samples from each bucket.
Balancing by count: We define 12 buckets—0, 1, …, 10 corresponding to zero to ten detections in an image and a twelfth bucket corresponding to images with more than ten detections. We yet again construct a dataset such that there are an equal number of samples from each bucket.
We employ each of these strategies on the medium scale. Since the above strategies can be composed with any starting pool, we additionally apply each of the above Detic-based strategies to our previous best medium scale filtered pool: Image-based CLIP score (L/14 30%). This yields five more datasets for 10 baselines in total.
Our results are summarized in the Table 13. In summary: 1) The Image-based CLIP score (L/14 30%) baseline still performs best. 2) Balancing data in the context of multimodal CLIP training remains an open problem. All balancing strategies lead to divergence of the CLIP contrastive loss and result in poor model performance. We hypothesize that this is due to the long-tailed nature of the data distribution, which leads to many repeated samples in our balanced data construction. This in turn, increases the likelihood that samples are contrasted with themselves in the loss computation.
Appendix N Training details
The full set of hyperparameters used for each scale is shown in Table 14. For choosing hyperparameters, we follow the OpenCLIP library , an open source reproduction of OpenAI’s CLIP. For the small, medium, and large tracks, these hyperparameters are equal to those in the CLIP paper, except with reduced batch size so that training runs on reasonable hardware. For the xlarge track, batch size is increased from that in OpenAI’s CLIP to accelerate training by allowing the use of many GPUs simultaneously with high utilization. For this run we also double the learning rate following prior work .
Appendix O Evaluation details
Models are evaluated over a wide range of 38 tasks to measure proficiency in various domains. We include 22 of the 27 classification tasks in the test suite of Radford et al. , excluding the few datasets that have license restrictions, are in video format, or are no longer available in their original form. We include 6 datasets that were designed to test generalization of models trained on ImageNet. We also include a majority of the Visual Task Adaptation Benchmark, excluding 3 datasets that are ill-suited for zero-shot evaluation . We include 3 datasets from the WILDS benchmark, which tests robustness to distribution shifts and spurious correlations . Finally, we include 2 additional datasets, Dollar Street and GeoDE, which test robustness of classification performance across income levels and geographical regions . Furthermore, we evaluate zero-shot image and text retrieval on the Flickr30k and MSCOCO datasets, and image association on the WinoGAViL dataset . The complete list of evaluation tasks is given in Table 15. We show a sample from each dataset in Figure 15.
Since we perform zero-shot evaluation, prompt and class name selection is important, and can have a significant impact on the results. To avoid heavy prompt engineering and overtuning to individual models, we opt to use the prompt templates used in Radford et al. whenever possible. Most datasets come with pre-defined class names, but some are overwritten with more descriptive labels, again based on previous literature. For datasets with no precedent in zero-shot evaluation, we reuse prompt templates from other datasets with a similar domain and task (e.g., SVHN is evaluated with MNIST prompts and class names).
For the majority of classification tasks, the primary evaluation metric is accuracy. For certain datasets with class imbalances, we instead compute mean per-class accuracy, as done in Radford et al. . On the WILDS benchmark datasets, we use the primary metric specified for each dataset on their leaderboard. Dollar Street and GeoDE test model generalization across socioeconomic and geographic diversity. Thus, for Dollar Street, we compute worst-group top-5 accuracy, with groups defined by income level, emulating Rojas et al. ; for GeoDE, we compute worst-group accuracy, with groups defined by region (Africa, Americas, West Asia, East Asia, Southeast Asia, and Europe), as defined in Ramaswamy et al. . For the image-text retrieval tasks, Flickr and MSCOCO, we compute both image and text recall (fraction of text captions for which the correct image was selected and vice versa), and plot their arithmetic mean. On WinoGAViL, we compute the Jaccard score (intersection-over-union) for each example, and show results for the harder samples (10 and 12 candidates). More information on WinoGAViL evaluation can be found in Bitton et al. .
For five of our evaluation tasks (the two CLEVR tasks, the two Camelyon tasks, and KITTI) the zero-shot performance of all evaluated models appears to be close to that of random guessing, and lack correlation to the type of filtering method used (see Figure 27). Consequently, we studied performance averaged only on the remaining 33 tasks, but found not substantial qualitative differences in our results. As a result, we opted to report the average on the full evaluation suite throughout our study.
One critical decision in DataComp is how exactly to evaluate models and whether or not to fine-tune models on evaluation tasks (i.e., supervised fine-tuning directly on task training sets). We opt for zero-shot evaluation, where a models are applied to downstream tasks directly to 1) ease computational burden on participants and 2) measure the out-of-the-box generalization capabilities of our models. To validate this design decision, we conduct linear probes on all models presented in Tables 3 and 18 on ImageNet. We follow a standard probing protocol and fine-tune the last linear layer from zero-shot initialization for 40 epochs with learning rate 1e-3, batch size 256, AdamW optimizer with default settings with the exception of weight decay (that we set to zero), and a cosine annealing schedule. As seen in Figure 16, zero-shot and linear probe performance follow similar trends for both filtering and BYOD tracks. Moreover the Spearman rank correlation between the two protocols over the models considered is 0.99 for the filtering track and 1.0 for BYOD. This suggests that better zero-shot models on ImageNet are correlated with better representations of linear probe fine-tuning on ImageNet.
O.1 Visual Question Answering
In addition to our evaluation suite containing multiple classification and retrieval tasks, we conducted experiments on visual question answering. More specifically, following Shen et al. , we use the CLIP models to contrast images with prompts formed by the questions and each candidate answer, without fine-tuning (i.e., in a zero-shot setting). Using the VQA v1 dataset , for each candidate answer, we construct a text prompt that also includes the question following the template Question: [question text] Answer: [answer text], as in Ilharco et al. . This text is then fed to CLIP’s text encoder. As previously noted by multiple authors, CLIP models struggle on this task, potentially due to the mismatch between the text in the downstream task and the captions seen during pre-training Shen et al. , Ilharco et al. , Song et al. . Nonetheless, we observe a strong correlation between VQA performance and ImageNet accuracy (0.877) and between VQA performance and average performance on our full evaluation suite. Full results are shown in Figure 17.
Appendix P Baseline details
Here we provide additional details on the creation of our baseline subsets. To highlight the qualitative differences between the filtering strategies we also provide visualization for No filtering (Figure 18), Basic filtering (Figure 19), and CLIP score (L/14 30%) (Figure 20), which can all be found in Table 3. Notice that No filtering gives relatively noisy data (e.g., matching a bicycle with a caption: “IMG_2187.jpg”), while CLIP score samples give qualitatively more descriptive cations.
For language detection, we use Fasttext 0.92, version lid.176, and cld3 - library gcld3 3.0.13. We count the number of words in each caption by splitting using whitespaces.
We use OpenAI pretrained CLIP ViT-B/32 and ViT-L/14 models to compute the cosine similarity text and image tower outputs as the CLIP scores. On the small and medium pools, we also experiment with baselines that filter out samples in the top few percentiles of CLIP scores. Specifically, we try baselines that use samples with top {1,2,5}-30% CLIP scores (ViT-B/32 model), and the performance is sightly better on the small pool (at most 0.5 gain of averaged accuracy) while slightly worse on the medium pool (0.4-0.8 loss of averaged accuracy). In Table 16, we show how the CLIP score thresholds relate to the fraction of the pool retained by the filter.
Each synset is represented by a synset offset that can be used to retrieve the synset from WordNet. In order to verify if a caption has a word corresponding to a synset from our set we iterate over every word and retrieve the synsets that this word can describe (using nltk.corpus WordNet). Following that, we retrieve the most likely lemma representing that synset, find its synset offset, and check if the number is part of the IN21K or IN1K sets.For the ImageNet 21K synsets, we have used the list in https://storage.googleapis.com/bit_models/imagenet21k_wordnet_ids.txt
This baseline uses text only to filter labels which mention concepts (synsets) appearing in IN21K, and applies a temperature parameter to control how equally-represented different concepts are in the dataset. For synset , let be the number of examples containing words matched to that synset, where as before for each word we only match the most likely synset. Furthermore, for image-text pair let be the set of synset matched to the caption.
The probability of sampling example is proportional to either (average synset score in the data point) or (maximum synset score in the data point), where is a “temperature” parameter controlling the flatness of the distribution. We sample examples with replacement but discard any example repeated more than 100 times.
We now provide a detailed description of the Image-based filtering procedure. First, since the core of the procedure concerns only image content, we begin with basic text-bsaed filtering: we remove from the pool only all examples with non-English captions (as determined by fasttext), and all examples whose captions have less than two words or less than six characters.
Next, we use clustering of image embeddings to select a subset of examples whose image content is related to a clean training set of interest. Let denote the CLIP image embeddings of the remaining examples in the pool. We cluster these embeddings into clusters using Faiss with 20 iterations, and let denote the resulting cluster centers. Due to memory constraints, for the large and xlarge pools, we perform the clustering on a random subset of about 160M examples (that pass the basic text-based filtering). For an embedding vector , let
denote the index of the cluster center nearest to as measured by inner product. Let denote the CLIP image embeddings of a clean supervised training set (we experiment with either ImageNet 1K or ImageNet 21K), and let
be the set of cluster indices who are nearest neighbors to some clean training set image. We then keep only images in the pool whose nearest cluster center is in . That is, out of the examples passing the text-based filtering, the output subset keeps the examples with indices
In addition to filtering methods, we experiment with cluster-based sampling methods. First, we compute the score of -th cluster as the number of ImageNet data assigned to this cluster. Then, for parameter we define a distribution over the pool by sampling cluster with probability and uniformly sampling an example for the cluster, rejecting any example repeated more than 100 times. We try 5 different , i.e., , and the best average accuracy is obtained when , while the performance is still worse than the image-based filtering on the small and medium pool. We therefore do not include this line of baselines in the experiments of large pool.
We rank the samples in the pool by the minimum embedding distance (1 minus cosine similarity) between its image and the ImageNet images; both embeddings are obtained from OpenAI pretrained CLIP ViT-L/14 model . Then we select top images by different fractions as in image-based filtering methods.
P.2 BYOD track
We experiment with the following data sources:
CC12M : images and HTML alt-text crawled and filtered from web pages.
YFCC15M: this is the 15M subset of the YFCC100M dataset that Radford et al. used for dataset ablation in their CLIP paper.
RedCaps : 12M images and corresponding captions were crawled from 350 manually curated subreddits between 2008 and 2020.
Shutterstock: 106M images and captions were obtained from the Shutterstock website in 2021 . We use the “photos” subset of this dataset, with 58M samples, which we found performed best, unless specified otherwise.
WIT : Image-text pairs from Wikipedia pages. We use the attribution fields as captions, which we found performed best.
COYO : A collection of 700M image-text pairs from Common Crawl.
LAION-2B : A 2.32 billion english subset of LAION-5B.
LAION-COCO: A dataset with 600M images from LAION-5B and synthetic captions.https://laion.ai/blog/laion-coco/
LAION-A: According to laion.ai, LAION-A is a 900M subset of LAION-2B with the aesthetic filtering procedure used in LAION-aesthetichttps://github.com/LAION-AI/laion-datasets/blob/main/laion-aesthetic.md and pHash deduplication .
In Table 17, we use some heuristics to measure the quality of some external data sources. First, following Nguyen et al. , we train a CLIP model on a 5M random subset from each source, and evaluate the performance of the resulting models on ImageNet and ImageNet-derived distributions — ImageNet-V2 , ImageNet-R , ImageNet-Sketch and ObjectNet . Moreover, for each data source, we use OpenAI’s pretrained CLIP ViT-B/32 and ViT-L/14 models to compute the cosine similarity between image and text embeddings of a data point, and obtain the average cosine similarity score for the whole dataset.
We present a series of results for the BYOD track in Table 18.
Appendix Q Fairness and biases
To study the biases displayed by our models, we include two diversity-related datasets, Dollar Street and GeoDE , in our evaluation suite, and perform further analysis on the face datasets FairFace and UTKFace with demographic labels, following Radford et al. .
We break down model performance on the Dollar Street and GeoDE datasets in Figure 21. Dollar Street consists of images of household items taken in homes around the world, and represents a wide socioeconomic range that includes homes with no Internet access . The objects belong to ImageNet categories, and the task is image classification. Standard ImageNet-trained models achieve monotonically increasing performance levels with higher household income levels . Here we use the income-based subgroups defined in Rojas et al. , and find a similar bias as discovered in their paper. While our trained models show a smaller worst-group performance gap than an ImageNet-trained ResNet-50, they underperform a model fine-tuned on Dollar Street. Models with higher average accuracy show a larger worst-group gap, which future work should try to address.
GeoDE consists of images of everyday items and objects, which again fall into ImageNet categories. The dataset represents six world regions equally, and primarily aims to promote geographic diversity of datasets . Both ImageNet models and our models show less bias under this distribution compared to Dollar Street, with a smaller worst-group accuracy gap. The trends show that performance across all regions improves steadily with increased scale, and the performance approaches that of a model fine-tuned on GeoDE. While we know that classifiers trained specifically on ImageNet can display geographic biases , these biases are not apparent in our GeoDE model evaluations. Future work is needed to investigate the extent to which our models have geographic biases not evaluated in GeoDE.
Q.2 Fairness
Emulating Radford et al. , we evaluate our best models from the filtering and BYOD tracks on the human face datasets FairFace and UTKFace, using zero-shot classification to predict the race, gender, and age annotated in these datasets. Following Hanna et al. and Hundt et al. , we acknowledge that these evaluations can be problematic as race and gender should not be considered fixed categories, but rather fluid attributes that may change for individuals, based on they way they identify at any given moment—regardless of appearance. We include these evaluations for continuity with prior work and as a probe into model behaviour, but hope future work will consider improved face fairness evaluation. We also note that race, gender, and age classification are not the intended end-goals of the models or benchmark, and we do not condone the use of CommonPool or models trained on CommonPool data for any decisions involving people.
As described in Appendix G, our filleting track models are trained on images with faces blurred. Nevertheless, these models still perform significantly above random chance on face classification. We hypothesize that this is due to a combination of faces bypassing our face blurring filter in the training data, contextual clues outside of the face region, or signal associated with skin color. The BYOD track model performs even better than the filtering track model. We hypothesize that this is because BYOD data is used off-the-shelf and hence contains non-blurred faces. In Table 19, we present overall accuracy for these three traits. Note that race is treated as a binary variable (white or non-white) to enable comparison to prior results, gender is a binary variable (male or female) according to annotations, and age is binned into 9 ranges according to the annotation precision of FairFace. The BYOD model, performs better at distinguishing the annotated gender, but is worse at distinguishing annotated race and age.
We further break down these statistics over the intersection of race and gender, examining gender classification accuracies in Table 20. We find that there are drastic differences in accuracy across different annotated subgroups, varying by both race and gender. The filtering models shows a tendency to misclassify Black, Southeast Asian, and East Asian males as females at 20.7%, 17%, and 19.3% respectively on FairFace. Furthermore, we find that while the BYOD model improves accuracy, on FairFace most of this improvement is on men (ranging from 1.7pp gain to 9.9pp gain), while on women, BYOD offers little change (ranging from 0.6pp gain to 6.2pp drop).
Following Radford et al. , we also examined associations of particular demographics with potentially harmful language. We replicate their setup with two classification tasks: (1) including race-gender intersection classes (e.g. “black woman”, “indian man”, etc.) and several harmful crime-related terms (“thief”, “criminal”, “suspicious person”); (2) including the same race-gender intersection classes and non-human terms (“animal”, “gorilla”, “chimpanzee”, “orangutan”). We compute the frequency of misclassification of people into one of the harmful categories and run these experiments on FairFace and UTKFace separately. The results are shown in Table 21. Unlike in Radford et al. , we find that our models have a very small probability of classifying human faces as non-human, with a max score across all subgroups of 0.1%. However, a significant proportion of people are misclassified as criminal. This again highlights the importance of dataset curation and the risks associated with zero-shot classification on models trained on web-scraped datasets.