WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
Conghui He, Zhenjiang Jin, Chao Xu, Jiantao Qiu, Bin Wang, Wei Li, Hang Yan, Jiaqi Wang, Dahua Lin
Introduction
Before the advent of large language models [4; 11; 12; 1; 8; 9; 17] and multimodal models [10; 7; 6; 3], the NLP and CV fields mostly used small amounts of high-quality manually annotated data to train models and conduct research on various specific tasks. Research at this stage focused more on high-quality domain-specific data, improving model performance through enhancing model structures, leading to a large number of excellent network structures. However, with the introduction of large models like Bert and BLIP, pretraining with a large amount of unsupervised internet data can give models good generalization capabilities. Based on large-scale pretraining, using a small amount of SFT fine-tuning and RLHF fine-tuning, some tasks even exceed the average human level. These studies all reflect the importance of large-scale pretraining datasets.
More and more research teams are releasing large-scale datasets, such as C4 , Pile in the NLP field, and LAION400M , LAION-5B , CC3M , CC12M , MMC4 in the multimodal research field. These datasets are far more significant in volume than supervised data, contain relatively affluent information, and effectively promote the development of LLMs and MLLMs.
In this work, we consider the diversity of dataset modalities, content diversity, content safety, and content quality. Based on these basic principles, we have collected, processed, and screened text, image-text, and video data from the Internet. The text data covers multiple fields, including technology, literature, media, education, and law; the image-text data covers a variety of fields, such as news events, people, natural landscapes, and social life; the video data covers military, arts, sports, nature, real world, knowledge, film art, media, food, history, science, and education, etc. These pretraining data have significantly improved the knowledge content, logical reasoning, and generalization abilities of the training model. Specifically, our contributions are as follows:
We build a large-scale training corpus WanJuan that includes multiple modalities: the text data includes more than 600 million documents, with a data storage volume exceeding 1TB; the image-text data is processed into documents, with a total of more than 22 million documents and a data size exceeding 200GB (images are provided via URL links); the video files total over 1000, with a data size exceeding 900GB.
In the construction process of the WanJuan dataset, we ensure data safety, high quality, and value alignment (filtering out pornography, violence, and bias) through algorithmic processing and manual verification.
We provide a unified JSON format processing, dataset download tool, and supporting documentation to facilitate users in quickly applying large model training.
Dataset Statistics
The WanJuan dataset includes multimodal data such as text, image-text, and video, all in Chinese or English, covering numerous fields and boasting high quality. Each modality has been carefully selected and processed to ensure the diversity and comprehensiveness of the dataset. Specifically, the composition of the dataset is as follows:
The text data comes from different sources such as web pages, encyclopedias, books, patents, textbooks, and exam questions. Through finely designed rules and algorithms, we filter and process the original data, remove invalid content, and ensure the content’s safety and high information content. The specific composition of the text data is shown in Table 1, which includes more than 600M documents with a storage volume exceeding 1TB. Figure 1 provides some examples of the text data. This extensive collection of text data provides a rich resource for training language models and studying various NLP tasks.
2 Interleaved Image-Text Data
The image-text multimodal data originates from public web pages. After processing, it forms interleaved image-text documents. The specific composition is shown in Table 2, which includes over 22 million documents and a data size exceeding 200GB (excluding images). Figure 2 presents some examples from the interleaved image-text dataset. The interleaved image-text data provides a valuable resource for studying the interaction between text and images, which is crucial for many multimodal tasks.
3 Video Data
The video data is sourced from the high quality of the China Media Group and Shanghai Media Group program footage. It encompasses over 1000 videos, with a data size surpassing 900GB. Figure 3 illustrates examples of the video data. The inclusion of video data allows for the exploration of tasks that require understanding and generating content across different modalities, such as video captioning and video question answering.
Methods
The video data, sourced from the China Media Group (CMG) and Shanghai Media Group (SMG), is of high quality. Thus, we mainly focused on cleaning text and text-image data, and this section introduces the data collection, cleaning, and value alignment process.
To ensure the comprehensiveness of the data, our text corpus combines data from eight different sources, as fully outlined in Table 2. For the English internet data provided by the Common Crawl, we employed a multi-step text extraction process, language detection, corpus filtering, and deduplication to obtain high-quality data.
Firstly, we extracted text from the original WARC files, then used different language detection tools (pycld2) to classify the extracted text, subsequently processing the Chinese and English texts differently. Given the abundance of invalid data on the internet, we then applied the following rules to filter and obtain high-quality data:
We removed irregular documents, including those with inappropriate average word and document lengths. If the most frequent word was non-alphabetic, or the frequency was too high, we considered it an uncommon document format and deleted it.
We removed documents with too little content, such as those with fewer than three sentences after processing; fewer than three paragraphs; fewer than three paragraphs longer than 200 words; or fewer than two stopwords.
We also cleaned paragraphs, removing special sections such as those containing words like JavaScript, sections outside of punctuation marks, and paragraphs with more than 1000 words. Even a small amount of duplicate data can significantly impact our observations. We noticed that the obtained data contained duplicates, so we tokenized the text data and used MinHashLSH and n-grams to evaluate similarity, deleting content with a similarity greater than 0.8.
Furthermore, due to the presence of harmful and low-quality content in internet data, we trained some models to assess quality and filter for different issues:
We trained content safety models for pornography, violence, gambling, attacks, and other toxic themes using the FastText model, separately for Chinese and English, to filter out potentially toxic data.
We trained data quality models for various low-quality data found online, such as auto-generated random data and advertising content, separately for Chinese and English, to reduce the proportion of low-quality data.
Based on the above filtering, we obtained safe, high-quality, value-aligned text data.
2 Image-Text Data Cleaning
The interleaved image-text data comes from four sources, as detailed in Table 2. Since the text-image data come from official sources, they are of high quality. Here, we selectively extracted the needed content to form the interleaved image-text data. For the formation of interleaved image-text data, we followed these steps:
To reduce the difficulty of cleaning and ensure data quality, we wrote specific parsing rules for each site. User-generated articles were sourced from a single open-source site.
We only extracted valid (ad-free, list-free, navigation bar-free, emoticon-free, comment-free) article content. We used a series of rules to filter. We also removed references, complex tables, lists, and other entry-related content for the text part of Wikipedia, retaining only the text paragraphs. For user-generated articles and authoritative media news, we used XPath, CSS selectors, and regular expressions to remove media sources, publishers, reposts, advertisements, and comments unrelated to the article theme, to obtain the article’s main body. Like with text data cleaning, we also performed deduplication based on similarity.
We believed the header images of Wikipedia articles are meaningful for image selection, so we only retained these. To ensure all images in user-generated articles had valid descriptions, we removed articles with more than 15 images and those where the number of text characters was less than twice the number of images. Based on this rule, we retained 55% of the valid articles.
We the valid text and images obtained from the above filtering. The format for Wikipedia was (the first paragraph, main image, and remaining paragraphs).
After applying these filters, we obtained high-quality interleaved text-image data. The language distribution was 62.3% Chinese and 37.7% English.
Conclusion
In this paper, we introduced the WanJuan dataset, a large-scale, multimodal Chinese-English dataset collected from diverse web sources. The dataset includes various modalities, such as text, image-text, and video, all meticulously processed to ensure safety, richness of content, and accuracy. This high-quality dataset provides a valuable resource for the training of large language models and the study of various multimodal tasks. We believe that the release of this dataset will significantly contribute to the advancement of research in the fields of Natural Language Processing and Computer Vision, especially for tasks that require understanding and generating content across different modalities.