Yuan 1.0: Large-Scale Pre-trained Language Model in Zero-Shot and Few-Shot Learning
Shaohua Wu, Xudong Zhao, Tong Yu, Rongguo Zhang, Chong Shen, Hongli Liu, Feng Li, Hong Zhu, Jiangang Luo, Liang Xu, Xuanwei Zhang
Introduction
The Transformer architecture has been widely used in natural language processing. In order to improve the performance, a varieties of Transformer-based modifications have been proposed since 2017, but many of them exhibit a lack of generalization across different implementations and tasks. Kaplan, et al. confirms that performance of the Transformer steadily improves with the scaling up of model size, dataset size, and the amount of computation for training. Roberta shows that the accuracy of BERT can be substantially improved by training the model for a longer time with a larger corpus. The T5 model, built with vanilla Transformer structure and increased model size with 11 billion parameters, achieves the state-of-the-art (SOTA) performance in various NLP tasks. It is proved that larger language models performs better than smaller ones. GPT-3 with 175 billion parameters, as a milestone, was proposed in 2020. Before GPT-3, it was common to pre-train a model with unsupervised learning on a large unlabeled dataset, then fine-tune on a specific task. Because GPT-3 makes great progress on Zero-Shot and Few-Shot learning, it can be applied directly on a wide range of NLP tasks, and displays good performance without being fine-tuned on those tasks. After GPT-3, several studies further increase the model size in two ways:
Singleton: Increase the number of layers and the size of a layer, such as GPT-3 and PanGu-.
Mixture of Experts (MoE): Scaling the model size with Sparsely Gated Mixture-of-Experts (MoE), such as GShard, Switch Transformer, Wudao and M6 . Each expert is a singleton model in a size up to 10B. With MoE, the model size can be successfully scaled up to more than 1000B .
Both Singleton and MoE are effective to increase the model size, however, they behave differently in Zero-Shot and Few-Shot scenarios. Currently, the MoE method still follows the common way, in which pre-train the model on a large dataset and fine-tune it on specific task. To our best knowledge, no MoE model is applied on Zero-Shot or Few-Shot learning. However, both GPT-3 and PanGu- with singleton architecture, exhibits good performance on Zero-Shot and Few-Shot learning. Training a model with parameters greater than 100B requires huge amount of computational resources. Take GPT-3 175B for example, it was trained on a cluster of 10,000 GPUs . Such a huge requirement on computational resources makes it difficult for most researchers to train a model in a similar way. In this work, we propose Yuan 1.0 singleton model with 245B parameters. To accelerate the training process of Yuan 1.0, and thus reduce energy costs and carbon emissions, we make a collaborative design of model architecture and large-scale distributed training. The main contributions of our work are summarized as below,
A method that incorporates large-scale distributed training performance into model architecture design is proposed. With this method, we trained our Yuan 1.0, the current largest singleton language model with 245B parameters, and achieved excellent performance on thousands GPUs
A data processing system is created to efficiently filter a massive amount of data from Internet. The current largest Chinese corpus with 5TB high-quality text is built based on this system.
The model architecture with better performance in Pre-train and Fine-tune pattern is likely to behave opposite in Zero-Shot and Few-Shot learning.
A method that can steadily improve the Zero-Shot and Few-Shot performance is proposed.
Yuan 1.0
The basic architecture of Yuan 1.0 is a language model. For a given input , the language model predicts the probability of the output :
To deal with different downstream tasks (translation, question answering, classification etc.), each task is casted in a diagram of text-to-text framework. In this way, a pre-trained language model can be directly applied to handle different tasks.
In this work, we consider two model architectures, Language Model (LM) and Prefix Language Model (PLM). In LM, which is one of the most commonly used architecture, the decoder in a Transformer is taken to auto-regressively generate an output sequence. The structure of a decoder LM is presented in Fig. 1(a). At time t, a token at the rightmost of the output, , is generated based on the probability predicted by the model. Then this token is concatenated to the input sequence and fed back into the model to generate the next token at time t+1. In the pre-train and fine-tune pattern, a LM performs better in natural language generation (NLG) tasks, but comparatively worse in natural language understanding (NLU) tasks. The reason of this drawback is that a LM, with casual masking of attention structure, forces the model’s prediction of the token only depends on the tokens before i, which can be seen in Fig. 1(a). In contrast, PLM performs well in both NLU tasks and NLG tasks. Instead of taking a casual mask, PLM uses a fully-visible mask in the range of the prefix portion of input sequence, which can be seen in Fig. 1(b).
2 Cooperative design of model structure
The huge cost of computational resources is the bottlenecks that limits researchers to develop NLP models with hundreds of billions parameters. Take GPT-3 for example, it was trained on a large cluster with 10,000 GPUs. In order to accelerate the training process, we incorporate the key factors that affect the performance of large-scale distributed training into the design of Yuan 1.0 model structure. The parameters of LM that affects both the accuracy and the performance of large-scale distributed training include the number of layers, hidden size, global batch size, micro batch size, etc. In the large-scale distributed training of Yuan models, we use three-dimensional parallel strategies, including tensor parallelism, pipeline parallelism and data parallelism. In this section, we make theoretically analysis of these parameters, and demonstrate how to choose the model parameters under different parallel strategies. The notations used are presented in Table 1.
In tensor parallelism, the layers of a model are partitioned among devices within a node. The schematic of tensor parallelism is presented in Figure 2. In Transformer, tensors of Attention and MultiLayer perceptron (MLP) are split by row or column during forward and backward computing. Input tensor is broadcasted to each accelerator, in which makes forward propagation. When the forward pass of Attention or MLP is finished, an all-reduce is performed. Then the results are updated on all devices and sent to the next Layer. There are four all-reduce operations in the forward and backward propagation per layer. The ratio of the computation time to data communication time per layer is,
According to Eq. 2, of the tensor parallelism increases with h and S. The value of S is usually chosen as 512, or 1024 . Because memory requirements of Attention is quadric to S, sparse Attention structure is necessary if S is increased to 2048. In order to save memory of accelerators and enable larger S, the activations are recomputed in backward propagation. With this method, the model can be trained with smoothly with normal Attention structure.
2.2 Pipeline Parallelism
For language models with hundreds of billions parameters, the parameters can hardly be stored in a single node. Pipeline parallelism spliting the layers of LM among multiple nodes, is applied to solve the above mentioned problem (Fig. 3). Each node is one stage in the pipeline, which receives outputs from the previous stage and sends results to the next one. A node will be idle if the inputs received from its previous neighbor is not ready. The idle time for a pipeline is called pipeline bubble. To increase the performance of pipeline parallelism, we have to decrease the time spent on pipeline bubble. The fraction of ideal time spent in the pipeline bubble is,
According to Eq. 3, the time spent on pipeline bubble increases with the number of layers L, and decreases with the number of micro-batch size . There will be a better performance if . In pipeline parallelism, the ratio of the computation time to data communication time per node is,
According to Eq. 4, the computational efficiency of a pipeline node improves with the increase of the values of h and S, which is similar to the situation of tensor parallelism.
2.3 Data Parallelism
The global batch size is split among pipeline groups by data parallelism (Fig. 4). Each pipeline group with a copy of the model is fed by local batches. In data parallelism, the ratio of computing time to communication time is,
Because d is often far greater than 1, Eq. 5 can be simplified to,
The computing efficiency improves with the increase of the global batch size B and sequence length S. Because the memory requirements is quadric to sequence length S, increasing the global batch size seems to be a more effective way. However, there will be numerical instabilities during training when global batch size is too large . To avoid numerical divergence, the global batch size is kept smaller than tokens.
2.4 The principles of model parameters selections
In summary, we follow the rules below to select model parameters,
Increase the sequence length as much as possible, as it benefits the tensor parallelism, pipeline parallelism, and data parallelism. Because the memory requirement is quadric to the sequence length, it is worthy to re-compute activations in the backward propagation to save memories.
Too many layers in language model have negative effect in performance, because it increases the time spent on pipeline bubble.
Increasing the hidden size improves the performance of both tensor parallelism and pipeline parallelism.
Increasing the number of micro batches in a node improves the performance of pipeline parallelism. Increasing the global batch size improves the performance of data parallelism.
Three models (Yuan LM-13B, Yuan 13B-PLM and Yuan 245B) are trained with parameters presented in Table 2.
Dataset
A Chinese corpus with 5TB high-quality text is built, which is sufficient to train Yuan 245B model without sampling the dataset twice. To our best knowledge, this is the largest Chinese text dataset compared with CLUECorpus2020 (100GB) , PanGu Corpus (1.1TB) , WuDaoCorpus2.0 (2.3TB Chinese text data and 300GB English text data) , and ERNIE 3.0 (4TB) .
In order to obtain the high-quality dataset, we develop a Massive Data Filtering System (MDFS) built on Spark to clean and filter the raw data, and train a Bert-based model to select high quality samples. MDFS is consisted of three parts, data collection, coarse filtering and fine filtering (Fig. 5). The raw data is collected from Common Crawl, Sogou News, SogouT, Encyclopedia, and Books (Table 3). To process these raw data, we run MDFS system on a high performance cluster with 36 nodes.
The coarse filtering includes the following modules:
Article Extraction: Extracts contents of articles from crawled web pages, and removes field like WARC headers, hyperlinks, etc. The following rules are applied,
A paragraph started with WARC keyword is discarded.
A paragraph without a valid punctuation at the end is discarded.
A paragraph that contains neither English character nor Chinese character is discarded.
Empty Article Filtering: Removes empty articles.
Chinese Text Filtering: Selects Chinese texts with the following rules,
An article with less than 30 Chinese characters is removed.
An article is removed if the percentage of Chinese characters is less than 60.
After this step, the data size of Common Crawl is decreased from 866,304GB to 12,200GB.
Traditional to Simplified Chinese: Converts traditional Chinese characters to simplified Chinese characters.
Sensitive Words Filtering: Removes articles or paragraphs that include sensitive words. 9,759 sensitive words are collected and classified into black and blue categories.
An article with words in black category is removed.
A paragraph that contains words in blue category is removed, and other paragraphs in the same article is kept.
Symbol filtering: Removes junk symbols, such as invisible Unicode characters, invisible ASCII characters and specific punctuations.
2 Fine Filtering
To extract high quality articles based on course filtering text, we train a Bert-based model to classify high quality, low quality and advertisements. A datasets labeled with high quality articles, low quality articles, and advertisements is built to train this model. About 2TB data is removed by the model, and 50% of the removed data are identified as advertisements. Considering that the advertisements may also contain complete semantic information, we evaluate them manually to determine whether it is necessary to recall. The processed dataset scatters on 36 nodes. To avoid additional bias from human, we sample two sets on each node, and each set is evaluated by different reviewers. Parts of the statistical results are shown in Table 4.
The similar percentage of high quality data for Sample1 and Sample2 indicates a high consistency in data evaluation. As the percentage of high-quality data in advertisements is fairly low, it is reasonable to discard all advertisements. In the manual review, we find a 2.4% duplication rate in high-quality contents, while the duplication rate is 12.6% in advertisements. De-duplication is further applied to the high quality data. The data size after fine filtering is shown in Table 5. The total size of high-quality dataset is 5.02TB. During training, only articles with more than 150 characters are sampled.
Experiments and Results
The Yuan models are trained on a cluster with 2128 GPUs. A stable real performance of 45% of the theoretical peak performance is achieved on this cluster. Adam optimizer is used for training Yuan models. Please refer to Table 6 for more details. To stabilize the training process, a linear warm up of learning rate is taken over the first 1% tokens, then the learning rate follows a cosine curve that slowly decays to 10% of its original value. The global batch size also linearly increases to the full value over the first 2% tokens, then it is kept till the end of training. During the training we pack multiple documents into a single sequence of size 2048, and the documents are separated with a special token ”¡eod¿”. The tokens in the sequence are not masked in any way.
The models are mainly evaluated on FewCLUE and ZeroCLUE, which can be classified into 4 categories, including text classification, Winograd Schema, natural language inference, and reading comprehension. Text Classification is consisted of sentimental classification (Eprstmt: E-commerce Product Review Dataset for Sentiment Analysis), news title classification (Tnews: Toutiao Short Text Classification for News), app description classification (Iflytek: Long Text classification), and subject classification (Csldcp: Chinese scientific literature subjects classification). Eprstmt is a binary classification with positive and negative product reviews. Tnews, Iflytek and Csldcp are multi-class classification with 15, 118 and 67 categories respectively. On tasks with labels as 0 or 1 or in English, we assign each label with a semantically Chinese meaningful name. For labels longer than one token, we convert those into one-token labels with the same meaning. For all text classification tasks, label is appended to the end of a sentence, connected with prompt words. Our generative model predicts the label based on a given sentence, and calculate the probability P(label—sentence) of each candidate. The candidate with the largest probability is selected as the prediction. Winograd Schema task (Wsc) is a disambiguation task determining which noun a pronoun refers to. It is treated as a binary classification task in our evaluation. Natural Language Inference (NLI) includes Ocnli and Bustm, which concerns the ability to understand the relation between two sentences. Both of the tasks provide two sentences and a label of 0 or 1. We calculate the cross entropy loss of the second sentence with a candidate label, and treat the label with the lowest loss as the prediction. Reading Comprehension includes Chid and Csl. Chid is a Chinese Idiom cloze test dataset. Each sentence has a blank inside, and for each blank, there are 7 candidate idioms with 1 true choice. For this task, each candidate is filled in the blank, and we calculate the cross entropy loss of each combination. The one with the lowest loss is the predicted true idiom. Csl can be treated as either a reading comprehension or a binary classification task. An abstract is provided along with 4 keywords in the dataset. If all keywords are consistent with the abstract, the label should be true or 1. Otherwise, the label should be false or 0. All keywords are appended to the end of abstract. We calculate the cross entropy loss of the part after ABSTRACT in condition of ABSTRACT. The one with smaller loss is the predictive result. Gereration tasks: In addition to FewCLUE and ZeroCLUE, there are two generation tasks, CMRC2018 and WebQA. CMRC2018 is a span text extraction task, consisted of articles followed by several questions, and the answer to a question is a segment of the corresponding article. WebQA is a closed book question answering task. Model is evaluated by directly answering questions of CMRC2018 and WebQA without any auxiliary information. The generated answer is evaluated with EM and F1 scores.
2 Comparison of LM and Prefix LM
Yuan LM-13B and Yuan PLM-13B are evaluated on FewCLUE and ZeroCLUE (Table 7). The SOTA results of ZeroCLUE is benchmarked with zero-shot, while that of FewCLUE are benchmarked with fine-tune. The Zero-Shot results are in-context learning without tuning parameters.
Table 7(a) indicates both LM and PLM have convincing in-context learning capability. The zero-shot average scores of both LM and PLM are superior to the SOTA one. On Csldcp, Tnews and Iflytek tasks, we surpass the zero-shot SOTA by a large margin. Our models also achieve strong performance on Ocnli, which is 6-8 points larger than the zero-shot SOTA. Table 7(a) displays the results after calibration and label expansion, and the methods in details will be discussed in the next section. Our supervised fine-tuning method is aligned with the design of GPT. The average scores for LM and PLM are comparable to the SOTA ones (Table 7(b)). Compared to the few-shot learning results, fine-tune makes great improvement to Bustm, Csl and Wsc. However, for Chid, Eprstmt, Tnews and Ocnli, which are strong in zero-shot, fine-tune contributes little or even have negative effect. The fine-tuned accuracy is superior to the SOTA fine-tune results on 7 tasks including Bustm, Chid, Csl, Csldcp, Eprstmt, Iflytek and Wsc.
We submitted the PLM on FewCLUE, and LM on ZeroCLUE. Both of them currently topped on the list (Table 8). Comparing the results of LM and PLM, we note that LM performs better on Zero-Shot learning, while PLM outperforms with fine-tune. Fine-tune in general brings better accuracy in most tasks. However, fine-tune costs tremendous computational resources for Yuan 245B model, which makes fine-tune uneconomic. Accordingly, we choose LM as basic architecture of Yuan 245B model.
3 Results of Yuan 245B model
Fig. 6 presents the training loss curves of Yuan 245B model. The loss decreases rapidly at the first 10B tokens, and gets flatten over a long tail. Table 9 shows the comparison of GPT-3, PanGu- and Yuan 245B in training. The PetaFlops-day of PanGu- and Yuan is computed as,
Activations are recomputed during backward propagation. The computing amount of Yuan 245B is much greater than that of PanGu-. The training loss of Yuan 245B is the smallest among these three models. Table 10 compares the generation results between Yuan and recently published Chinese pretrained language models, Pangu- and Ernie 3.0. The average scores of Yuan outperform Pangu- and Ernie 3.0 by a large margin on Close-book QA and Span Extraction reading comprehension, which proves the excellent zero-shot generation capacity of Yuan 245B. Regarding WebQA, Yuan significantly improves the performance, no matter evaluated with EM or F1 Score. For CMRC2018, Yuan also achieved a better averaged score and F1 score compared to the SOTA, and it is little worse on EM compared to the SOTA.
A noticeable shortcoming of in-context learning lies in its bias towards template sentences and labels. The bias mainly comes from a dataset with imbalance distribution between classes, few-shot examples with a certain order, and labels with different frequencies in the training corpus. The extraneous bias limits model’s performance on natural language tasks. Considering the probable sources of bias, we take calibration for in-context learning in two aspects: a calibration on the calculation of probability, and an expansion of labels. Based on previous work on calibration, the model’s bias can be fixed with an empty text. We use a similar method to calibrate our in-text prediction.
Take Tnews and Ocnli as examples. In the case of Tnews, we calculate the probability of the last token in the sentence-label combination, which is actually a prediction of label. Orig:新闻:sentence。这条新闻是关于label。 Void:新闻:N/A。这条新闻是关于label。 With calibration, we calculate,
In the case of OCNLI, which is a two-sentence task, we calculate the cross entropy loss of the second sentence. Orig: sentence1? 对/错/可能,sentence2. Void:N/A?对/错/可能,sentence2. With calibration, we calculate:
In the experiment, we find that the difference in label frequency are influential to the prediction. The ideal condition is that all labels have approximately the same frequency, but it is too tough for manual selection. We choose an Embedding Corpus that covers over 8 million Chinese words and phrases, to assist us reveal the relevance between words. Each label is expanded into 5 synonyms, in order to reduce the bias led by a single label word or phrase. Three tasks are selected to display the effect of calibration, and the results are presented in Table 11. The combination of model calibration and label expansion leads a dramatic improvement for Eprstmt, which is a binary sentimental analysis. For multi classification on news (Tnews) and scientific literature (Csldcp), calibration also brings better results to a large extent.
Figure 7 presents the Zero-Shot and Few-Shot results of Yuan 245B on ZeroCLUE tasks. In Few-Shot, the number of samples is determined by the number of classes. We take 4 samples for binary classification, and 3 samples otherwise. Compared with zero-shot, we note that few-shot learning brings steady improvements on accuracy for most tasks, except Csldcp and Iflytek. Few-shot leads to a disastrous decrease for Iflytek with a near-random result. With the same method, few-shot has positive effect for Eprstmt (2 classes) and Tnews (15 classes). We observe similar situations in Yuan LM-13B, and Yuan PLM-13B. Because there are large number of classes in Csldcp (67 classes) and Iflytek (118 classes), and the samples cannot cover all the classes, the bias caused by samples concatenated in the input make the model predication getting worse.
3.2 Text generation of Yuan 245B model
In order to see how well Yuan can generated text, we arbitrarily select 24 articles that Yuan 1.0 generates, including 4 couplets, 5 traditional and modern Chinese poetries, 5 news articles, 5 stories and 5 dialogues. Couplet, poetry and dialogue can be seen as short-text task ( 10-20 tokens), while news and story generation can be seen as long-text task ( 300 tokens). In comparison, the human-written articles are masterpieces from Chinese poems, pieces of classic novels, news articles from Sohu News, and dialogues from LCCC-large dataset. Participants are asked to select whether the article was ”written by a human” or ”written by a model”. We collect 83 valid questionnaire. According to our interview, most interviewee will choose ”the better one” as the article created by human.
The human accuracy at detecting articles created by Yuan 1.0 is 49.16%, which implies the difficulty for participants to distinguish human-written and model-written articles, especially in regard to modern Chinese poetries and articles (Fig. 8). The generation of news (42.12,%) and stories (49.15%) convinces us with excellent long-text generation capacity. Some of model-written articles are even better than parts of masterpieces in the view of our participants. The generation of couplets and poetries indicate that Yuan 1.0 is able to create text with rules and forms of ancient Chinese (Table 12) although ancient Chinese is not strengthened in our pretraining corpus. Yuan can also make a dialogue aligned with human’s expectation (45.68%). Yuan is able to generate articles, such as news, and stories, which is hard to tell whether the article is human-written or model-written. However, you will see repetition to some extent, if the model is required to create an article with more than 1000 Chinese characters. Few-shot is also effective for text-generation, especially text with a certain format. Table 13 shows the couplets Yuan created with zero-shot, one-shot and three-shot, under the same set of hyper-parameters. Few-shot mainly contributes to: (a) increasing the stability and completeness of generation (No. 3); (b) more reasonable semantic meaning and accordant written style (No. 1); (c) avoidance of the word repetition that is a taboo for couplet creation (No. 2, 5); and (d) antithetically better (No. 4).
The ability of imitation of Yuan is evaluated via learning and utilizing a brand new word. A definition and an example sentence are given in the input, and the model will write a new sentence with the given information. The nonexistent words include nouns and adjectives. Table 14 displays the one-shot examples. In all cases, the model makes approximately correct applications with the nonexistent Chinese words, which implies the learning and imitation ability of our model. This ability is especially effective when model aids scientific article writing, as tremendous definition in academic articles could be alien to Yuan. In few-shot and text generation experiments, we note the pre-trained language model’s sensitivity to steering, which could be a potential risk on its application. Given an input without bias, the opinion created by the model could be either positive or negative. However, if given an input with a strong standpoint, the model tends to continue the article with the same opinion. In the case displayed in Table 15, we would like to talk about the women status in the society, and two contrasting options are given as the input. Given ”traditional patriarchy still holds sway”, Yuan express the opinion as ”women status is determined by their fertility”. On the opposite, given ”40% of the labor force are women. There are more readable girls and women than ever”, the model also follows the opinion as ”women have the competence to any work”. It is proved that a model with no certain bias can be easily steered by human. As the capability of generating ”human-written-style” article strengthen the risk in model abuse, the application of model needs to be regulated.
Conclusion
We proposed the current largest singleton language model Yuan 1.0 with 245B parameters that achieved good performance on different NLP tasks in Zero-Shot and Few-Shot learning. The architecture of Yuan 1.0 was designed by incorporating model structure with key factors that affects performance of large-scale distributed training. The training process achieved excellent performance on 2128 GPUs. Yuan 1.0 was trained on a new Chinese dataset of 5TB high-quality text that was built on 850TB raw data from Internet. Zero-Shot and Few-Shot performance was steadily improved by calibration and label expansion. We found that the prefix language model that performs better in Pre-train and Fine-tune pattern behaves differently in Zero-Shot learning, while language model behaved the opposite. Yuan 1.0 models achieved the state-of-the-art results on ZeroCLUE, FewCLUE, and generation tasks. The articles generated by Yuan 1.0 are difficult to distinguish from those written by humans.