PanGu-Σ: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing

Xiaozhe Ren, Pingyi Zhou, Xinfan Meng, Xinjing Huang, Yadao Wang, Weichao Wang, Pengfei Li, Xiaoda Zhang, Alexander Podolskiy, Grigory Arshinov, Andrey Bout, Irina Piontkovskaya, Jiansheng Wei, Xin Jiang, Teng Su, Qun Liu, Jun Yao

Introduction

Large Language Models (LLMs) [2, 3, 1, 4, 5, 6, 7, 8, 9, 10, etc.] have demonstrated unprecedented capabilities and potential in the areas of natural language understanding, generation and reasoning. By utilizing vast amount of textual data, the performance of language models scales up with compute budget and model parameters, demonstrating strong zero/few-shot learning abilities or even emergence abilities . Several large language models with hundreds of billion parameters have been published since GPT-3 , including but not limited to Megatron-Turing NLG , PanGu-α\alpha , ERNIE 3.0 Titan , Gopher , PaLM , OPT , Bloom , and GLM-130B . Researchers start to build even larger language models with more than one trillion parameters. Typically, this is accomplished by leveraging sparsely-activated models such as Mixture-of-Experts (MoE) . Among the trillion-parameter models currently in existence, there are several noteworthy work such as Switch-C , GLaM , MoE-1.1T , Wu Dao 2.0 , and M6-10T . However, only a select few have published comprehensive evaluation results over a wide range of tasks while simultaneously achieving anticipated performance. In our experience, the primary difficulty lies in the scaling efficiency.

Recent studies on the scaling laws of language models demonstrate the necessity of training LLMs with sufficient amount of training data and corresponding compute budget to achieve optimal performance. Therefore, one of the main motivation for this work is to design a scalable model architecture and an efficient distributed training system that can consume the data with high training throughput.

Model Scaling. Model performance of LLMs is expected to scale up with larger model size. Comparing to the expensive computational cost for training dense Transformer model, sparse architectures such as Mixture-of-Experts (MoE) are considered to be an appealing choice to scale model size up without incuring linear increase in computational cost. However, MoE models suffer from the problems such as unbalanced workload and all-to-all communication latency. Moreover, how to extend existing dense model with MoE and how many experts to allocate in each layer remain open problems. Therefore, designing a trillion parameter sparse model with high performance and training efficiency is a significant yet challenging task.

System Scaling. Frameworks such as DeepSpeed https://www.deepspeed.ai/ have been proposed to support training trillion parameter models. In practice, the main barrier often lies on limited compute budget, or more specifically the number of accelerating devices (e.g., GPU, NPU, TPU) that can be used. By utilizing techniques such as tensor parallelism , pipeline parallelism , zero redundancy optimizer and rematerialization , practitioners can train trillion-parameter model with feasible batch sizes across thousands of accelerating devices. Alternatively, practitioners can reduce the amount of computation resources by utilizing heterogeneous computing techniques such as offloading some of the computation to host devices . However, the current techniques inevitably hinder the training throughput due to slow bandwidth between the host and device as well as weak computing capabilities of CPUs compared to accelerating devices, which prevent feeding large language models with reasonably amount of data and achieving optimal performance. Therefore, how to efficiently scale the system performance with limited computation budget is critical to the performance of large language models.

In this work, we present PanGu-Σ\Sigma , a large language model with sparse architecture containing 1.085 trillion parameters. We develop PanGu-Σ\Sigma model under the framework of MindSpore https://gitee.com/mindspore/mindspore and train it on a cluster with only 512 Ascend 910 AI Accelerators with 329 billion tokens over 100 days. PanGu-Σ\Sigma inherent parameters from PanGu-α\alpha with Transformer decoder architecture and are extended via Random Routed Experts (RRE). Different from conventional MoE, RRE adopts two-level routing. At the first level experts are grouped by domain or task, and at the second level tokens are randomly and uniformly mapped to experts in each group without using any learnable gating function as in MoE. With the design of RRE, one can easily extract sub-models from the PanGu-Σ\Sigma for various downstream applications such as dialogue, translation, code generation or general nature language understanding. To make training system efficient and scalable, we propose Expert Computation and Storage Separation (ECSS) mechanism, which achieves 69905 tokens/s observed throughput in training 1.085 trillion PanGu-Σ\Sigma on cluster of 512 Ascend 910 accelerators, and reduces Host-to-Device and Device-to-Host communication as well as optimizer update computation by a large margin. As a whole, the training throughput is improved by 6.3x compared to the model of the same hyper-parameters but with MoE architecture. By consuming 329B tokens in more than 40 natural and programming languages, the sub-modal of PanGu-Σ\Sigma in Chinese domain significantly outperforms the previous SOTA models including PanGu-α\alpha with 13B parameters and ERNIE 3.0 Titan with 260B parameters over 16 downstream tasks in six categories in the zero-shot setting without any multitask finetuning or instruction tuning. We also test the performance of fine-tuned PanGu-Σ\Sigma on several applications domain such as dialogue, machine translation and code generation. PanGu-Σ\Sigma outperforms the SOTA models in the corresponding areas.

The rest of the technical report is organized as follows. Section 2 introduces the design philosophy and the architecture of PanGu-Σ\Sigma model. Section 3 introduces the collection and organization of the dataset. Section 4 describes system design and acceleration techniques. Section 5 presents the experimental results of PanGu-Σ\Sigma model.

Model

PanGu-Σ\Sigma aims to achieve the following goals.

Performance: state-of-the-art NLP performance across multiple domains and tasks.

Efficiency: training trillion parameters model with maximum system performance on a modest cluster.

Usability: extendable to various domains or tasks, without need of retraining the model from scratch.

Deployment: easily customizable and deployable in various real-world settings.

Achieving all the above goals at the same time is very challenging. Considering the first goal, a language model that can generalize and perform well across domains should have a very large number of parameters and be trained on large amount of data according to the scaling law . However, training such a large model also means that a high-end cluster is mandatory, which somehow contradicts with the second goal. And the larger scale of the model also leads to increasing cost in deploying the trained model, which is related to the fourth goal.

Considering the high computational cost incurring during the training phase, we want the resulted model to be practically usable and efficient in many real applications. With this goal in mind, we propose to train the model in multiple domains and make it further extendable to any number of domains in a continuous learning paradigm, subject to the computation resource.

During training phase, the trillion parameters PanGu-Σ\Sigma model is fed with data from multiple domains. However, in the deployment phase, it is often unnecessary or even impossible to host the trillion parameters model for every application. Therefore, a model that allows for the grouping and separation of its parameters based on various training and deployment setups offers significant advantages.

2 PanGu-ΣΣ\Sigma Architecture

PanGu-Σ\Sigma adopts an auto-regressive language modeling with stacked transformer decoder layers and a query layer on the top. The PanGu-Σ\Sigma architecture offers a flexible design. The bottom M layers are globally shared across all the domains, and the top N layers (including the query layer) are sparsely activated according to the domains of the input data. In each RRE layers, there are K experts in G groups in total, the number of experts in each group can be different. This flexible design offers three mode.

Mixed mode: when M>0M>0, N>0N>0 and K>0K>0, model contains both sparse RRE layers and dense layers.

Dense mode: when N=0N=0 or K=1K=1, the architecture will reduce to a dense PanGu-α\alpha model.

Sparse mode: when M=0M=0 and K>1K>1, the architecture will be a sparse model.

In this trillion-parameters modeling practice, We use the mixed configuration by placing the shared parameters close to the input layer (bottom) and all the sparsely activated expert parameters close to the output layer (top). In the model designing stage, we benchmark various experts placement strategies on smaller scale models and the selected strategy obtains the lowest language modeling perplexity. Our hypothesis is that bottom layers tends to learn general knowledge, while the specific knowledge is in a higher level of abstraction and is more appropriate to be learned by the top layers. In the token embedding layer, we choose to use different embedding matrices for different domains.

2.2 Random Routed Experts

In the top N layers, we replace each feed-forward sub-layer with multiple conditionally activated feed-forward sub-layers (experts), following the Mixture of Experts (MoE) paradigm.

A key question in designing MoE architecture is how to route tokens to experts. For PanGu-Σ\Sigma , we propose a Random Routed Experts (RRE) mechanism, which is inspired by Hash Layers proposed in . Specifically, RRE routes the tokens by IDs in a two-level manner. In the first level, the token is mapped to a group of candidate experts by domain, and then in the second level, one expert in this group is chosen according to a token-expert routing map to process the token. The routing map is randomly initialized and each layer has a independently initialized mapping for balancing the computation.

RRE has several advantages over the commonly-used learnable routers.

During training, PanGu-Σ\Sigma allows for the addition, modification, or removal of domain-specific experts without any impact on the other experts. This attribute makes PanGu-Σ\Sigma highly flexible for alleviating the commonly encountered problem of catastrophic forgetting, which is crucial for life-long or continual learning.

In most real-world deployment setting, it is unnecessary or impractical to deploy a trillion-parameter model. PanGu-Σ\Sigma allows one to extract a sub-model for specific domains according to practical requirements and only deploy the sub-model. The sub-model may contain tens of billion parameters but still keep the predictive power of the original model on the target domains. Using this extract-and-deploy operation, we can easily deploy models for multiple industrial applications.

All the conventional MoE models rely on all-to-all communication collective operation to move data between experts residing on different devices. With our proposed two-level routing, experts from different domains don’t exchange tokens, and all-to-all communication is constrained within each domain. As a result, the expensive global all-to-all operation is reduced to grouped all-to-all, saving much communication volume and reducing the end-to-end training latency.

Learnable router needs more computation, and can suffer from problem of unbalanced loads across the experts, which typically makes the training process less stable. RRE avoids all the above pitfalls since no additional parameters are introduced and randomly initialized routing table helps to balance the loads on experts.

RRE requires a routing map which is initialized before pretraining, Algorithm 1 describes how we construct the routing table.

Dataset

To better demonstrate the capability of PanGu-Σ\Sigma model to efficiently and independently learn from multiple domains, we collect datasets in 40 domains, with a large amount of data in four major domains: Chinese, English, Bilingual (Chinese and English) and code. The remaining domains with smaller portion consists of 26 other monolingual natural languages, 6 programming languages, and textual data from finance, health, law, and poetry domains, respectively.

For Chinese texts, we collect the WuDaoCorpora 2.0 which contains 200GB and the CLUECorpus2020 which contains 100GB. For English texts, the Pile dataset which contains 800GB and C4 dataset which contains 750GB were collected. For code, we use the Python code (147GB) which has been used in PanGu-Coder , as well as the Java code (161GB) from GHTorrent , which are then filtered by file size (<<1MB), average number of characters per line (<<200), maximum number of characters per line (<<1000) and their compilablity. Then, these collected English, Chinese and code texts data was sampled and distributed to the four major domains. Finally, we get more than 300B tokens for the four major domains. The detailed statistics of data distribution and data sources in four major domains are presented in Table 1.

For the remaining 36 domains, the data for 26 monolingual domains are mainly from CCAligned and CCMatrix . Similar to the code domain mentioned above, the data for 6 programming language domains are collected through GHTorrent and filtered in the similar way. Finance domain data is filtered from the WuDaoCorpora 2.0 using the tags. Health domain data is from Chinese MedDialog Dataset . Law domain data is sampled from CAIL2018 . Poetry domain dataset is from Werneror-Poetery https://github.com/Werneror/Poetry. Finally, we sampled more than 25B tokens for the 36 domains.

2 Format

For the four major domains, each can be adapted to different downstream tasks. In order to better support domain-specific downstream tasks, this paper uses different data format for different domains. For Chinese and English domains, the ¡EOT¿ token which indicates the end of training text is inserted at the end of each training sample.

For Bilingual domain, the ¡EN¿ or ¡CN¿ token is inserted into the head of the training sample according to the source of the training sample (either from the Chinese dataset or the English dataset), and the ¡EOT¿ token is inserted at the end of each training sample.

For the code domain, the ¡Python¿ or ¡Java¿ token is inserted into the head of the training sample based on the programming language type of the training sample, and the ¡EOT¿ token is inserted at the end of each training sample.

For the remaining 36 domains, the data formats of 26 monolingual domains, finance, health, law, and poetry domains are the same as the Chinese and English domains, and the data format of 6 programming language domains is the same as the code domain.

For a formatted data set DD, suppose it contains n training samples D={s1,s2,…,sn}D=\left\{s_{1},s_{2},\dots,s_{n}\right\}. To make full use of the computing power of the Ascend 910 cluster and accelerate training in the pre-training phase, we concatenate all samples in the data set into a sequence, and then intercept training instances in the concatenated sequence according to the fixed length (1024), as shown in Figure 6. In the fine-tune phase, for each training sample in the formatted dataset, if the length is less than the fixed length, we pad the sample to the fixed length with a special token ¡Pad¿. If the length is greater than the fixed length, the extra part is truncated. Figure 7 shows the process. Different to PanGu-α\alpha model, each training sample of PanGu-Σ\Sigma model contains two field: input sequence of token IDs which are training instance and their domain ID. The domain ID indicates which domain the training instance belongs to. The RRE layers of the PanGu-Σ\Sigma model decide which experts the training tokens is routed to by the domain ID.

System

PanGu-Σ\Sigma is implemented with MindSpore 1.6 framework https://www.mindspore.cn/versions/en and trained on 512 Ascend 910 accelerators (also know as Ascend 910 NPU).

Training a trillion parameters language model poses multiple challenges. First, it requires enormous amount of memory in training. Although the sparse architecture can effectively save computation, it doesn’t reduce the memory consumption and we still need to store all the parameters and optimization states inside the accelerator memory. Assuming Adam optimizer with mixed-precision training is used, a 1T model typically consumes 16TB memory in total just for parameters, gradients and optimizer states. During training, the model needs extra memory for input data, network activations, communication buffers and temporary variables. We estimate that training a PanGu-Σ\Sigma model with 1 trillion parameters with a reasonably batch size needs more than 32TB memory and requires more than 1,000 Ascend 910 accelerators or NVIDIA V100 GPUs with 32GB High Bandwidth Memory (HBM).

Instead of pouring lots of hardware resources to scale-up the model, we aim to train PanGu-Σ\Sigma with a reasonably-sized cluster of 512 Ascend accelerators. To this end, we adopt the heterogeneous training and offload the optimizer states to CPU. After enabling heterogeneous training, all optimizer states are moved from accelerator to the host with 750GB host memory and KunPeng 920 CPU https://www.hisilicon.com/en/products/Kunpeng/Huawei-Kunpeng/Huawei-Kunpeng-920, and we can fit the entire training process into the cluster.

Second, the system throughput is unacceptable after enabling vanilla optimizer offloading. The root cause is again the sheer amount of parameters. Gradients and updated parameters need to be exchanged via the slow host-to-device and device-to-host communication, and CPUs need to iterate thorough all parameters and update them. To improve the training throughput, we leverage the sparse nature of PanGu-Σ\Sigma architecture. Since PanGu-Σ\Sigma use a sparse architecture and most of its parameters are conditionally activated, the optimizer only need to update part of experts in one iteration. So we propose Expert Computation and Storage Separation (ECSS) method as illustrated in Figure 8.

In Expert Computation and Storage Separation, we consider experts as knowledge database to store specific knowledge of different tasks or domains. In each iteration, experts are sparsely activated by different token IDs with specific domain. In MindSpore, we use lookup operator to select parts of activated experts, and sparsely update their parameters in the backward computation. In optimizer CPU offload computing, MindSpore copy FP16 parameters from host CPU to NPU, compute the gradients on NPU, move FP16 gradients from NPU to CPU, and compute optimizer states and update parameters in the host CPU. With a lower experts sparsity ratio such as 0.10.1, the computation cost is only near 10%\% of full model.

Besides ECSS with Ascend-KunPeng sparse heterogeneous computing, we also adopt other parallel training and accelerating techniques provided by MindSpore and CANN https://www.hiascend.com/en/software/cann. We use 8-ways model parallel for all the attention and feed-forward layers, 64-ways expert parallel without replica and 64-ways data parallel for non-expert parts. To further optimize memory footprint, rematerialization and optimizer parallel are also adopted to reduce the peak memory consumption. We also use FastGelu and fused LayerNorm to accelerate point-wise computation. By combining all the techniques together, we achieved 6.3 times throughput promotion compared to vanilla PanGu-Σ\Sigma heterogeneous training, as shown in Figure 9.

Experiments

We use the following PanGu-Σ\Sigma configuration for this work. The configuration mostly follows the 13B version of PanGu-α\alpha model. In this way, we can effective inherit the knowledge already learned by PanGu-α\alpha.

1.2 Pretraining settings

We use a cluster of 64 nodes, with each node equipped with 8 Ascend 910 accelerators and MindSpore framework. High performance collective communication library Huawei Collective Communication Library (HCCL) is used to facilitate high speed high bandwidth communication for distributed training.

There are two stages in PanGu-Σ\Sigma pretraining process. In the first stage, we activate four main domains’ experts to consume data from all the four main domains including bilingual, Chinese, English and codes. In the second stage, we let all the experts to consume all domain’s data. Figure 11 shows how 640 experts are assigned to 40 domain groups. We train PanGu-Σ\Sigma with global batch size of 512 with sequence length of 1024 for each sample. The pretraining lasts about 100 days. Figure 10 shows the loss curve of PanGu-Σ\Sigma pretraining.

Mixed-Precision training is enabled to speedup the training process. Apart from vocabulary embedding layer, loss function, Softmax operation, LayerNorm layers and Adam optimizer, all other operations adopt FP16 format.

Failure recovery is very important for long term large scale distributed training, especially for huge models like PanGu-Σ\Sigma . Therefore, a rigorous process of saving checkpoints and restarting from previous checkpoints is indispensable. For a trillion-parameter model, one set of checkpoint storing all parameters and optimizer states for a single iteration already has a jaw-dropping 10TB size. Uploading checkpoints of such size to our long term object store is a challenging task, since uploading all checkpoints at the same time quickly saturates the network bandwidth and inevitably lead to training failure. To solve this issue, we launch the upload process in a round-robin style and limit the number of the simultaneously running process. This solution proves to be effective and stable for our entire training process.

1.3 Hybrid Hyper-parameter ADAM Optimizer

We design a Hybrid Hyper-parameter ADAM Optimizer to provides further stability for PanGu-Σ\Sigma during the pretraining phase.

To better understand PanGu-Σ\Sigma training process, we inspected the statistics of training states and find out that the gradients of RRE layers are much smaller than a non-sparse model. To tackle such a problem, we first set a very small ϵ1\epsilon_{1} for all model parameters, then we go one step further and set an even smaller ϵ2\epsilon_{2} only for the RRE layers, since compared to the dense layers, sparse layers received smaller effective batch due to its conditionally-activated nature. Specifically, we set hybrid hyper-parameters for ADAM optimizer below:

2 Inheritance Learning

To improve the training efficiency, accelerate model convergence, and reduce carbon emissions during training, the PanGu-Σ\Sigma model inherits the capabilities of the existing model, and then continues to train in four domains simultaneously. In this paper, PanGu-Σ\Sigma inherits the PanGu-α\alpha 13B version.

Because PanGu-α\alpha’s vocabulary is mainly designed to support Chinese texts, we extend its vocabulary to support both Chinese and English texts. PanGu-Σ\Sigmauses Byte-level BPE instead of BPE adopted by PanGu-α\alpha, the vocabulary is formulated by adding T5 small vocabulary to PanGu-α\alpha’s vocabulary, then remove repeated sub-words. Some special tokens are added to the vocab. These special tokens are classified into two types: control tokens (e,g., ¡python¿, ¡Java¿, ¡CN¿, ¡EN¿) and spaces tokens for representing whitespace runs of different lengths.

2.2 Inheriting and Extending model parameters

In order to inherit the capability of the existing model as much as possible, PanGu-Σ\Sigma’s word embedding and all experts in RRE layer are initialized with the corresponding embedding and feed-forward layers from PanGu-α\alpha, and other parameters are initialized with corresponding parameters. For example, to initialize the word embedding parameters of PanGu-Σ\Sigma , we first create a word embeddings Ws∈Rvs×hW_{s}\in R^{v_{s}\times h}, if a sub-word of PanGu-Σ\Sigma exists in PanGu-α\alpha, its word embedding is initialized with those of PanGu-α\alpha. And if not, they are randomly initialized with a standard normal distribution. For the experts parameters in the RRE layer of PanGu-Σ\Sigma , each expert is initialized with the FFN parameters of the corresponding layer in the PanGu-α\alpha model.

In order to reduce the mutual interference between English and code domain in the training process, we make the code domain and other domain updated in different embedding slots. Therefore, we further extend the PanGu-Σ\Sigma word embedding Ws∈Rvs×hW_{s}\in R^{v_{s}\times h} to Ws′∈Rvs′×h,(vs′=2×vs)W_{s^{{}^{\prime}}}\in R^{v_{s^{{}^{\prime}}}\times h},\left(v_{s^{{}^{\prime}}}=2\times v_{s}\right). The slots [vs,2×vs]\left[v_{s},2\times v_{s}\right] of word embeddings Ws′W_{s^{{}^{\prime}}} belongs to code domain and the slots [0,vs]\left[0,v_{s}\right] belongs other domain. Figure 12. shows how PanGu-Σ\Sigma inherits the PanGu-α\alpha’s parameters and extends it.

2.3 Extracting domain specific sub-model

It is expensive to deploy a trillion parameters model like PanGu-Σ\Sigma directly. In order to transfer abilities of PanGu-Σ\Sigma to various downstream tasks and reduce the consumption of serving resources, we propose a loss-free expert pruning method by leveraging the RRE design. Domain models can be separately extracted for further fine-tuning, evaluation and deployment. Figure 13 illustrates how to extract the the domain specific sub-model from PanGu-Σ\Sigma . For the word embedding, the word embedding slots which belongs to the domain are extracted. For the experts in the RRE layers, the experts allocated for the specific domain are extracted. Other parameters of PanGu-Σ\Sigma are copied seamlessly.

3 Chinese Downstream Tasks Evaluation

Following PanGu-α\alpha, we evaluate PanGu-Σ\Sigma at zero-shot settings on 16 datasets of six tasks. For each dataset, if the test set is available, we use it to evaluate the model. Otherwise, we use the validation set. The following describes each task in turn.

Machine reading comprehension. This task contains four datasets: CMRC2018 , DRCD , DuReader , and C3 . The first three datasets CMRC2018, DRCD, and DuReader are span extraction tasks. We formulate each of them into a text generation task, using the model to generate answers based on given passages and questions. And we use F1, exact match (EM), and ROUGE-1 as the evaluation metrics. In addition, for the DuReader dataset, which is aligned with PanGu-α\alpha, only the Zhidao subset is selected to evaluate the model performance. The last dataset, C3, is a multi-choice reading comprehension task. Given a passage, a question, and multiple candidate answers, the purpose is to select one of the candidate answers as the predicted answer to the question.

Natural language inference. There are two datasets: OCNLI and CMNLI . Given two sentences, one as a premise and the other as a hypothesis, the aim is to determine whether the relation between the premise and the hypothesis is entailment, neutral, or contradiction. We convert this task into a three-class classification problem to solve.

Text classification. We use TNEWS and IFLYTEK datasets. The total number of categories for TNEWS and IFLYTEK is 15 and 119, respectively. Following PanGu-α\alpha, for each instance, we randomly sample three negative categories plus one ground-truth category to form a new set of candidate categories, then simplify this task into a four-class classification task for processing.

Semantic similarity. We use two datasets: AFQMC and CSL . AFQMC aims to determine whether two sentences are semantically the same or different. Given an abstract of a paper and a set of keywords, the goal of CSL is to judge whether the set of keywords contains pseudo keywords according to the abstract. Hence we convert each of them into a two-class classification problem to solve.

Winograd schema challenge. This task contains only the CLUEWSC2020 dataset. CLUEWSC2020 is a coreference resolution task. Given a sentence, together with a pronoun and a noun in the sentence, the aim is to determine whether the pronoun refers to the noun. We merge multiple instances with the same sentence and the same pronoun into a single instance that contains a sentence, a pronoun, and multiple nouns. Then the goal becomes to select one of the multiple nouns as the object the pronoun refers to.

Cloze and completion. There are five datasets: CHID , CMRC2019 , PD , CFT , and CMRC2017 . Both CHID and CMRC2019 are multi-choice completion tasks. Given a passage with multiple blanks and multiple candidate answers, for each blank in the passage, the goal is to select the appropriate one from all the candidate answers to fill in the blank. For CHID, we use the Hungarian algorithm to post-process the model prediction results to ensure that different blanks in the same passage are filled in different idioms. On the CMRC2019 dataset, following ERNIE 3.0 Titan , for each blank, we randomly sample three negative candidate answers plus one ground-truth answer to form a new set of candidate answers, and moreover, beam search is also used in the model prediction process to find an optimal combination of answers for multiple blanks in a passage. CMRC2017 contains two subsets, one for completion and the other for reading comprehension. As with PanGu-α\alpha, we also evaluate PanGu-Σ\Sigma only on the completion subset. For CMRC2017, PD and CFT, given a passage with a blank, the goal is to fill in the blank with the appropriate words. Aligned with ERNIE 3.0 Titan, we also convert PD, CFT and CMRC2017 into multi-choice completion tasks, and the choices are all words that appear in the passage where the blank is located.

3.2 Evaluation Details

Each dataset of all Chinese downstream tasks can be evaluated using either a generation-based method or a scoring-based method. We use the generation-based method to evaluate CMRC2018, DRCD, DuReader, and the scoring-based method to evaluate other datasets. For each instance, a text sequence is obtained by filling it into a manually designed template, and then the text sequence is fed into PanGu-Σ\Sigma for prediction to get the result. The templates we used for all datasets are shown in Table 4.

For each instance to be predicted, it is filled into the corresponding template to obtain a text sequence. After that, the text sequence is used as the input to PanGu-Σ\Sigma to generate the answer. We use a greedy decoding strategy to generate the answer.

Each instance to be predicted contains multiple candidate answers. For each candidate answer, a text sequence is obtained by filling the candidate answer together with the sample into the corresponding template, and the perplexity of the text sequence is calculated by PanGu-Σ\Sigma. Finally, the candidate answer corresponding to the text sequence with the smallest perplexity is selected as the predicted answer for the instance to be predicted.

3.3 Result

We choose PanGu-α\alpha and ERNIE 3.0 Titan as the baseline for comparison. The performance of each Chinese downstream task is shown in Table 5. Compared to ERNIE 3.0 Titan with 260 billion parameters, PanGu-Σ\Sigma surpassed on 11 out of 16 datasets, with an average score of 3.96 points higher on all datasets.

4 Chinese Dialogue Generation

To verify the ability of PanGu-Σ\Sigma on Chinese dialogue generation, in this subsection, we compare with several high performance Chinese dialogue systems, including CDialGPT , EVA , EVA 2.0 and PanGu-Bot . The PanGu-Σ\Sigma model is fine-tuned on about 51.5M dataset including social media data, knowledge-grounding dialogue and question answering data, which is consistent with PanGu-Bot. PanGu-Σ\Sigma consistently outperforms baselines on self-chat, topic-grounded dialogue generation and question answering in terms of automatic evaluation and human evaluation.

CDialGPT: A GPT-based Chinese dialogue model trained on a large-scale cleaned Chinese conversation dataset LCCC, which contains about 104M parameters.

EVA: An encoder-decoder-based Chinese dialogue model trained on WDC-Dialog corpus including 1.4B Chinese context-response pairs. This model contains about 2.8B parameters.

EVA2.0: An improved version of EVA. A well designed data processing pipeline is explored to construct training data based on WDC-Dialog corpus, and various decoding strategies are utilized to improve generation. Furthermore, EVA2.0 designs better model architecture for open-domain Chinese dialogue, including attention scale strategy, deeper decoding network, and role embedding.

PanGu-Bot: The PanGu-α\alpha based Chinese dialogue models trained on collected 51.5M dialog sessions, which contain two versions of 350M and 2.6B parameters, respectively. To improve training efficiency, multiple dialogue sessions are concatenated with a special token, and resetting strategies on position ids and attention masks are utilized to distinguish different samples.

4.2 Self-chat evaluation

Self-chat is a common method for evaluating the quality of dialogue systems. During the evaluation, the conversation goes based on given prompts, with dialogue system playing both roles of user and bot. In this subsection, we provide 50 prompts to trigger multi-turns conversation with each containing 9 turns. We use top-5 random sampling with repetition penalty set to 1.2 during decoding. Three human annotators are asked to judge whether each turn conforms to the following six criteria: 1) Sensibility evaluates the semantic-consistency with the context of response; 2) Specificity evaluates the specificity and informativeness of response; 3) Interestingness evaluates the interest of response and the ability to catch people’s attention; 4) SSI averages values of Sensibility, Specificity and Interestingness; 5) Hallucination evaluates factual mistakes contained in response; 6) Safety evaluates the avoidance of unsafe behavior of dialogue system, e.g. response with social bias, toxicity, harmfulness and offensives.

As shown in Table 6, in self-chat evaluation, the overall response quality of PanGu-Σ\Sigma is much higher than the baselines, especially in terms of Specificity. This is because PanGu-Σ\Sigma inherits the 13B version of PanGu-α\alpha model, and the sub-model for dialogue generation contains about 38B parameters, which can memorize a wealth of knowledge. The improvements in terms of Hallucination and Safety indicate that PanGu-Σ\Sigma can learn the patterns of knowledge and safe expression in human dialogue effectively, and therefore generate factually correct and safe responses. A case of self-chat is shown in Figure 14, where the conversation goes smoothly with rich knowledge. More self-chat cases are shown in Appendix A.1.

4.3 Topic-grounded dialogue evaluation

A well-designed dialogue system should be able to incorporate relevant knowledge with characteristic of chit-chat. Therefore, in this subsection, we aim to evaluate the performance on topic-grounded dialogue, where the dialogue history contains abundant knowledge and topic information. We randomly sample 2,000 dialogues from topic-grounded corpus NaturalConv , and keep each context containing at least 5 turns. We use nuclear sampling with top-p set to 0.5 during decoding. The following metrics are used for automatic evaluation: 1) Semantic consistency measures the consistency between generated response and context, which is scored by a BERT-based binary classifier model trained on NaturalConv with an accuracy of 0.906; 2) Distinct-1 and distinct-2 are the ratios of distinct unigrams and bigrams in response, respectively, for evaluating the diversity; 3) Bleu can evaluate the n-gram overlap degree between generated response and golden response.

The results of automatic evaluation and human evaluation are shown in table 7 and table 8, respectively. Compared with baselines, PanGu-Σ\Sigma can generate more diverse, semantic-consistent, knowledgeable and interesting responses. This is because PanGu-Σ\Sigma can well response with consideration of topic and knowledge contained in dialogue history. A case of topic-grounded dialog is shown in Table 9, where the response of PanGu-Σ\Sigma introduces knowledge about 郎平(Lang Ping). More topic-grounded cases are shown in Appendix A.2.

4.4 Open domain question-answering evaluation

For evaluating the PanGu-Σ\Sigma’s ability to answer fact-based question in conversation, 6 categories of questions collected from PanGu Bot are utilized for evaluation. The greedy search decoding strategy is applied. The results of open domain question-answering evaluation is shown in table 10. PanGu-Σ\Sigma model can well answer factual questions with highest accuracy, which can further verify the advantages of PanGu-Σ\Sigma on knowledge memorization. A case of question-answering is shown in Table 11, where the answer of PanGu-Σ\Sigma is the the most accurate. More cases of question-answering are shown in Appendix A.3.

4.5 Natural language generation of base PanGu-ΣΣ\Sigma model

For evaluating base PanGu-Σ\Sigma’s abilities on open-ended text generation, we present three categories of cases about character dialog, question-answering and text generation with few-shot prompt learning, which are shown in Table 12, Table 13 and Table 14 respectively. More cases of character dialog and open-end text generation are shown in Appendix A.4 and Appendix A.5.

5 Machine Translation

To verify the generative and multilingual ability, we compare the performance of PanGu-Σ\Sigma with the state-of-the-art model CeMAT, and the benchmark pre-trained large models (mT5-XXL, CPM-2, ERNIE3.0) on the machine translation task. Following the existing pre-training methods, we use the PanGu-Σ\Sigma model to fine-tune directly on the dataset of translation tasks and use SacreBLEU as an evaluation metric. We perform validation on two mainstream datasets, WMT17 and WMT20, covering two different translation reversals, Chinese-English and English-Chinese, respectively. The experiments find that PanGu-Σ\Sigma has a large improvement over both baseline models, and in low-resource translation experimental scenarios even outperforms significantly the results of full data fine-tuning of other pre-trained models.

mT5 is a multilingual variant of T5, which leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. mT5 was pre-trained on a new Common Crawl-based dataset covering 101 languages and achieved the state-of-the-art performance on many multilingual benchmarks, such as machine translation. mT5-XXLarge is a largest version of mT5 with 13B parameter size.

CPM-2 is a large-scale cost-efficient pre-trained language model. CPM-2 accelerate the pre-training process by dividing the pre-training process into three stages: Chinese pre-training, bilingual pre-training, and MoE pre-training. To test the cross-lingual generation ability, we use the bilingual version model.

ERINE3.0 is a unified framework for pre-training large-scale knowledge enhanced models. Which fuses auto-regressive network and auto-encoding network, and can be easily tailored for both natural language understanding and generation tasks with zero-shot learning, few-shot learning or fine-tuning.

CeMAT is a universal Conditional Masked Language Pre-training for both Autoregressive and non-Autoregressive machine translation tasks. Which is also a multi-lingual pre-trained language model consist of 32 languages.

During the fine-tuning, we used the language tag ”¡Language ID¿” as the prefix for Chinese and English text sequences respectively, and then spliced the source and target sequences together as the input to the model, with the source sequence at the beginning of the sequences and the two sequences separated by ”¡EOT¿”. We first verified PanGu-Σ\Sigma on WMT20 Chinese-English dataset, almost large scale pre-trained language model eval the cross-linugal genetation ability on that. In addition to this, we also compare the performance of PanGu-Σ\Sigma with the current SOTA translation pre-trained model CeMAT on WMT17 datasets, and covering two different translation reversals, Chinese-English and English-Chinese, respectively.

As shown in Table Table 16, on the WMT20 Chinese-English translation task, PanGu-Σ\Sigma exceeded the mT5-XXL model by 12.6 BLEU, which also showed a significantly higher improvement of 9.8 BLEU compared to the Ernie3.0, which is the Chinese-English SOTA pre-trained large model, indicating that the PanGu-Σ\Sigma model was able to learn stronger cross-language understanding and generation abiliby from the pre-trained data. To further verified PanGu-Σ\Sigmaś performance in low-resource scenarios, we using a randomly sampled 30w training dataset, as shown in Table 2. Using only a small amount of training data, the PanGu-Σ\Sigma model still outperformed large models such as Ernie3.0 by more than 3.19 BLEU, which used a full 26M of training data.

Compared to the translation pre-trained language model CeMAT, PanGu-Σ\Sigma also shows a meaningful quality improvement. As show in Table Table 16, the PanGu-Σ\Sigma pre-trained model exceeds the CeMAT model by 3.0 BLEU on the English-Chinese task and also has a 0.7 BLEU improvement on the Chinese-English task, both achieving SOTA results. Also, in the specific translation case, we found that the translation results of PanGu-Σ\Sigma model have higher fidelity compared with CeMAT, as shown in Table Table 17.

6 Code Generation

In order to measure the performance of PanGu-Σ\Sigma on code downstream tasks, we evaluated the performance of PanGu-Σ\Sigma ’s code domain model on MBPP tasks. MBPP is a benchmark to measure the ability of pre-trained models to generate Python programs from natural language descriptions. The MBPP datasets contain 374 programming problems for fine-tuning and 500 programming tasks as test dataset. Each sample in fine-tuning dataset contain function description, three test cases which check for functional correctness, and function code which is a ground-truth solution that passes all test cases. Figure 15 shows a sample in the MBPP fine-tune dataset.

PanGu-Coder introduces additional datasets which contain APPS and Code Contests (CC) datasets from a more similar distribution. The additional datasets provide a large number of competitive programming problems. APPS includes 10, 000 programming tasks that generate or complete code given the problem description. Code Contests (CC) containing over 13k programming problems. PanGu-Σ\Sigma also introduces these additional datasets. For each problems in APPS and CC, we up-sample 5 different correct solutions. Then, we filter the samples with text length over 1024. Finally, we get 56k instances for fine-tuning.

To make it easier for the model to distinguish between task descriptions and solutions, we format training instances for fine-tuning. For these instances in MBPP, we concatenate function description and three test cases to form prompt, and then add a ¡comment¿ token to the head of the prompt and a ¡python¿ token to the end of the prompt. Function code is appended to the ¡python¿ token and the ¡EOT¿ token is add to the end of function code. Similar for these instances in APPS or CC, the only different is that the function description is treated to prompt directly.

All these formatted instances from the MBPP fine-tune dataset, APPS and CC constitute the fine-tune datasets. We fine-tune the code domain model, which is extracted from PanGu-Σ\Sigma , for 5 epochs on the fine-tuning datasets.

6.2 Results

For all sample in MBPP test dataset, function descriptions are augmented with three test cases as prompt, which is similar to the data format used during fine-tuning. We use greedy decoding to generate function code based on the formatted prompt. If the generated code passes all three of the given test cases, then the generated code passes the test. To evaluate the performance of the fine-tuned code domain mode of PanGu-Σ\Sigma , we use the pass@1 as the estimator. The pass@1 of model refers to generating only one function code for each sample in the test dataset, and then counting the percentage of the generated function code that passes the test.

Table 18 shows the comparison of existing models, as well as PanGu-Σ\Sigma on the MBPP dataset, along with model size and number of tokens trained by model. The PanGu-Σ\Sigma outperforms the current state-of-the-art model PanGu-Coder by 1.4 point on the pass@1 for MBPP tasks. The training data of PanGu-Σ\Sigma is less than PanGu-coder, which contains only 75B code data, while Python code data related to MBPP tasks is only 50B data. This suggests that PanGu-Σ\Sigma makes more efficient use of data.

7 English Natural Language Understanding

In order to compare with other large language models on English tasks, we evaluate PanGu-Σ\Sigma model on the SuperGLUE benchmark . SuperGLUE consists of 8 natural language understanding tasks. We use accuracy as the performance metric except for MultiRC dataset where F1{\rm F1}-score over the set of answer options is used (denoted by F1a{\rm F1}_{a}). We cast each task to a multiple-choice classification problem. The prediction is chosen based on the maximum log-likelihood score, log⁡P(completion ∣ context)\log{\rm P}(completion\,|\,context), of each available completion given the context. For some of the datasets, we normalize this score by the token length of the completion, but for COPA and RECORD non-normalized scores yield better results. We generally view binary classification in such a way that the completion options are “Yes” and “No”, except for the COPA for which the model chooses between two appropriate sentence continuations. In the table 19, we report model’s performance on each of the SuperGLUE datasets along with the average score. We focus on the zero-shot setup and make a comparison with the GPT-3 model which has a similar evaluation setup.

In the Table 19, evaluation results are presented. We see that, even with only 112B English tokens, the performance of PanGu-Σ\Sigma English sub-model with 38B parameters roughly meets the performance of the GPT-3 13B model and gets a higher average score.

Conclusion and Future Work

In this work, we have present a trillion parameters language model architecture PanGu-Σ\Sigma . With the Random Routed Experts (RRE) and Expert Computation Storage Separation (ECSS), PanGu-Σ\Sigma achieves high system performance under the MindSpore framework using Ascend 910 AI accelerators. By extending and continually training from PanGu-α\alpha with 329B tokens, PanGu-Σ\Sigma has successfully achieved state-of-the-art results in a bunch of downstream tasks such as few-shot NLU, open-domain dialogue, question answering, machine translation, and code generation. Despite these achievements, there remain some worthwhile problems to pursue in the future work.

Sparse models offer the benefits of a larger model size at reduced computation cost. Despite the existing advancements, numerous algorithmic and system challenges persist within the sparse architecture. Addressing these challenges and creating a user-friendly, high-performing sparse architecture system continues to be an open problem.

Large language models are designed to be applied in real-world scenarios. Therefore, to enhance model evolution, the system should receive accurate feedback from the open environment. Although InstructGPT and ChatGPT https://openai.com/blog/chatgpt provide promising approaches, they require a substantial amount of data labeling, which can be time-consuming and costly. Consequently, devising an efficient method to generate valuable signals that can align with the real world is a crucial research topic worth exploring.

Large-scale language models provide intelligent foundations and various modalities alignment objectives for artificial intelligence systems. Therefore, utilizing language models as a foundation and incorporating multiple modalities for perception input in a multimodal model will be one of the most important topics, as already demonstrated by Flamingo and GPT-4 .

Large language models have great potential for real time applications, but their deployment cost remains a major hurdle to overcome. To make them more accessible for commercialization, researchers should focus on two directions: 1) to explore techniques to compress the large language model’s size while preserving its emergence abilities; 2) to optimize the system software and/or hardware to accelerate the model’s performance. Both of these directions are valuable for the deployment of large language models.

Online knowledge updates are also critical for optimal performance of the large language model system. However, effectively storing and updating knowledge online is a significant challenge that requires advanced system infrastructure and algorithms. As large-scale language models continue to develop, the issue of online learning will undoubtedly become increasingly crucial and a key topic for the future research.

Acknowledgements

We would like to thank Bin Zhou, Zhiwei Wang, Yasheng Wang, Liangyou Li, Bin He and Fanyi Du for their great support for this work; Chen Li, Yifan Yao, Kaisheng Wang, Zhenzhang Yang, Zhongzhe Hu, Zhepeng Sun, Zhijian Guo, Jun Wang, and Ziqiang Chen for their help to handle infrastructure issues.

References

Appendix A Natural Language Generation Examples

A.2 Topic-grounded dialog generation

A.3 Open domain question-answering of dialog model

A.4 Character dialog generation

A.5 Open-ended text generation