M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining
Junyang Lin, An Yang, Jinze Bai, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Yong Li, Wei Lin, Jingren Zhou, Hongxia Yang
Introduction
The expeditious growing of foundation models on broad data highly contributes to the development of the whole deep learning and artificial intelligence community. Foundation models with self-supervised learning on big data have become an emerging paradigm of artificial intelligence systems , as they mostly possess high transferability to a wide range of downstream tasks and even multiple modalities. The scale of foundation models across domains, including natural language processing, computer vision, and cross-modality representation learning, have been growing tremendously from millions to trillions of parameters [9; 26; 27; 29; 36; 40; 4; 32; 17; 13; 19; 50; 49; 45; 47; 41] thanks to the concurrent advancement in distributed training framework [23; 30; 34; 31; 2; 14; 33] and hardware design, and these studies have made a demonstration of the neural scaling law . However, the training of these transformer-based models incurs high financial costs and even environmental damage due to the massive carbon footprint and thus training extreme-scale models under a decent amount of resources but with high efficiency should be a fundamental goal for both the research and industrial communities to achieve, which promotes the progress of greener AI [25; 37].
Generally there are two tracks of research in large-scale pretraining, dense models and sparse expert models respectively. A typical case of large-scale dense models is GPT-3 . It is a 175-billion-parameter transformer model trained with GPUs for months, incurring striking financial and environmental costs. Researchers have been looking for methods to training large-scale models with a decent amount of costs. Solutions include effective management of memory with gradient and optimizer state partitioning or more efficient model parallelism and pipeline parallelism [40; 23; 14]. A series of following studies apply those techniques to realize fast training of 10-billion-parameter transformers with hundreds of GPUs in months. [19; 49; 47; 41] Sparse expert models with large model capacity are capable of fast training owing to the combination of data parallelism and expert parallelism [38; 17; 13; 35], and it is even accessible to train a 1-trillion-parameter transformer with no more than GPUs .
Be there as it may, a question emerges in our mind: is it possible to train an extreme-scale model with only a decent amount of resources, e.g., training a 10-trillion-parameter model with GPUs? Such training requires large memory for parameters, including weights, gradients, and even optimizer states. Tackling the problem requires the utilization of external memory except for GPU memory, for instance, CPU memory or even NVMe storage [34; 31]. These methods resolve the problem of high memory footprint, but instead, their extra cost is low training efficiency caused by the frequent swap in-and-out between memories.
In this paper, we provide a solution to training large models that require high memory footprint, and we demonstrate a successful practice of pretraining unprecedented extreme-scale models with over trillion parameters, an order of magnitude larger than the previous state-of-the-arts [13; 45]. The whole pretraining was conducted on solely NVIDIA-V100 GPUs and lasts around days. A simple and effective training strategy called “Pseudo-to-Real” enables sharing and delinking parameters. This training strategy is compatible with architectures built by stacking layers with an identical structure, including dense models like GPT [26; 27; 4], BERT , or sparse expert models like M6 [19; 45]. It is essentially a two-stage training, where in the first stage, we apply cross-layer parameter sharing that requires much less memory footprint for efficient convergence, and in the second, we delink the parameters for better performance. It first trains a relatively small model but with a computation graph of a large one with the utilization of cross-layer parameter sharing, and we name it “Pseudo Giant”. Then it builds a correspondingly large model and delinks the parameters of the shared layer for second-stage model initialization. In this way, we achieve fast convergence in the first stage as the training costs much less memory and speeds up with large batches. Parameter sharing that addresses the communication overhead improves training speed as well. The second-stage training is responsible for the final convergence for better performance.
We unlock the secret of pretraining an unprecedented extreme-scale model with over 10 trillion parameters on limited resources of GPUs. Compared with the previous M6-T on around GPUs, we do not have a significant increase in computation resources but level up the model scale by an order of magnitude. Besides the application of the “Pseudo-to-Real” training strategy, we provide a faster offloading mechanism for both management of CPU memory for parameter storage and utility of GPUs. We successfully train the M6-10T within days to reach strong performance in log perplexity evaluation and outperform the baseline M6-T.
We illustrate the training difficulty of extreme-scale models on limited resources and provide a simple but effective solution called “Pseudo-to-Real”. Upstream and downstream evaluation demonstrates the effectiveness of the strategy.
We further demonstrate a successful practice of pretraining a 10-trillion-parameter model on GPUs and reach an outstanding performance within days.
Related Work
In recent years, pretrained language models with growing magnitudes of parameters have been proposed, keeping to raise the validated upper limit of scaling law for model capacity w.r.t the number of parameters . Earlier milestones of extreme-large models come from GPT-2 and Megatron-LM , which demonstrates that scaling the transformer model up to billions of parameters can result in improvement on language modeling benchmarks [22; 24]. Turing-NLG , as a successor, implements a 17-billion-parameter transformer and achieves further lower perplexity. Similar phenomena are also observed on classification language tasks by T5 model . The GPT-3 pushes the boundary of model scale to 1,750 billion parameters and demonstrates its striking effectiveness on downstream tasks in even zero-shot settings. Furthermore, large-scale pretraining has recently demonstrated success in other fields, including pretraining on other languages or cross-lingual pretraining [44; 49; 47; 7], cross-modal pretraining [28; 32; 19; 10; 51] and code generation . Along with the benefit from increasing the model scale, the concern of unaffordable pretraining cost in time, computation resource and energy keeps emerging [25; 3], resulting in the strong demand for more efficient and greener large-scale pretraining .
Our work is about designing algorithm to train extreme-scale models efficiently. Researchers have been demonstrating different types of methods to reach this objective. To reduce the computational cost during training, some studies introduce sparsity to the model. Mixture-of-Experts (MoE) was proposed and was reintroduced in Mesh Tensorflow . It shows that MoE with sparsity can significantly improve the model scale efficiently without increasing computation, and the models achieved state-of-the-art performance in language modeling and machine translation. GShard extended it to a tremendously large scale of 600B parameters with the sophisticated collaborated design of model architecture and demonstrated its effectiveness across over 100 languages. Similarly, Switch Transformer reached trillion parameters and showed its strong performance on NLU tasks. Those models with high sparsity are computationally efficient, and therefore though they possess large model capacity they still can be trained with high efficiency. Most other studies still focus on training dense models to validate the scaling law, and thus the emerged problem is the distributed training of large models. The most influential distributed framework should be DeepSpeed that proposed ZERO that partitions optimizer states and gradients to multiple GPU devices, and ZERO-offload as well as ZERO-Infinity can even offload parameters to CPU memory and NVMe storage. The memory footprint management makes training extremely large models on limited resources possible. However, such offloading mechanisms still have some defects that they may fail to fully utilize the fast hardware. For example, when offloading all parameters to the CPU, the GPU memory can be idle. In this work, we tackle this issue by proposing a granular offloading mechanism that can determine which parameters to be offloaded.
In addition to reducing the amount of computation in a single iteration, another route to speedup training is to reduce the needed iterations for model convergence. Child et al. and Xiong et al. propose to put forward the layernorm operations in transformer blocks for more stabilized and faster convergence. You et al. employs a layer-wise adaptive optimizer to enable super-large pretraining batches. Zhang & He progressively increases the layer dropping rate in a stochastic manner which significantly speedups pretraining.
Approach
This section describes our proposed two-stage training strategy, “Pseudo-to-Real”, and shows experimental results to validate its effectiveness for training high-memory-footprint-required large models.
Choice in model architecture depends on several factors. First, the architecture should contain a sequence of stacking layers, as the sequential structure enables cross-layer parameter sharing. We prefer a simple encoder or decoder architecture, instead of an encoder-decoder framework where there are cross attentions that bring extra parameters and incur difficulties in activation checkpointing. Second, a model of such architecture should be compatible with different types of downstream tasks including understanding and generation, and it is even better that it can be compatible with multiple modalities. Third, as we mention that dense models and sparse expert models are two main tracks of large-scale pretraining, we prefer the model that can flexibly become whether dense or sparse expert models. Therefore, we select M6 as an option, and we evaluate the effects of our method on M6 of different scales and types.
M6 is built with stacking transformer layers, which includes self attention and feed-forward neural nets (FFN). For the transformation from dense models to sparse expert models, we should only replace FFN layers with the Mixture-of-Expert (MoE) layers. MoE consists of multiple experts, which are usually FFNs distributed on different devices. A gating network decides the dispatching and combining behaviors of each token and thus tokens can be processed in diverse devices. Such mechanism is a combination of data parallelism and expert parallelism, and thus it is highly efficient though with large model capacity. For the training, to realize the learning of both understanding and generation, the model is trained with text denoising and language modeling on plain text data and with image-based text denoising and image captioning on multimodal data. The model is compatible with different types of downstream tasks and can process information of multiple modalities.
2 Pseudo-to-Real
This section demonstrates the details of “Pseudo-to-Real” two-stage training strategy that enables fast training of high-memory-footprint-required transformer models. The strategy consists of two stages. The first stage trains a model with many fewer parameters but with a large computation graph (“Pseudo Giant”), and the second stage trains a corresponding large model (“Real Giant”) initialized with the delinked weights of the shared layer. Thus we name the strategy “Pseudo-to-Real”, and the general idea is demonstrated in Figure 1.
The core of “Pseudo” stage is to train a Pseudo-Giant that shares parameters across layers. Cross-layer parameter sharing has proved successful in maintaining satisfactory performance while keeping a much smaller amount of parameters. The method was first mentioned in the original Transformer , and Dehghani et al. and Bai et al. both illustrated that it is effective for vanilla transformer with encoder-decoder work. Lan et al. introduced the method to pretraining and proposed a lite BERT with different sharing techniques. Owing to its effectiveness, we introduce it to training an extreme-scale model, and we hypothesize that the first-stage training can gain benefits from cross-parameter sharing as it can address communication overhead and it consumes much less memory footprint. Also, as Pseudo Giant with much fewer parameters is not bounded by memory, it can be trained with large batches for acceleration.
Suppose we build an M6 model with layers that share parameters across all layers. The Pseudo Giant though consists of a computation graph of a -layer transformer, its number of weight parameters and optimizer states should be of those of the original one. As to the gradients, we can accumulate the gradients of each layer in the backward computation process, and therefore the amount of gradients becomes of the original one. Such saving in memory enables much faster training with larger batches. Also, it is capable to use fewer resources even due to lower memory consumption.
This can also be applied to MoE models, as their architecture is stacking transformer layers. However, different from dense models, MoE models partition their weights to all devices and redistribute token representations by sparse activation. This limits the flexible usage of GPU resources. It is available to choose different numbers of GPU devices at different stages in training dense models. In contrast, the restoration of MoE model parameters requires the same number of GPU devices used at the last stage. To tackle this problem, we take advantage of the memory efficiency of the Pseudo stage, and distribute more experts to a single GPU and partition experts to more GPUs. Suppose we train a model with experts at each MoE layer. It is possible to train a Pseudo Giant with only GPUs, where there are experts on each GPU, and then train a Real Giant on GPUs where there is only expert on each GPU.
2.2 “Real” Stage: Delinking the Shared Parameters
We name a large model without cross-layer parameter sharing “Real Giant”, in comparison with Pseudo Giant. Both Pseudo Giant and Real Giant share a computation graph, but they have different numbers of parameters. In this work, we discover how to build a connection between the two models. Given a Pseudo Giant fully trained until convergence, we apply the delinking of cross-layer shared parameters to accelerate Real Giant training. There is no need to train a large model from scratch. The model can start its convergence from low perplexity.
Embedding initialization can be directly restored, but the layer weights should be treated specially. In practice, there is only one layer of weights in Pseudo Giant, and there are layers of weights in Real Giant. Thanks to their identical structure, each layer of Real Giant can be initialized with . Without further training, this model is equivalent to a fully-trained Pseudo Giant.
This extremely simple training strategy is highly beneficial for the high-memory-footprint-required large models, especially extreme-scale models like the 10-trillion-parameter M6. While the first stage of training saves much time for faster convergence, we can use a decent amount of computational resources in this stage as lower efficiency in this stage becomes acceptable. Therefore, in the practice of training an extreme-scale M6, we apply CPU offloading to utilize CPU memory. Therefore, we can use a limited amount of resources, e.g., GPUs, to train an unprecedented 10-trillion-parameter model efficiently, which is an order of magnitude larger than the state-of-the-arts.
2.3 Timing for Switching
A question naturally emerges: when should we switch from the Pseudo stage to the Real stage? As mentioned above, the greatest advantage of Pseudo stage for training is the significantly faster convergence speed. Yet the performance of Pseudo Giant is bounded by its limitation in the number of parameters. Training Pseudo Giant until its final convergence apparently incurs much waste of time.
In practice, we present a simple strategy to determine the training step to switch from Pseudo to Real based on the convergence speed. During the training of the Pseudo stage, we evaluate a training step in a fixed interval by attempting to transfer it into the Real stage and training for a small while (e.g., minutes). After that, we will revert the model parameters to the evaluated step and continue the training of the Pseudo stage for the same training time as the Real stage. If the decreasing speed of loss in the Real stage surpasses that of the Pseudo stage, we determine the evaluated training step as the best switching point for the next-stage training.
3 Experiments
In this section, we provide a series of experiments to evaluate the effectiveness of the training strategy by observing the models’ upstream and downstream performance.
We conduct experiments for pretraining and finetuning to analyze model competence in upstream and downstream tasks. Following the classical data setup for pretraining and finetuning, we pretrain the model on BookCorpus and English Wikipedia , which are corpora with around 16GB of plain texts.
We validate the effectiveness of “Pseudo-to-Real” training strategy by implementing a medium-size M6 model to conduct extensive analyses on both upstream and downstream performance. To satisfy the requirements of high memory footprint where the two-staged training can make difference in training efficiency, we conduct experiments on a 1.4B-parameter model, the largest model trained on a single NVIDIA V100-32GB . The corresponding Pseudo Giant is a model with the same computation graph but sharing parameters across layers.
Following Radford et al. and Lewis et al. , we use a vocabulary of around subwords. Each sample consists of sentences from an identical passage, and we use a sequence length of and correspondingly truncate or pad the sequence. We build an M6 model with layers of transformer, whose hidden size is and intermediate size is . As our evaluation is conducted on plain text data, we adopt the two tasks, text denoising and language modeling, for pretraining. We name the one with cross-layer parameter sharing “Pseudo” and the one without it “Real” in Table 1. The total number of parameters of Real Giant is around billion and that of its corresponding Pseudo Giant is around million. We use “P2R” referring to “Pseudo-to-Real” to represent the model trained with both “Pseudo” and “Real” stages. “Pseudo” and “Real” refer to the models that are pretrained from scratch, in contrast with “P2R”. Furthermore, we build an M6 model of M parameters as a baseline. It has a smaller intermediate size of and consists of layers. We name it “Base” in Table 1.
Following the common practice in pretraining [9; 20], we apply AdamW optimizer for optimization. To determine the most suitable learning rate of the two stages for fast convergence, we have made some preliminary tests and finally used the peak learning rate of for “Pseudo” stage and for “Real” stage, respectively. We use a cosine decaying mechanism for learning rate scheduling and a warm-up ratio of . For both Pseudo and Real Giant, we train them with a total batch size of . In practice, we use a micro-batch size of and a gradient accumulation step of , and we train our models on NVIDIA-V100 GPU devices. We create two setups for better comparison. The one is “limited budget”, where models have been trained for the same duration of time. Correspondingly, “Pseudo” has been trained for around steps, Real has been trained for around steps, and “P2R” has been trained for steps. The other is “training for longer”, which demonstrates a wall-time performance of M6-1B for better comparison.
3.2 Downstream Evaluation
We compare the model performance by the downstream evaluation on language modeling and text summarization tasks, which cover capabilities including language modeling and generation. To be more specific, the downstream datasets include:
WikiText103: a classical language modeling evaluation benchmark dataset that consists of Wikipedia articles. We evaluate the Perplexity (PPL) of the pretrained model, which is a per-token exponential cross-entropy loss that reflects probability distribution over texts.
Gigaword: a dataset for summarization which consists of around 3.8M articles and summaries, and an effective benchmark to evaluate the model’s capability in text generation for abstractive summarization.
For better comparison, we also present the other pretrained models as baselines to demonstrate that the models can achieve strong performance in downstream tasks. Furthermore, we also compare the model quality on the time basis to reflect the advantage of Pseudo-to-Real training strategy.
To view their training efficiency, we focus on their training speed on the condition that we attempt to exhaust the GPU utility. Table 1 demonstrates the training speed of the models. We report their training speed on NVIDIA-V100 GPU devices with their consumed samples per second. Obviously training Real Giant model from scratch is highly time-consuming, and Pseudo Giant training has an advantage of around times of convergence speed over Real Giant training.
We also observe the loss convergence of both P2R and Real Giant trained from scratch. In Figure 2 we present the development of pretraining language modeling loss, which is the log perplexity, on the time basis. The log perplexity of P2R decreases much faster than that of the Real Giant, with an advantage larger than .
We add other strong baseline pretrained models for the two tasks for comparison. M6-base model matches the performance of pretrianed models, including UniLM, Megatron-LM GPT, etc. [12; 40]. Experimental results are consistent with our hypothesis that Pseudo-to-Real training can speed up training effectively. In the setup of limited budget, the M6 model trained with Pseudo-to-Real can outperform the Pseudo Giant and Real Giant in both language modeling and text generation. We also add an M6-1B model trained for a longer time to show its wall-time performance on downstream tasks. However, we did not reach the best performance in Gigaword. As it is not related to the training strategy, we leave this issue to future work for further discussion about how to finetune large models on downstream tasks.
Towards a 10-Trillion-Parameter Model
Previously, training large-scale models brings tons of challenges to the collaboration algorithm design, distributed training, as well as hardware design, etc. Training a GPT-3 of B parameters with a combination of data parallelism on over GB of data should cost around GPU-years. Later with the emergence of partitioning on optimizer states, gradients, and even weights, GPU memory can be fully utilized without performance degradation. Now we can even use the CPU memory or even NVMe storage to store the parameters, but we have to bear the costs of efficiency. Therefore, we attempt to tackle the difficulty of extreme-scale model training from the perspective of algorithm design and thus we apply the aforementioned Pseudo-to-Real training strategy to train an extreme-scale model.
We design a 10-trillion-parameter M6 model with the combination of existing methods and proposed strategies to demonstrate a case of how to train an extreme-scale model efficiently. In comparison with the previous studies of trillion-parameter models [13; 45], this one is almost 10 times larger. To efficiently utilize the memory, we adopt Mixture-of-Experts and we replace every FFN layer with the memory-efficient MoE layers. Notably, we remove the auxiliary loss that consumes memory and demonstrates little effects on model quality, and we follow Yang et al. to apply expert prototyping for improved model quality and training stability. To be more specific, the hidden size is and the intermediate size is . The number of model layers is . For each MoE layer, there are experts distributed on multiple devices. We use prototypes of experts based on our experience in preliminary experiments. The training batch size per GPU device is . We implement models on EFLOPs, a distributed training platform with an advanced server architecture and a new network topology . Specifically our models are trained on a cluster of 8-GPU workers connected by RDMA networks with a bandwidth of 100Gb. The CPU memory of each worker is around 750GB. The expert distribution and Granular Offloading strategies are supported by Whale framework.
2 Granular CPU offloading
To utilize CPU memory for larger models with fewer resources, we apply the Granular CPU offloading which has a higher efficiency compared with the conventional CPU offloading. Previously we note that conventional offloading mechanisms offload all parameters, which may fail to effectively utilize GPU. To improve the efficiency of CPU offloading, we propose a new CPU offloading mechanism called Granular CPU offloading. The training process is composed of phases including “Forward (Fn)”, “Backward(Bn)” and “Apply (An)”. Offloading all model parameters to CPU in Fn and Bn requires loading parameters from CPU to GPU memory twice. Activation checkpointing that brings recomputation needs the parameters loaded in Bn. An requires the gradients offloaded from GPU memory to CPU memory. Assume the model parameter size is , the above processes bring parameter movement of size .
In offloading, PCIE is the bottleneck of the whole training process. We observe that when offloading all parameters with recomputation, the GPU memory is idle. We can fill up the GPU memory by selective offloading, leaving part of the model in GPU memory. In this way, the model can be accelerated by reducing across-device memory copy. In our preliminary experiment, with the setting of training a 48-layer 78B-parameter M6 model on 8 NVIDIA-V100 GPU devices, the step-time costs 89 seconds when fully offloading the parameters into CPU memory. In comparison, granularly offloading the first 24 layers into CPU memory and leaving the remaining 24 layers on GPU reduces the step-time to only 45 seconds. The significant difference in training step-time indicates that the time-cost of parameter movement between CPU and GPU will dominate the training step-time when offloading is employed, thus the granularity of offloading and the utilization of GPU should be seriously considered in extreme-scale pretraining. In addition, offloading the whole model can result in OOM error in the CPU when the model is extremely large.
With Granular CPU offloading, we successfully implement a 10-trillion-parameter M6 model on solely 512 NVIDIA-V100 GPUs. Furthermore, at the Pseudo stage, we can train a Pseudo Giant with a computation graph of 10 trillion parameters only with 256 GPU devices without the utilization of CPU memory for offloading. Thus in our practice, we train a Pseudo Giant with only 256 GPUs, and then partition the experts and redistribute them to 512 GPUs. This saves the usage of GPU resources, which is more resource-efficient and also more environmentally friendly.
3 Analysis
We pretrain the M6-10T with Pseudo-to-Real training strategy for around days, and we additionally train an M6-10T from scratch without the strategy for around days for comparison. The Real stage training can be facilitated without Out-of-Memory errors with the help of Granular CPU offloading, but its step-time is only around s. In contrast, the step-time of Pseudo stage is only s without the cost of offloading, which greatly boosts the training efficiency of M6-10T P2R. The M6-10T with P2R has been trained for steps, but the one from scratch has been trained for solely steps. We record the log perplexity of both models trained on the M6-Corpus on the time basis in Figure 4(a). Results show that within the same length of time M6-10T with P2R can outperform the one trained from scratch by a large margin.
We have trained M6-10T with Pseudo-to-Real strategy for around days, and the model converges to a low level of log perplexity based on the upstream evaluation on the M6-Corpus. For better comparison, we also show the convergence performance of the 1T-parameter M6-T model proposed in the previous work. As shown in Figure 4(b), the observation is consistent with our intuition that the model with a larger capacity can converge faster on the sample basis, and it should achieve the best performance on language modeling. What leaves open is whether it can positively lead to better downstream performance concerning different types of downstream tasks. Finetuning extremely large models should be difficult and there is still much room for us to discover the potential of extreme-scale models. However, the contribution of this work is leveling up their training efficiency, which can be regarded as an initial step to investigate the secrets of super large models.
Conclusion and Future Work
Pseudo-to-Real training strategy is a simple and effective way to train large-scale models that are highly memory consuming, and it is also essential to training extremely large models with limited resources with significantly higher training efficiency. We unlock pretraining unprecedented extreme-scale models with 10 trillion parameters with limited resources of 512 GPUs in days. Besides the application of Pseudo-to-Real training strategy, we further provide Granular CPU offloading to enhance GPU utility while breaking the GPU memory wall with a cost in efficiency. The advances take a leap towards extreme-scale model training beyond implementation on limited resources. With only a few GPU cards, training large models with tens or hundreds of parameters has become accessible to many researchers. We believe that this can motivate low carbon dioxide production and encourages the progress of green AI.
Ethics Statement
This work is highly concerned with large-scale language models and multimodal pretrained models. These models have been pretrained on broad data of plain texts and image-text pairs, which might contain harmful information, such as hate speech, terrorism, pornography, etc. We have put much efforts to remove these kinds of data in our datasets by quality evaluation on texts and images. However, this problem cannot be eliminated and ignored, and it is common in the pretraining community. For those models that are not trained on commonly-used public datasets, we will carefully release the model checkpoints before careful evaluation, and also limit the access to avoid misconduct.
Reproducibility Statement
This work is generally reproducible. Following the description in Section 3, researchers can easily implement the training strategy on the codebases for pretraining, including Huggingface Transformer https://github.com/huggingface/transformers, Fairseq https://github.com/pytorch/fairseq.