Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers

Zhewei Yao, Xiaoxia Wu, Conglong Li, Connor Holmes, Minjia Zhang, Cheng Li, Yuxiong He

Introduction

Large-scale transformers have been demonstrated to have supreme performance on natural language processing (tenney2019bert, radford2019gpt, colin2019t5), computer vision (dosovitskiy2020image), and other applications (gong2021ast, guo2021pct). However, both the pretraining procedure and some downstream finetuning tasks (e.g., long document summary) are time-consuming and resource-hungry. Thus, there is a need to speed up the training and reduce the compute cost for large-scale transformer pretraining and finetuning.

Recently, hou-etal-2022-token adopt the token pruning/dropping/bypassing technique (kim2021learned, goyal2020power, kim2020length) from BERT inference to BERT pretraining by skipping the compute of part of the input tokens at some middle layers. The results of (hou-etal-2022-token) (referred to as TokenBypass) show that it can theoretically reduce the pretraining cost by 25% for both BERTbase{}_{\text{base}} and BERTlarge{}_{\text{large}} without losing accuracy on finetuning tasks. Although achieving great speedup, TokenBypass (1) needs an import-score metric to determine the dropped tokens and special token treatment to keep important tokens (e.g., [CLS]), both of which require manual designs; (2) has to keep the first half layers and the last layer (in total, half of the depth) in full sequence length training, which limits its layer-bypassing ability. (3) solely focuses on BERT Masked-LM pretraining tasks and has not been applied to other tasks, e.g., causal-LM. In this work, we address those challenges and introduce our random and layerwise token-dropping method (random-LTD). In summary, our contributions are as follows:

All tokens are treated equally without any special token treatment or import-score measurement, i.e., no manual design, and are dropped in a purely random manner. Meanwhile, instead of fully bypassing the dropped token for all middle layers (hou-etal-2022-token), each layer in random-LTD drops tokens independently from the other layers. This helps the multi-head attention in the middle layers capture the dependency relation across different tokens suggested in (vig2019analyzing).

random-LTD applies token dropping at all middle layers except the very first and last layers, which further reduces manual design and increases training efficiency. We also propose a new monotonic sequence length growth method as training evolves to (1) reduce the gradient noise introduced by random-LTD for better convergence and (2) close the training and inference (autoregressive generation) gap, since random-LTD breaks the autoregressive manner in middle layers during training, for GPT models.

To reduce the tuning effort for the newly proposed training procedure, we introduce a new LayerToken learning rate schedule, which scales the learning rate based on the sum of consumed tokens of each layer for pretraining tasks.Note that the numbers of consumed tokens for different layers are different. We show its superb performance for random-LTD on GPT/BERT pretraining compared to the standard iteration-based learning rate schedule.

We extensively test random-LTD on both pretraining tasks, including GPT and BERT pretraining, and finetuning tasks, including causal-LM finetuning for GPT and image classification for ViT. For all tasks, random-LTD achieves similar accuracy as the original baseline method with up to 33.3% theoretical cost saving and up to 25.6% wall-clock time saving.

Finally, we show that random-LTD has a potential regularization effect, which can be used for both pretraining and finetuning problems.

Background

Token dropping (or token bypassing/pruning) (kim2021learned, goyal2020power, kim2020length, press2021train, wang2021spatten) was originally proposed for BERT inference to reduce the computational overhead. In this case, if a token ii (Xj,iX_{j,i}) is decided to be dropped at layer jj (LjL_{j}), the compute cost of this token through all remaining layers (LkL_{k} where k>jk>j) is eliminated. As such, the sequence length sis_{i} of the ii-th layer’s input Xi−1X_{i-1} will be a non-increasing array, i.e., s0≥s1 ... ≥sls_{0}\geq s_{1}~{}...~{}\geq s_{l}. However, such a configuration has been shown instability for adaptive token-dropping inference (kim2020length). Therefore, kim2020length utilize the sandwich rule and distillation from (yu2019universally) to stabilize training and boost accuracy. But these two methods also significantly increase the training cost. Thus, such techniques cannot be applied to speed up the pretraining procedure. Recently, hou-etal-2022-token extended token dropping from inference to BERT pretraining (referred to as TokenBypass). hou-etal-2022-token use several importance scores/metrics to determine the dropped tokens, e.g., cumulative loss and frequency of each token. To overcome the training instability issue, the authors proposed two main mechanisms: (1) the sandwich token dropping rule, where the first (layer 1 to i)i) and the last few layers (layer Ll−jL_{l-j} to LlL_{l}) of the BERT capture all tokens (i.e., no token dropping) and the middle layers bypass s′≤ss^{\prime}\leq s tokens from LiL_{i} to Ll−jL_{l-j}. Particularly, the authors (only) test on the encoder transformer (12-layer BERTbase{}_{\text{base}} and 24-layer BERTlarge{}_{\text{large}}), and let i=l/2−1i=l/2-1, j=1j=1, s′=s/2s^{\prime}=s/2. (2) special token treatment, where special tokens (e.g., [MASK], [CLS], [SEP]) are never dropped.

Compared to TokenBypass from (hou-etal-2022-token), our random-LTD (1) does not require importance score metric, special token treatment, or the sandwich token dropping rule, which dramatically reduces the manual design effort; (2) has been broadly tested on pretraining tasks, including GPT and BERT, as well as finetuning tasks, including ViT classification and GPT causal-LM. Meanwhile, we found out that directly applying TokenBypass to causal-LM leads to severe accuracy degradation. Please see the detailed description of random-LTD in Section 3 and our extensive evaluation in Section 4 and 5. We also include a thorough discussion of other efficient training methods in Appendix LABEL:sec:other_efficient_training_approaches.

Methodology

Layerwise Token Dropping Mechanism. As pointed out in Section 2, existing inference and training token dropping methods either permanently drop tokens from the compute graph at intermediate layers, or at least make some tokens fully skip a consecutive series of middle layers. However, several works (vig2019analyzing, michel2019sixteen, voita2019analyzing) have shown that MHA focuses on different tokens at different layer depths and the attention map aligns with the dependency relation most strongly in the middle of transformer architectures. Therefore, TokenBypass used in hou-etal-2022-token, i.e., fully skipping middle layers, may hinder the learnability/generalization of the architecture during pretraining/inference. We conjecture that this might be why multiple first/last layers need to be kept and the special token treatment is needed in (hou-etal-2022-token). To further verify if this fully skipping middle layer mechanism (hou-etal-2022-token) causes any learnability issue, we apply TokenBypass on GPT finetuning tasks and observe much lower performance as compared to baseline. See more details in Section 5.1.

Random Token Dropping. Various important score-based metrics are used to determine the token dropping criterion. Most of them can be categorized in two ways: attention score related metrics or loss/frequency-based metrics. However, both of them introduce challenges that make LTD less practical. Particularly, for attention score-based metrics, the compute cost for LTD is too high since the metric has to be calculated for every layer; for loss-/frequency-based metrics, generally accumulated loss or frequency is used and this accumulated metric would not be changed within the same iteration (a.k.a. one forward pass of the network). Therefore, the unchanged loss/frequency metric leads the dropped token to be the same for different layers, making the token dependency not be captured by the MHA of middle layers (vig2019analyzing).

To satisfy the independent requirement of LTD, we propose to use purely random token dropping assignment. For each transformer layer, we randomly (uniformly) select a small batch of tokens to proceed with the compute and drop the rest. In more details, assume Mi=M_{i}={mi(1)m_{i}(1), mi(2)m_{i}(2), …, mi(s)m_{i}(s)} is a random shuffle of S=S={1, 2, …, s}. Then the dropped token set is Ji=J_{i}={mi(1)m_{i}(1), mi(2)m_{i}(2), …, mi(ai)m_{i}(a_{i})} for the input of Li+1L_{i+1}.

Random and Layerwise Token Dropping. Combining layerwise token dropping with random token dropping, we have our final random and layerwise token dropping method (random-LTD), which can efficiently apply token dropping for each individual layer and can capture the attention dependency of each token with other others in middle layers with high probability.

The illustration of the comparison between standard baseline training and random-LTD is shown in Fig. 2 (an additional comparison with (hou-etal-2022-token) in Fig. LABEL:fig:illustration_of_random_ltd_and_baseline_and_tokenbypass). The pseudo-code is given in Fig. 2. For each layer, as compared to the baseline, random-LTD randomly selects (function “gather” in Fig. 2) a subset of the tokens and feeds (function “Layer” in Fig. 2) them into the transformer layer. Afterward, we combine (function “combine” in Fig. 2) the output of transformer layer with the dropped tokens to recover the full sequence length. Thus, the next layer still receives the full sequence and can repeat this process.

Since the dropped tokens of each layer are independent, there is no need for random-LTD to treat special tokens (e.g., [MASK], [CLS], [SEP], [PADDING]) differently from other normal tokens, which can further reduce the cost of computing the dropping criterion. Meanwhile, we show that special token treatment does not bring extra benefits for random-LTD on BERT pretraining in Section 5.2.

2 Dropping Schedule of random-LTD

Layers without Token Dropping. While TokenBypass (hou-etal-2022-token) needs to keep half of the layers in full sequence length training, random-LTD has no such limitation. Thanks to the attention-capture feature of random-LTD, we can apply random-LTD to most of the transformer layers except the first and last transformer layers.

Keeping the first and last layers in full sequence length training usually leads to better performance since (1) the first layer directly connects to the embedding, and it can help refine the raw feature; (2) the last layer directly connects to the final prediction; a feature realignment for all tokens can improve the model quality. We also provide a detailed study to show the importance of keeping the first and last layers without token dropping in Section 5.3.

In order to reduce the gradient variance introduced by random-LTD for better training, we monotonically increase the kept sequence length throughout training (referred to as MSLG) with a linear schedule. Particularly, the dropped token set JiJ_{i} for the ii-th layer gradually shrinks and the kept token set KiK_{i} gradually grows as the training proceeds. Denote the size of JiJ_{i} (KiK_{i}) at step tt is ai,ta_{i,t} (bi,tb_{i,t}), its final size is (ss), and the total training iterations is TT. Assume we want to gradually reduce the size of JiJ_{i} to zero at iteration T′T^{\prime} and the decreasing strength is sdecs_{dec}. Then the decreasing step size is Tdec=T′/(a0,t/sdec)T_{dec}=T^{\prime}/({a_{0,t}}/{s_{dec}}), i.e., for every TdecT_{dec} iterations, the size of JiJ_{i} (KiK_{i}) reduces (increases) by sdecs_{dec}. Please see Fig. 3 for an illustration of KiK_{i} on GPT pretraining. We also show that MSLG outperforms the constant drop schedule with similar compute savings in Section 5.4.

3 New Learning Rate Schedule for Pretraining

When performing pretraining on language models, we oftentimes use a decaying learning rate schedule based on iteration with a warmup period. Particularly, at the first few thousand or hundred iterations, warming up the learning rate is critical for distributed pretraining tasks due to its instability (goyal2017accurate, li2021curriculum). However, an iteration-based schedule is not optimal for random-LTD.

First, random-LTD reduces the effective batch size of middle layers at the initial warmup phase. The effective training tokens for dropped token layers become much smaller than the baseline training. Second, for most of our training cases, MSLG does not reach the full length until >2/3>2/3 of training iterations for large compute saving. At such time, the iteration-based learning rate is considerable small. And this small learning rate cannot provide efficient training dynamics for random-LTD. Therefore, to stabilize the initial training phase and to have a large enough learning rate in the later training phase, we need to increase the warmup iterations and slow down the learning rate decay. Here, we propose a new learning rate schedule based on the layerwise tokens consumption, called layer-token learning rate (LayerToken LR). Please see Appendix LABEL:sec:formal_layertokenlr_description for the formal and detailed description of LayerToken LR.

We emphasize that one can always tune the learning rate schedule by increasing the maximum learning rate or the warmup iterations. However, it would require a lot of engineering effort. Therefore, we propose this LayerToken LR schedule, which is more suitable for our random-LTD than the standard one. We also include a detailed comparison between the standard learning rate schedule and LayerToken LR in Section 5.5.

Main Results

In this section, we first provide the results of random-LTD for pretraining on GPT and BERT models. We then extend random-LTD on the computer vision domain to demonstrate its broader applications. Similar to Section 3.3 and Appendix LABEL:sec:formal_layertokenlr_description, we use the LayerToken compute the cost to measure the total training budget.Similar to (hou-etal-2022-token), we do not include (1) the final prediction layer and (2) the attention compute difference between the different lengths of sequence for the compute cost comparison. We also provide the real training time saving for GPT and BERT pretraining. Kindly note that the real-time saving depends on various factors, e.g., the implementation and hardware.

We train GPT-3-style models with 350 million parameters (GPT-3350M{}_{\text{350M}}) and 1.3 billion parameters (GPT-31.3B{}_{\text{1.3B}}) on PILE dataset (gao2020pile) and the total number of training tokens is 300 billion. For random-LTD, the initial dropped token length for all middle layers is 1920 (i.e., 128 tokens are kept for compute), and it decreases by 16 for every 1.75B training tokens. After 210B training tokens, random-LTD degrades to standard training procedure with full sequence length. Theoretically, this can save 1/3 of the LayerToken training budget. See Appendix LABEL:sec:gpt_experimental_setup for more details.

The evaluation loss curves for baseline and random-LTD are shown in Fig. 3. As can be seen, for both GPT-3350M{}_{\text{350M}} and GPT-31.3B{}_{\text{1.3B}}, random-LTD has similar evaluation losses as the baseline with 1/3 less LayerToken consumption. We also provide the zero-shot evaluation results in Tab. 4.1. For both GPT-3350M{}_{\text{350M}} and GPT-31.3B{}_{\text{1.3B}}, random-LTD achieves comparable results as the baseline, Besides, random-LTD can save 14.3% wall-clock training time on GPT-3350M{}_{\text{350M}} and 25.6% wall-clock training time on GPT-31.3B{}_{\text{1.3B}}.

We reiterate that the LayerToken consumption saving ratio cannot directly transfer to GPU wall-clock training time saving ratio due to the implementation/hardware. Meanwhile, note that the saving number we reported here is not the maximum potential saving (with fixed implementation/hardware, etc) since we can reduce the training GPU numbers for random-LTD at the initial training phase, which has a shorter effective training sequence length. Also, although random-LTD has the same theoretical compute saving for both GPT-3350M{}_{\text{350M}} and GPT-31.3B{}_{\text{1.3B}}, the real wall-clock time saving varies a lot because GPT-31.3B{}_{\text{1.3B}} has a larger hidden dim size, which means the model spends more time on real computing than other operators, e.g., data movement and gradient communication.

Model Method LayerToken Saving GPU Cost (saving) Ave. GPT-3350M{}_{\text{350M}} baseline None 64×\times2.59 (0.0%) 38.6 random-LTD 33.3% 64×\times2.22 (14.3%) 38.9 \cdashline1-15 GPT-31.3B{}_{\text{1.3B}} Baseline None 64×\times5.42 (0.0%) 42.7 random-LTD 33.3% 64×\times4.03 (25.6%) 42.5

Method LayerToken Saving GPU Cost (Saving) Ave. baseline None 64×\times5.89 (0.0%) 85.42 \cdashline1-9 random-LTD-1 26.2% 64×\times5.42 (7.95%) 86.95 random-LTD-2 31.1% 64×\times5.21 (11.5%) 86.42

We pretrain BERTlarge{}_{\text{large}} on PILE dataset for 2M iterations with batch size 1024 and sequence length 512 following (shoeybi2019megatron). We apply random-LTD with two variants, random-LTD-1 and random-LTD-2. Particularly, for random-LTD-1 (random-LTD-2), the initial kept token length is 200 (128), and it increases by 16 for every 48B (38B) training tokens. As such, we save 26.2% (31.1%) LayerToken consumption for random-LTD-1 (random-LTD-2). We evaluate the trained model on four downstream tasks as (shoeybi2019megatron), i.e., MNLI, QQP, RACE-m, RACE-h. Note that we apply standard finetuning without token dropping to have a fair comparison for both pretrained models from the baseline and random-LTD. Please see Appendix LABEL:sec:bert_experimental_setup for more details.

Tab. 4.1 summarizes the results along with the full results in Tab. LABEL:table:bert-std. Although random-LTD is slightly worse than baseline on a certain task (QQP, see Tab. LABEL:table:bert-std), it gives much higher accuracy on other tasks while saving 26.2–31.1% of the theoretical computation overhead in pretraining. Overall, random-LTD-1 achieves 1.54 points higher average accuracy over baseline and random-LTD-2 achieves 1 point higher average accuracy over baseline.

Meanwhile, random-LTD-1 (random-LTD-2) saves about 7.95% (11.5%) wall-clock time as compared to baseline. Note that similar to GPT pretraining, the saving depends on the implementation/hardware, and random-LTD has potentially larger savings if elastic training is performed. Also, although GPT-3350M{}_{\text{350M}} and BERTlarge{}_{\text{large}} have similar model sizes as well as similar theoretical compute saving, the final wall-clock training time saving varies by about 3%. This is caused by BERTlarge{}_{\text{large}} having a shorter final sequence length (i.e, 512) than GPT (i.e, 2048), and the real compute time for a sentence with sequence length 128 is not 1/4 (or 1/16) of a sentence with 512 (2048) tokens. This leads the overall compute time-saving for GPT-3350M{}_{\text{350M}} to be larger than that for BERTlarge{}_{\text{large}}.

3 ViT Finetuning

We perform the vision transformer (ViT) on both ImageNet (with a 12-layer pretrained ViT) and CIFAR (with a 24-layer pretrained ViT). For random-LTD, the initial sequence length is 66 and linearly reaches the 197 full sequence length at 80% of the total training iterations such that 22.3% layer-token saving is achieved. See training details in Appendix LABEL:subsec:fine-tune-vit. We summarize the result with standard deviation in Tab. 4 along with the full details in Tab. LABEL:table:vit-main-full. As can be seen, random-LTD can achieve comparable results as the baseline on all three datasets. This demonstrates the broader applications of random-LTD.

Discussion

In this section, we present several import ablation studies and the potential regularization effect of random-LTD. Besides the three tasks used in previous sections, we also include GPT finetuning on causal-LM problems using the GPT-2350M{}_{\text{350M}} from Huggingface (wolf2019huggingface). Please see Appendix LABEL:subsec:fine-tune-gpt for the training details. Also, we reduce the iterations of BERT pretraining from 2M to 200k due to resource limitations. Please see Appendix LABEL:sec:bert_experimental_setup for more details.

Although TokenBypass (hou-etal-2022-token) demonstrates its great ability on BERT pretraining, its skipping policy may still hurt the performance of other tasks, e.g., causal-LM. The reason is mentioned in Section 3, i.e., MHA of middle layers focuses on different tokens at different depths. Fully skipping those layers may lead the causal-LM task to lose attention capability. However, random-LTD does not have this issue as it randomly selects kept tokens for each layer.

To verify this, we provide an ablation study on the comparison between random-LTD and TokenBypass with GPT-2350M{}_{\text{350M}} finetuning on PTB (marcus-etal-1993-building). We make two sets of experiments:

Set 1. Following (hou-etal-2022-token), we bypass half of the tokens based on their empirically moving average loss from L12L_{12} to L23L_{23}. Similarly, we apply constant random token drop to the layers from the middle to the second last (L12L_{12} to L23L_{23}).

Set 2. We apply TokenBypass or constant random token dropping for half of the tokens starting from the second layer (L2L_{2}) until the second last layer (L23L_{23}).

The validation curves of two cases are shown in Fig. 4. As can be seen, for both cases, random-LTD performs much better than TokenBypass. Particularly, for the Set 2 comparison, the perplexity of random-LTD is about 10 points lower than TokenBypass, demonstrating that random-LTD can be applied to more layers than TokenBypass.

2 With/Without Special Token Treatment

Different from the GPT pretraining task, which has consecutive sentences/paragraphs, the BERT pretraining data consists of two sentences that could be unrelated. The special tokens [CLS] and [SEP] play critical roles for the model in determining the beginning/end of each sentence so that the model can predict the relationship between the two sentences. Thus, there could be potential gain in keeping the tokens for all layers, which has a more detailed discussion in (hou-etal-2022-token).

Here we present an ablation study on whether keeping those special tokens helps random-LTD or not. We perform a straight comparison for random-LTD: (1) one with purely random selection and (2) the other with an additional criterion, i.e., keeping the special tokens for all layers. See training details in Appendix LABEL:sec:bert_experimental_setup

The results of MNLI/QQP are shown in Tab. 4 along with the full details in Tab. LABEL:table:bert-ablation-study-full. As can be seen, for both pretraining loss and downstream finetuning, special token treatment does not provide any benefit for random-LTD. Note that random-LTD without special token treatment is also more compute-friendly for token dropping.

3 Why we need to keep the first and last layer?

To understand why keeping the first and last layers in full sequence length training, we present single layer sensitivity analysis shown in Fig. 5 for GPT-2350M{}_{\text{350M}} finetuning on Wikitext-2 and Wikitext-103. Particularly, we apply constant token dropping for one layer and keep all the other layers in the standard training mode. After the training, we measure the PPL and use it as the sensitivity metric, i.e., higher PPL indicates high sensitivity and vice versa. The U-shape of both curves implies that the first and last ones are most sensitive to token dropping.

To further understand if random-LTD can be applied to all layers when using MSLG, we include other three scenarios, i.e., applying random-LTD to (1) all but not last layer, (2) all but first layer, and (3) all layers. We perform finetuning tasks on both causal-LM and image classification. See the full training details in Appendix LABEL:subsec:fine-tune-vit and Appendix LABEL:subsec:fine-tune-gpt. From Tab. 6 and LABEL:table:main-layers-full, we can clearly see that keeping the first and the last layers intact leads to a substantial improvement (beyond standard deviation) over the rest three scenarios.

4 Why we need sequence length growth?

We now give an ablation study on why MSLG schedule is necessary. Again, We perform finetuning tasks on both causal-LM and image classification.

We specially set a constant token dropping rate that matches the token saving of MSLG schedules with all other hyperparameters fixed. We present the results in Tab. 6 and LABEL:table:seq-schedules-full. It can be clearly seen that given the almost same amount of LayerToken saving (33%−35%33\%-35\%), the constant dropping schedule has worse performance than MSLG. MSLG schedule can actually be even better or comparable to those constant ones whose saving is 10%10\% smaller.

5 LayerToken LR Schedule effect

To study the effectiveness of LayerToken LR, we compare three training scenarios for GPT-3350M{}_{\text{350M}} with 300B training tokens (see Appendix LABEL:sec:gpt_experimental_setup for training details): (1) the baseline training with the standard learning rate, (2) random-LTD with the standard learning rate, and (3) random-LTD with LayerToken LR. The validation curves and their corresponding learning rates with respect to iterations are plotted in Fig. 6. As can be seen, the green curve (random-LTD with LayerToken LR) can achieve comparable validation loss as the baseline, which is better than random-LTD with the standard learning rate. This confirms that the small learning rate introduced by the standard learning rate schedule slows the learning of random-LTD at the later training phase. A similar observation is made for the BERT pretraining, which will be deferred to Appendix LABEL:sec:bert-lr due to the space limit.

6 Comparison on GPT-31.3B1.3B{}_{\text{1.3B}} with Various Training Budgets

In this section, we perform various training budgets to train GPT-31.3B{}_{\text{1.3B}} and compare the performance of baseline and random-LTD to verify if random-LTD can consistently save the training cost.

To fully understand the effectiveness of random-LTD, we train GPT-31.3B{}_{\text{1.3B}} using a baseline with 120B (i.e., 2880 billion LayerToken consumption) to 360B tokens (i.e., 8640 billion LayerToken consumption). For random-LTD, we follow our Section 4 setting to save 1/3 of the total training budget and apply it to three training budgets, which are equivalent to 2880, 4800B, and 5760B LayerTokens.

The results are summarized in Tab. 5.7. The first noticeable result here is that the total amount of training LayerTokens directly affects the model’s final quality. Longer training usually leads to better accuracy. Meanwhile, random-LTD can use 2/3 of the training budget to achieve similar evaluation results as the baseline. For example, random-LTD with training budgets 2880, 4800, and 5760 billion LayerTokens can lead to similar accuracy as the baseline with training budgets 4320, 7200, and 8640 billion LayerTokens.

7 The interplay between random-LTD and Dropout

Token dropping can be viewed as a special (coarse-grained) case of dropout. As such, it may have the potential to work as dropout for certain tasks or further help dropout for certain tasks. To investigate this, we apply random-LTD with/without standard dropout on both pretraining and finetuning tasks.

BERTlarge{}_{\text{large}} Pretraining. Please see Appendix LABEL:sec:bert_experimental_setup for training details. For random-LTD-3 (random-LTD-4), the initial kept token length for all middle layers is 256, and it increases by 16 for every 3.8 (6) billion training tokens such that we eventually save 14.1% (22.3%) LayerToken consumption.

The results are summarized in Tab. 5.7 with details in Tab. LABEL:table:bert-ablation-study-full and we include the pretraining perplexity to better understand the dropout effect. Clearly, turning off dropout results in lower evaluation/test perplexity (shown in Tab. 5.7). Meanwhile, the performance of those no-dropout models (baseline*, random-LTD-3*, random-LTD-4*) on MNLI and QQP show an obvious improvement over their dropout counterparts (baseline, random-LTD-3, random-LTD-4). However, for RACE finetuning, there is no learning for the no-dropout baseline model, which somewhat is surprising but shows the importance of dropout for pretraining. In contrast, when turning off the dropout for random-LTD, we see a compelling better accuracy on RACE, which exceeds the standard baseline pretraining by >1% on RACE-m. Thus, random-LTD brings not only the efficiency but also the potential regularization effect to BERT pertaining.