MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers

Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, Furu Wei

Introduction

Pretrained Transformers Radford et al. (2018); Devlin et al. (2018); Dong et al. (2019); Yang et al. (2019); Joshi et al. (2019); Liu et al. (2019); Bao et al. (2020); Radford et al. (2019); Raffel et al. (2019); Lewis et al. (2019a) have been highly successful for a wide range of natural language processing tasks. However, these models usually consist of hundreds of millions of parameters and are getting bigger. It brings challenges for fine-tuning and online serving in real-life applications due to the restrictions of computation resources and latency.

Knowledge distillation (KD; Hinton et al. 2015, Romero et al. 2015) has been widely employed to compress pretrained Transformers, which transfers knowledge of the large model (teacher) to the small model (student) by minimizing the differences between teacher and student features. Soft target probabilities (soft labels) and intermediate representations are usually utilized to perform KD training. In this work, we focus on task-agnostic compression of pretrained Transformers Sanh et al. (2019); Tsai et al. (2019); Jiao et al. (2019); Sun et al. (2019b); Wang et al. (2020). The student models are distilled from large pretrained Transformers using large-scale text corpora. The distilled task-agnostic model can be directly fine-tuned on downstream tasks, and can be utilized to initialize task-specific distillation.

DistilBERT Sanh et al. (2019) uses soft target probabilities for masked language modeling predictions and embedding outputs to train the student. The student model is initialized from the teacher by taking one layer out of two. TinyBERT Jiao et al. (2019) utilizes hidden states and self-attention distributions (i.e., attention maps and weights), and adopts a uniform function to map student and teacher layers for layer-wise distillation. MobileBERT Sun et al. (2019b) introduces specially designed teacher and student models using inverted-bottleneck and bottleneck structures to keep their layer number and hidden size the same, layer-wisely transferring hidden states and self-attention distributions. MiniLM Wang et al. (2020) proposes deep self-attention distillation, which uses self-attention distributions and value relations to help the student to deeply mimic teacher’s self-attention modules. MiniLM shows that transferring knowledge of teacher’s last layer achieves better performance than layer-wise distillation. In summary, most previous work relies on self-attention distributions to perform KD training, which leads to a restriction that the number of attention heads of student model has to be the same as its teacher.

In this work, we generalize and simplify deep self-attention distillation of MiniLM Wang et al. (2020) by using self-attention relation distillation. We introduce multi-head self-attention relations computed by scaled dot-product of pairs of queries, keys and values, which guides the student training. Taking query vectors as an example, in order to obtain queries of multiple relation heads, we first concatenate query vectors of different attention heads, and then split the concatenated vector according to the desired number of relation heads. Afterwards, for teacher and student models with different attention head numbers, we can align their queries with the same number of relation heads for distillation. Moreover, using a larger number of relation heads brings more fine-grained self-attention knowledge, which helps the student to achieves a deeper mimicry of teacher’s self-attention module. In addition, for large-size (2424 layers, 10241024 hidden size) teachers, extensive experiments indicate that transferring an upper middle layer tends to perform better than using the last layer as in MiniLM.

Experimental results show that our monolingual models distilled from BERT and RoBERTa, and multilingual models distilled from XLM-R outperform state-of-the-art models in different parameter sizes. The 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 (66 layers, 768768 hidden size) model distilled from BERTLARGE{}_{\text{LARGE}} is 2.0×2.0\times faster, meanwhile, performing better than BERTBASE{}_{\text{BASE}}. The base-size model distilled from RoBERTa\textsclarge{}_{\textsc{large}} outperforms RoBERTa\textscbase{}_{\textsc{base}} using much fewer training examples.

We generalize and simplify deep self-attention distillation in MiniLM by introducing multi-head self-attention relation distillation, which brings more fine-grained self-attention knowledge and allows more flexibility for the number of student’s attention heads.

We conduct extensive distillation experiments on different large-size teachers and find that using knowledge of a teacher’s upper middle layer achieves better performance.

Experimental results demonstrate the effectiveness of our method for different monolingual and multilingual teachers in base-size and large-size.

Related Work

Multi-layer Transformer Vaswani et al. (2017) has been widely adopted in pretrained models. Each Transformer layer consists of a self-attention sub-layer and a position-wise fully connected feed-forward sub-layer.

2 Pretrained Language Models

Pre-training has led to strong improvements across a variety of natural language processing tasks. Pretrained language models are learned on large amounts of text data, and then fine-tuned to adapt to specific tasks. BERT Devlin et al. (2018) proposes to pretrain a deep bidirectional Transformer using masked language modeling (MLM) objective. UniLM Dong et al. (2019) is jointly pretrained on three types language modeling objectives to adapt to both understanding and generation tasks. RoBERTa Liu et al. (2019) achieves strong performance by training longer steps using large batch size and more text data. MASS Song et al. (2019), T5 Raffel et al. (2019) and BART Lewis et al. (2019a) employ a standard encoder-decoder structure and pretrain the decoder auto-regressively. Besides monolingual pretrained models, multilingual pretrained models Devlin et al. (2018); Lample and Conneau (2019); Chi et al. (2019); Conneau et al. (2019); Chi et al. (2020) also advance the state-of-the-art on cross-lingual understanding and generation.

3 Knowledge Distillation

Knowledge distillation has been proven to be a promising way to compress large models while maintaining accuracy. Knowledge of a single or an ensemble of large models is used to guide the training of small models. Hinton et al. (2015) propose to use soft target probabilities to train student models. More fine-grained knowledge such as hidden states Romero et al. (2015) and attention distributions Zagoruyko and Komodakis (2017); Hu et al. (2018) are introduced to improve the student model.

In this work, we focus on task-agnostic knowledge distillation of pretrained Transformers. The distilled task-agnostic model can be fine-tuned to adapt to downstream tasks. It can also be utilized to initialize task-specific distillation Sun et al. (2019a); Turc et al. (2019); Aguilar et al. (2019); Mukherjee and Awadallah (2020); Xu et al. (2020); Hou et al. (2020); Li et al. (2020), which uses a fine-tuned teacher model to guide the training of the student on specific tasks. Knowledge used for distillation and layer mapping function are two key points for task-agnostic distillation of pretrained Transformers. Most previous work uses soft target probabilities, hidden states, self-attention distributions and value-relation to train the student model. For the layer mapping function, TinyBERT Jiao et al. (2019) uses a uniform strategy to map teacher and student layers. MobileBERT Sun et al. (2019b) assumes the student has the same number of layers as its teacher to perform layer-wise distillation. MiniLM Wang et al. (2020) transfers self-attention knowledge of teacher’s last layer to the student last Transformer layer. Different from previous work, our method uses multi-head self-attention relations to eliminate the restriction on the number of student’s attention heads. Moreover, we show that transferring the self-attention knowledge of an upper middle layer of the large-size teacher model is more effective.

Multi-Head Self-Attention Relation Distillation

Following MiniLM Wang et al. (2020), the key idea of our approach is to deeply mimic teacher’s self-attention module, which draws dependencies between words and is the vital component of Transformer. MiniLM uses teacher’s self-attention distributions to train the student model. It brings the restriction on the number of attention heads of students, which is required to be the same as its teacher. To introduce more fine-grained self-attention knowledge and avoid using teacher’s self-attention distributions, we generalize deep self-attention distillation in MiniLM and introduce multi-head self-attention relations of pairs of queries, keys and values to train the student. Besides, we conduct extensive experiments and find that layer selection of the teacher model is critical for distilling large-size models. Figure 1 gives an overview of our method.

Multi-head self-attention relations are obtained by scaled dot-product of pairsThere are nine types of self-attention relations, such as query-query, key-key, key-value and query-value relations. of queries, keys and values of multiple relation heads. Taking query vectors as an example, in order to obtain queries of multiple relation heads, we first concatenate queries of different attention heads and then split the concatenated vector based on the desired number of relation heads. The same operation is also performed on keys and values. For teacher and student models which uses different number of attention heads, we convert their queries, keys and values of different number of attention heads into vectors of the same number of relation heads to perform KD training. Our method eliminates the restriction on the number of attention heads of student models. Moreover, using more relation heads in computing self-attention relations brings more fine-grained self-attention knowledge and improves the performance of the student model.

We use A1,A2,A3\mathbf{A}_{1},\mathbf{A}_{2},\mathbf{A}_{3} to denote the queries, keys and values of multiple relation heads. The KL-divergence between multi-head self-attention relations of the teacher and student is used as the training objective:

2 Layer Selection of Teacher Model

Besides the knowledge used for distillation, mapping function between teacher and student layers is another key factor. As in MiniLM, we only transfer the self-attention knowledge of one of the teacher layers to the student last layer. Only distilling one layer of the teacher model is fast and effective. Different from previous work which usually conducts experiments on base-size teachers, we experiment with different large-size teachers and find that transferring self-attention knowledge of an upper middle layer performs better than using other layers. For BERTLARGE{}_{\text{LARGE}} and BERTLARGE-WWM{}_{\text{LARGE-WWM}}, transferring the 2121-th (start at one) layer achieves the best performance. For RoBERTa\textsclarge{}_{\textsc{large}} and XLM-RLARGE{}_{\text{LARGE}}, using the self-attention knowledge of 1919-th layer achieves better performance. For the base-size teacher, we also find that using teacher’s last layer performs better.

Experiments

We conduct distillation experiments on different teacher models including BERTBASE{}_{\text{BASE}}, BERTLARGE{}_{\text{LARGE}}, BERTLARGE-WWM{}_{\text{LARGE-WWM}}, RoBERTa\textscbase{}_{\textsc{base}}, RoBERTa\textsclarge{}_{\textsc{large}}, XLM-RBASE{}_{\text{BASE}} and XLM-RLARGE{}_{\text{LARGE}}.

We use the uncased version for three BERT teacher models. For the pre-training data, we use English Wikipedia and BookCorpus Zhu et al. (2015). We train student models using 256256 as the batch size and 6e-4 as the peak learning rate for 400,000400,000 steps. We use linear warmup over the first 4,0004,000 steps and linear decay. We use Adam Kingma and Ba (2015) with β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999. The maximum sequence length is set to 512512. The dropout rate and weight decay are 0.10.1 and 0.010.01. The number of attention heads is 1212 for all student models. The number of relation heads is 4848 and 6464 for base-size and large-size teacher model, respectively. The student models are initialized randomly.

For models distilled from RoBERTa, we use similar pre-training datasets as in Liu et al. (2019). For the 12<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>76812<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 student model, we use Adam with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98. The rest hyper-parameters are the same as models distilled from BERT.

For multilingual student models distilled from XLM-R, we perform training using the same datasets as in Conneau et al. (2019) for 1,000,0001,000,000 steps. We conduct distillation experiments using 88 V100 GPUs with mixed precision training.

2 Downstream Tasks

Following previous pre-training Devlin et al. (2018); Liu et al. (2019); Conneau et al. (2019) and task-agnostic distillation Sun et al. (2019b); Jiao et al. (2019) work, we evaluate the English student models on GLUE benchmark and extractive question answering. The multilingual models are evaluated on cross-lingual natural language inference and cross-lingual question answering.

General Language Understanding Evaluation (GLUE) benchmark Wang et al. (2019) consists of two single-sentence classification tasks (SST-2 Socher et al. (2013) and CoLA Warstadt et al. (2018)), three similarity and paraphrase tasks (MRPC Dolan and Brockett (2005), STS-B Cer et al. (2017) and QQP), and four inference tasks (MNLI Williams et al. (2018), QNLI Rajpurkar et al. (2016), RTE Dagan et al. (2006); Bar-Haim et al. (2006); Giampiccolo et al. (2007); Bentivogli et al. (2009) and WNLI Levesque et al. (2012)).

Extractive Question Answering

The task aims to predict a continuous sub-span of the passage to answer the question. We evaluate on SQuAD 2.0 Rajpurkar et al. (2018), which has been served as a major question answering benchmark.

Cross-lingual Natural Language Inference (XNLI)

XNLI Conneau et al. (2018) is a cross-lingual classification benchmark. It aims to identity the semantic relationship between two sentences and provides instances in 1515 languages.

Cross-lingual Question Answering

We use MLQA Lewis et al. (2019b) to evaluate multilingual models. MLQA extends English SQuAD dataset Rajpurkar et al. (2016) to seven languages.

3 Main Results

Table 3 presents the dev results of 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>3846<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>384 and 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 models distilled from BERTBASE{}_{\text{BASE}}, BERTLARGE{}_{\text{LARGE}} and RoBERTa\textsclarge{}_{\textsc{large}} on GLUE and SQuAD 2.0. (1) Previous methods Sanh et al. (2019); Jiao et al. (2019); Sun et al. (2019a); Wang et al. (2020) usually distill BERTBASE{}_{\text{BASE}} into a 66-layer model with 768768 hidden size. We first report results of the same setting. Our 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 model outperforms DistilBERT, TinyBERT, MiniLM and two BERT baselines across most tasks. Moreover, our method allows more flexibility for the number of attention heads of student models. (2) Both 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>3846<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>384 and 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 models distilled from BERTLARGE{}_{\text{LARGE}} outperform models distilled from BERTBASE{}_{\text{BASE}}. The 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 model distilled from BERTLARGE{}_{\text{LARGE}} is 2.0×2.0\times faster than BERTBASE{}_{\text{BASE}}, while achieving better performance. (3) Student models distilled from RoBERTa\textsclarge{}_{\textsc{large}} achieve further improvements. Better teacher results in better students. Multi-head self-attention relation distillation is effective for different large-size pretrained Transformers.

We report the results of 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 students distilled from BERTBASE{}_{\text{BASE}} and BERTLARGE{}_{\text{LARGE}} on GLUE test sets and SQuAD 2.0 dev set in Table 2. 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 model distilled from BERTBASE{}_{\text{BASE}} retains more than 99%99\% accuracy of its teacher while using 50%50\% Transformer parameters. 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>7686<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>768 model distilled from BERTLARGE{}_{\text{LARGE}} compares favorably with BERTBASE{}_{\text{BASE}}.

We compress RoBERTa\textsclarge{}_{\textsc{large}} and BERTLARGE{}_{\text{LARGE}} into a base-size student model. Dev results of GLUE benchmark and SQuAD 2.0 are shown in Table 3. Our base-size models distilled from large-size teacher outperforms BERTBASE{}_{\text{BASE}} and RoBERTa\textscbase{}_{\textsc{base}}. Our method can be adopted to train students in different parameter size. Moreover, our student distilled from RoBERTa\textsclarge{}_{\textsc{large}} uses a much smaller (almost 32×32\times smaller) training batch size and fewer training steps than RoBERTa\textscbase{}_{\textsc{base}}. Our method requires much fewer training examples.

Most of previous work conducts experiments using base-size teachers. To compare with previous methods on large-size teacher, we reimplement MiniLM and compress BERTLARGE-WWM{}_{\text{LARGE-WWM}} into a 12<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>38412<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>384 student model. Dev results of SQuAD 2.0, MNLI-m and SST-2 are presented in Table 4. Our method also outperforms MiniLM for large-size teachers. Moreover, we report results of distilling an upper middle layer instead of the last layer for MiniLM. Layer selection is also effective for MiniLM when distilling large-size teachers.

Table 5 and Table 6 show the results of our student models distilled from XLM-R on XNLI and MLQA. For XNLI, the best single model is selected on the joint dev set of all the languages as in Conneau et al. (2019). Following Lewis et al. (2019b), we adopt SQuAD 1.1 as training data and evaluate on MLQA English development set for early stopping. Our 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>3846<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>384 model outperforms mBERT Devlin et al. (2018) with 5.3×5.3\times speedup. Our method also performs better than MiniLM, which further validates the effectiveness of multi-head self-attention relation distillation. Transferring multi-head self-attention relations can bring more fine-grained self-attention knowledge.

4 Ablation Studies

We perform ablation studies to analyse the contribution of different self-attention relations. Dev results of three tasks are illustrated in Table 7. Q-Q, K-K and V-V self-attention relations positively contribute to the final results. Besides, we also compare Q-Q + K-K + V-V with Q-K + V-V given queries and keys are employed to compute self-attention distributions in self-attention module. Experimental result shows that using Q-Q + K-K + V-V achieves better performance.

Effect of distilling different teacher layers

Figure 2 presents the results of 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>3846<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>384 model distilled from different layers of BERTBASE{}_{\text{BASE}}, BERTLARGE{}_{\text{LARGE}} and XLM-RLARGE{}_{\text{LARGE}}. For BERTBASE{}_{\text{BASE}}, using the last layer achieves better performance than other layers. For BERTLARGE{}_{\text{LARGE}} and XLM-RLARGE{}_{\text{LARGE}}, we find that using one of the upper middle layers achieves the best performance. The same trend is also observed for BERTLARGE-WWM{}_{\text{LARGE-WWM}} and RoBERTa\textsclarge{}_{\textsc{large}}.

Effect of different number of relation heads

Table 8 shows the results of 6<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>3846<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>384 model distilled from BERTBASE{}_{\text{BASE}} and RoBERTa\textscbase{}_{\textsc{base}} using different number of relation heads. Using a larger number of relation heads achieves better performance. More fine-grained self-attention knowledge can be captured by using more relation heads, which helps the student to deeply mimic the self-attention module of its teacher. Besides, we find that the number of relation heads is not required to be a positive multiple of both the number of student and teacher attention heads. The relation head can be a fragment of a single attention head or contains fragments from multiple attention heads.

Discussion

MobileBERT Sun et al. (2019b) compresses a specially designed teacher model (in the BERTLARGE{}_{\text{LARGE}} size) with inverted bottleneck modules into a 2424-layer student using the bottleneck modules. Since our goal is to compress different large models (e.g. BERT and RoBERTa) to small models using standard Transformer architecture, we note that our student model can not directly compare with MobileBERT. We provide results of a student model with the same parameter size for a reference. A public large-size model (BERTLARGE-WWM{}_{\text{LARGE-WWM}}) is used as the teacher, which achieves similar performance as MobileBERT’s teacher. We distill BERTLARGE-WWM{}_{\text{LARGE-WWM}} into a student model (2525M parameters) using the same training data (i.e., English Wikipedia and BookCorpus). The test results of GLUE and dev result of SQuAD 2.0 are illustrated in Table 9. Our model outperforms MobileBERT across most tasks with a faster inference speed. Moreover, our method can be applied for different teachers and has much fewer restriction of students.

We also observe that our model performs relatively worse on CoLA compared with MobileBERT. The task of CoLA is to evaluate the grammatical acceptability of a sentence. It requires more fine-grained linguistic knowledge that can be learnt from language modeling objectives. Fine-tuning the model using the MLM objective as in MobileBERT brings improvements for CoLA. However, our preliminary experiments show that this strategy will lead to slight drop for other GLUE tasks.

2 Results of More Self-Attention Relations

In Table 9 and 10, we report results of students trained using more self-attention relations (Q-K, K-Q, Q-V, V-Q, K-V and V-K relations). We observe improvements across most tasks, especially for student models distilled from BERT. Fine-grained self-attention knowledge in more attention relations improves our students. However, introducing more self-attention relations also brings a higher computational cost. In order to achieve a balance between performance and computational cost, we choose to transfer Q-Q, K-K and V-V self-attention relations instead of all self-attention relations in this work.

Conclusion

We generalize deep self-attention distillation in MiniLM by employing multi-head self-attention relations to train the student. Our method introduces more fine-grained self-attention knowledge and eliminates the restriction of the number of student’s attention heads. Moreover, we show that transferring the self-attention knowledge of an upper middle layer achieves better performance for large-size teachers. Our monolingual and multilingual models distilled from BERT, RoBERTa and XLM-R obtain competitive performance and outperform state-of-the-art methods. For future work, we are exploring an automatic layer selection algorithm. We also would like to apply our method to larger pretrained Transformers.

References

Appendix A GLUE Benchmark

The summary of datasets used for the General Language Understanding Evaluation (GLUE) benchmarkhttps://gluebenchmark.com/ Wang et al. (2019) is presented in Table 11.

Appendix B SQuAD 2.0

We present the dataset statistics and metrics of SQuAD 2.0http://stanford-qa.com Rajpurkar et al. (2018) in Table 12.

Appendix C Hyper-parameters for Fine-tuning

For SQuAD 2.0, the maximum sequence length is 384384. The batch size is set to 3232. We choose learning rates from {3e-5, 6e-5, 8e-5, 9e-5} and fine-tune the model for 3 epochs. The warmup ration and weight decay is 0.1 and 0.01.

GLUE

The maximum sequence length is 128128 for the GLUE benchmark. We set batch size to 3232, choose learning rates from {1e-5, 1.5e-5, 2e-5, 3e-5, 5e-5} and epochs from {33, 55, 1010} for different student models. We fine-tune CoLA task with longer training steps (25 epochs). The warmup ration and weight decay is 0.1 and 0.01.

Cross-lingual Natural Language Inference (XNLI)

The maximum sequence length is 256256 for XNLI. We fine-tune 1010 epochs using 6464 as the batch size. The learning rates are chosen from {3e-5, 4e-5, 5e-5, 6e-5}.

Cross-lingual Question Answering

For MLQA, the maximum sequence length is 512512. We fine-tune 44 epochs using 3232 as the batch size. The learning rates are chosen from {3e-5, 4e-5, 5e-5, 6e-5}.