Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers
Machel Reid, Edison Marrese-Taylor, Yutaka Matsuo
Introduction
Recent improvements in NLP tasks can be attributed to the Transformer (Vaswani et al., 2017) model. The model has led to better deeply contextualized representations (Devlin et al., 2019), better machine translation systems (Vaswani et al., 2017), and language models (Baevski and Auli, 2019; Dai et al., 2019). Despite their success, one main drawback of training these models is their computational cost, being a greatly limiting factor for many, with training times and memory usage ballooning as model sizes increase to attain better performance.
With this in mind, there has been recent interest in making the Transformer more parameter-efficient to reap its performance benefits while making the model more computationally efficient and able to scale better. Many approaches have focused on automating model design with neural architecture search that aim at finding more efficient Transformer variations using gradient descent (Wu et al., 2020; So et al., 2019; Mehta et al., 2020a). As such, these techniques are expensive, requiring a significant amount of GPU hours to find good designs. In contrast to these approaches, the model by (Lan et al., 2020) has directly tackled model parameter reduction in the context of deeply contextualized word representations, still attaining similar (or better) performance. In this paper, we take a similar approach and look to explore whether these ideas can be applied to sequence-to-sequence models in a simple manner by designing the Subformer.
The Subformer incorporates two novel techniques: (1) SAFE (Self-Attentive Factorized Embeddings), in which we use a small self-attention layer to reduce embedding parameter count, and (2) Sandwich-style Parameter Sharing, in which we develop a simple and intuitive technique for parameter sharing to be effective in Transformer models.
We evaluate the Subformer on machine translation, abstractive summarization, and language modeling, showing that our model can achieve similar or better performance compared with a base/big Transformer with a 40% parameter reduction and minimal modification to the original architecture, reinforcing the existing over-parameterization claims (Fan et al., 2020; Mehta et al., 2020a; Lan et al., 2020). On WMT’14 EN-DE we achieve a BLEU score of 29.3, compared to Transformer-big’s 28.6 with 13M fewer parameters, and we also outperform the Transformer-XL model with a significant 3.6 lower perplexity and 37% fewer parameters.
The Subformer
Let us start by defining the notation to be used throughout the paper. We refer to the model dimension as , feed-forward projection dimension as , the vocabulary size as , and the number of layers as . Note that, unlike standard Transformer models, in which the embedding dimension is kept the same as , we disentangle to embedding dimension to reduce parameter count. For this reason, we denote the embedding dimension to be .
We propose to reduce the number of parameters in our embedding layers, which can take up to 25% of the total parameter count (in the case of Transformer-base, Vaswani et al., 2017), using a small self-attention layer. Specifically, we look to reduce the embedding size by disentangling the model dimension from the embedding dimension, reducing the embedding dimension , and then projecting this to the model dimension using a small self-attention sub-layer followed by a feed-forward module.
Given vocabulary size , the usage of a standard embedding layer would result in parameters. However, considering that the power of Transformers lies in their ability to learn contextual representations with attention, using a smaller for non-contextual embeddings and then projecting to is intuitively an effective method for parameter reduction (Lan et al., 2020). This results in a significant parameter reduction for values of .
Current models (Baevski and Auli, 2019; Dai et al., 2019; Lan et al., 2020) often use a single linear projection, i.e. . In contrast, we empirically show that simply contextualizing this projection with a small self-attention layer results in stronger performance with a minimal addition of parameters —especially in the encoder-decoder case, where the input embedding layer and output projection are often tied (Press and Wolf, 2017) (Table 1).
2 Sandwich-style Parameter Sharing
Weight sharing techniques, despite being surprisingly effective, have been relatively unexplored in the context of generative Transformers. However, this has been shown to be powerful for leveraging models with large capacity and less memory usage/computation (Wu et al., 2019; Lan et al., 2020).
Given that the output of each Transformer layer depends directly on its two sub-layers —multiheaded attention and the feedforward module— when discussing alternatives for parameter sharing across transformer layers there are several options. As we aim to leverage the aforementioned properties of weight sharing, we performed preliminary experiments, investigating the capabilities of weight sharing in the following five settings: (1) All-shared Naively sharing all encoder and all decoder layers —that is including both of their sub-layers, following Lan et al. (2020); Dehghani et al. (2019). (2) All-Shared (Independent FFN) Naively sharing all encoder and all decoder layers but allowing each layer to have an independent feed-forward sub-layer. (3) All-Shared (except last) Sharing weights across layers such that layer remains independent. (4) Every 2 layers shared Sharing every two layers, i.e. in the case of a 6-layer transformer. (5) Sandwich Finally, we only share the middle or central layers (i.e. ), leaving layers and to have independent sets of parameters.
Table 2 summarizes the results of our exploratory study. As can be seen, naive parameter sharing/tying approaches do not offer any advantages, hurting performance significantly (50%) when compared to the regular Transformer. However, our results also show that when combined properly, using Sandwich-style parameter sharing, we can attain a good balance of parameter reduction and performance. When compared to tasks such as pre-training deep contextualized word representations, generative tasks such as machine translation require informative token-level representations for each input token to be accurately translated. In this context, we surmise that the success of Sandwich-style parameter sharing on this sequence-to-sequence task is a result of it being able to have the input and output layers (arguably, the most important layers) be trained independently, allowing them to learn different operations than the sandwich layers.
Having introduced our proposed parameter reduction techniques, we will now explain the Subformer architecture. The Subformer is composed of four main components, for both the encoder and decoder: the embedding layer, the model layers, the sandwich module and the projection layers. We disentangle the sandwiched layer dimension from that of the model layer, allowing the sandwich layer width to be larger than the rest of the model. For this reason, we denote the dimension of the sandwiched layer to be and its corresponding feed-forward dimension to be .
Experimental Setup
We apply our method to a variety of sequence modeling tasks: neural machine translation, summarization, and language modeling. We compare Transformer-base and big (Vaswani et al., 2017) with the Subformer trained in the same setting. We also include simple sandwich-style parameter sharing (denoted as Sandwich-{base,big}) and the usage of only SAFE as an ablation of what these techniques do in their naive forms when decoupled. Additional implementation and training details with hyperparameter settings are in the appendix.
We evaluate our model on two standard translation benchmarks: WMT’14 English-German (EN-DE) and WMT’16 English-Romanian (EN-RO). Following previous work, we evaluate all models using tokenized BLEU (Papineni et al., 2002). In order to better contextualize our results, we consider parameter-efficient models DeLighT (Mehta et al., 2020a) (contemporaneous work), and the Evolved Transformer (So et al., 2019).
Abstractive Summarization
We test the model’s ability to process long documents on the CNN-DailyMail summarization benchmark (Hermann et al., 2015; Nallapati et al., 2016), comprising over 280K news articles paired with multi-sentence summaries. We also compare effects of BART (Lewis et al., 2020) pretraining (details in Appendix B). For this task we contexutalize our results with specialized architectures such as Pointer-Generator Networks (See et al., 2017), and methods leveraging pretraining: RobertaShare (Rothe et al., 2020), BertExtAbs Liu and Lapata (2019), and BART (Lewis et al., 2020). We evaluate using ROUGE 1,2,L (Lin, 2004).
Language Modeling
We evaluate on the large-scale Wikitext-103 dataset (Merity et al., 2017). Models are evaluated in terms of perplexity on the test portion. In order to better contextualize our results, we consider the QRNN (Merity et al., 2018), Transformer-XL (Dai et al., 2019) and Deep Equilibrium Model (DEQ) (Bai et al., 2019), which also employs parameter sharing.
Results
Tables 3 and 4In all tables, results from other work used to contextualize our results are placed above the double bar. summarize our results on machine translation. Firstly, we note that our Transformer baselines outperform Vaswani et al. (2017) (base: 27.3 27.7, big: 28.4 28.6). We surmise that this is due to training for longer and with a larger batch size.
The Subformer outperforms our Transformer baselines when trained in the same setting, with similar or fewer parameters. Subformer-small reduces parameters by 40%, matching performance with our Transformer baselines. Subformer-base and mid, outperform our model significantly (0.4, 0.8 BLEU) when using a fewer/similar number of parameters. Furthermore, we note that Subformer-mid performs similarly to the Transformer-big model (210M params.) in Table 4, despite a 70% parameter reduction.
For our large set of models (Table 4), Sandwich-big achieves the same performance as our Transformer-big reimplementation, but with 40% fewer parameters. We believe this to be an indication of the capability of Sandwich-style parameter sharing as the encoder/decoder layers get wider, while also providing further evidence for the overparameterization of the Transformer. Subformer-large, with 197M parameters achieves a significant 0.7 BLEU score gain over Transformer-big, despite using 13M fewer parameters.
Language Modeling
Results for language modeling can be seen in Table 5. The Subformer uses adaptive input embeddings (Baevski and Auli, 2019) instead of SAFE embeddings, following common practice. We also train two Transformer baselines with the same setup —one with the same amount of parameters and another with a similar parameter count to Transformer-XL— to provide better context for comparison. Task-specific techniques that can be applied during training, such as caching (Dai et al., 2019) or other methods applied during inference time (Khandelwal et al., 2020; Krause et al., 2018) can further improve all models so we do not focus on these.
As seen in Table 5, the Subformer outperforms the baselines by a significant margin (between 1.4 and 6.5 perplexity), with a significant reduction in parameters.
Abstractive Summarization
For the CNN/Daily Mail summarization task we use Subformer-base. We also pretrain a Transformer and Subformer model with the same architecture on Wikipedia (details in Appendix B). As can be seen in Table 6, the Subformer outperforms two Transformer baselines with both the same parameter count and its respective Transformer-base configuration in both settings, demonstrating the Subformer’s performance on a variety of tasks and with longer sequences.
Discussion on Speed and Convergence
We found training time to consistently speed up by 10-30%, and inference speed to either be faster by 10-20% (keeping ) to be slightly slower by 10-30% (when ) (due to more operations, similar to Lan et al. (2020)). The Subformer converges faster most likely due to fewer parameters to optimize. Given the fewer number of parameters, it can be expected for the models to converge with fewer iterations. We test this on the task of language modeling, where we found that the Subformer converged 65% faster than its Transformer counterpart, as shown in Table 7. We also measure inference speed on our machine translation models (Table 8).
Conclusion
In this paper we have presented the Subformer, a parameter-efficient Transformer-based generative model primarily based on two simple parameter factorization/sharing techniques. Our empirical results on three sequence modeling tasks show that the Subformer can achieve similar or better performance compared with a base/big Transformer with a 40% parameter reduction. Furthermore, the simplicity of our approach allows the Subformer to be used in conjunction with other parameter reduction techniques in the literature, for even smaller but performant models. We hope our work incites interest in parameter sharing techniques for an even wider range of Transformer models, ultimately helping reduce their computational cost in general.
Ethical Considerations
This work has impact in the field of natural language processing, and develops a more efficient approach for learning performant generative models. As with much of language technology has the potential to be both used for good and used maliciously. We also experiment with pretraining, learning representations in an unsupervised way, which is likely to capture and amplify biases found in the data. However, our approach has a potential positive impact given the lower cost and energy expenditure needed to train our proposed model.
Acknowledgments
We thank Jorge Balazs, Yusuke Iwasawa, Jungo Kasai, Cristian Rodriguez-Opazo, Alfredo Solano, Yutaro Yamada, and Victor Zhong for their helpful feedback and discussions over this work. MR is grateful to the Masason Foundation for their support.
References
Appendix A Designing the Subformer
The Subformer is a play on words, referencing its small size - i.e. sub-, as well as the Sandwich-style parameter sharing technique, referencing the type of sandwich.
Architecture
When using SAFE, our parameter count would result in parameters, where the first term represents the embedding layer, the value groups the query, key, and value projections and 2 output feed-forward layers, and represents the linear projection from the embedding dimension to the model dimension. As mentioned in the paper, this results in a significant parameter reduction for values of .
As we tie the decoder’s output projection layer (returning a distribution over the vocabulary) with the decoder’s input embedding matrix, we project the decoder’s last hidden state (with dimension ) to using a two layer MLP. Also, when we perform encoder attention in the decoder’s Sandwich Module, we simply linearly project the query from the decoder from to and then project it back to once the attention operation is complete.
Memory Footprint
Table 9 summarizes the memory footprint of our proposed techniques. In this table, the benefits of Sandwich-style parameter sharing can be seen as the number of independent layers is controlled to be , however, Transformers generally need to be deeper (with a standard of ) to learn more meaningful representations with the parameter count scaling linearly with respect to the layer count. Similarly, the benefits of disentangling the model dimension from the embedding dimension can be seen as well. Due to the parameter reduction attained by these techniques, the models can be trained in memory-constrained scenarios with a larger batch size.
Appendix B Data and Training Details
Training was done on 8 GPUs on a single DGX-1 Machine, while training on 16 GPUs was done using multiple compute nodes of a compute cluster. We train all base/small models on 8 NVIDIA Tesla V100 GPUs. For all big/large models, we train on 16 NVIDIA Tesla V100 GPUs. All models were trained with mixed precision (Micikevicius et al., 2018) and are implemented in PyTorch (Paszke et al., 2019) using our modification of fairseq (Ott et al., 2019).
We train using 8192 tokens per GPU an update frequency of 2, for small, base models. For large models, we train on 16 GPUs with 4096 tokens per GPU with an update frequency of 2. We follow the training setup of Ghazvininejad et al. (2019): we use the same weight initialization scheme as BERT (Devlin et al., 2019), sampling weights from , initializing biases to zero and setting layer normalization parameters and to be and , respectively. For regularization we use the best of dropout, weight decay of , while using label-smoothed cross entropy loss with . We train using an effective batch size of 128K tokens. The models are trained using Adam (Kingma and Ba, 2015), with hyper-parameters and . We warm up the learning rate to a peak of within 10K iterations and then decay the learning rate with the inverse square root schedule. When creating the final model, we use the checkpoint with the lowest loss on the development set and generate using a beam size of 5 (Vaswani et al., 2017), tuning the length penalty of … in the validation set. We perform early stopping, training for a maximum of 250K iterations.
We use the following settings for our models: (1) Subformer-small has , , and , (2) Subformer-base has , , , , (3) Subformer-mid has , , and (4) Subformer-large has , and . For WMT’16 EN-RO, our small model has , and and our base model has , , and .
In terms of datasets, we make use of the same pre-processed data used by Ghazvininejad et al. (2019) for WMT’14 EN-DE with a 32K BPE (Sennrich et al., 2016) vocabulary and during evaluation we perform de-hyphenation (Vaswani et al., 2017). We use the same data as Lee et al. (2018) for WMT’16 EN-RO with a 35K BPE vocabulary.
Abstractive Summarization
We follow Edunov et al. (2019) and use the official ROUGE-1.5.5.pl script with parameters -m -a -n 2. As mentioned in the paper, our model configuration is the same as Subformer-base, but we set . Articles are truncated to 400 tokens (See et al., 2017) and we use a BPE vocabulary of 32K types (Edunov et al., 2019). We follow the training schedule of Edunov et al. (2019). During inference, we tune generation length in the range of {40, 50, 60} and use tri-gram blocking, following standard practice. When pretraining, we pretrain on Wikipedia (14GB) for 100K iterations, using a batch size of 512K tokens. We use a learning rate of 7e-4, warmed up over 10K iterations.
Language Modeling
When training our language models, we use 8 GPUs with 3072 tokens per GPU and an update frequency of 3, following Baevski and Auli (2019). Models are trained using Nesterov’s accelerated gradient optimizer (Sutskever et al., 2013), warming up the learning rate to 1.0 for 16K iterations, and then annealing for 270K iterations using a cosine annealing schedule. We use three configurations: (1) , (2) and and (3) and . All models use . Our considered dataset, Wikitext-103, contains 103M tokens and has a vocabulary of nearly 270K.
Appendix C Extended Related Work
Given the effectiveness of the Transformer, improving the architecture has been of much interest to the NLP community. Within this domain, one branch of research concerns the reduction of the quadratic complexity (w.r.t. sequence length) of the Transformer’s core self-attention mechanism (Wu et al., 2019; Kitaev et al., 2020), pushing it down to linear or log-linear complexity. The second branch of research regards improving the expressiveness of Transformer models, by using more layers (Dou et al., 2018), or by improving the architecture (Wu et al., 2019; So et al., 2019). The third branch of research regards improving the parameter efficiency of Transformers. Approaches towards this goal include neural architecture search approaches (So et al., 2019; Wu et al., 2020), where new Transformer-based architectures are learned using gradient descent, more manually crafted approaches (Dehghani et al., 2019; Mehta et al., 2020a), as well as weight-sharing approaches (Lan et al., 2020; Wu et al., 2019). The work most similar to ours is ALBERT (Lan et al., 2020) in which complete weight sharing is used to pre-train deep contextualized word representations (Peters et al., 2018; Devlin et al., 2019). Different from this work, we focus on common NLP generative/sequence-to-sequence tasks versus large-scale pre-training and develop an approach to increase model capacity while reducing parameter footprint tailored to this setting.
Compressing Transformers
We find prior work on pruning and quantizing Transformer models to reduce their size with a focus on sequence-to-sequence settings like machine translation (Prato et al., 2019), on encoder-based methods like BERT (Zafrir et al., 2019; Ganesh et al., 2020) or with a more generic scope in mind (Cheong and Daniel, 2019; Lee et al., 2019). Our approach is orthogonal to these since we directly aim at reducing the number of parameters of Transformer models by proposing architecture modifications and weight sharing techniques.
Reducing Embedding Dimensionality in Sequence Models
As embeddings can substantially increase the parameter count as the vocabulary size increases, especially in sequence modeling scenarios, embedding reduction techniques have been proposed, including using a linear projection to project to a lower dimension (Baevski and Auli, 2019; Dai et al., 2019) or using combinations of block sparse transformations (Mehta et al., 2020b, a). We propose a self-attention based projection layer, SAFE, which we empirically show to outperform the aforementioned linear projection methods with a similar parameter count.