Orthogonal Language and Task Adapters in Zero-Shot Cross-Lingual Transfer
Marko Vidoni, Ivan Vulić, Goran Glavaš
Introduction
Multilingual representation spaces aim to capture meaning across language boundaries and that way (at least conceptually) enable cross-lingual transfer of task-specific NLP models from resource-rich languages with large annotated datasets to resource-lean languages without (m)any labeled instances Joshi et al. (2020). When such a multilingual representation space is defined via parameters of a deep neural language encoder, e.g., a transformer network Vaswani et al. (2017); Devlin et al. (2019), it can be used to transfer knowledge across tasks as well as across languages (Pruksachatkun et al., 2020; Ponti et al., 2020b; Pfeiffer et al., 2020b; Phang et al., 2020, inter alia) .
Massively multilingual transformers (MMTs), pretrained on large multilingual corpora via language modeling (LM) objectives Devlin et al. (2019); Conneau et al. (2020) have recently overthrown (static) cross-lingual word embedding spaces Ruder et al. (2019); Glavaš et al. (2019) as the state-of-the-art paradigm for zero-shot cross-lingual transfer in NLP. However, MMTs are constrained by the so-called curse of multilinguality, a phenomenon where the quality of language-specific representations starts deteriorating when the number of pretraining languages exceeds the MMT’s parameter capacity Conneau et al. (2020). Languages with smallest pretraining corpora are most affected: the largest transfer performance drops are observed with those target languages Lauscher et al. (2020); Wu and Dredze (2020).
Additional LM training of a pretrained MMT on monolingual corpora of an underrepresented target language (i.e., target language LM-fine-tuning) is a partial remedy towards satisfactory downstream transfer Wang et al. (2020a); Ponti et al. (2020a). However, this approach does not increase the MMT capacity and, consequently, might deteriorate its representation quality for other languages.
As a solution, Pfeiffer et al. (2020c) propose the MAD-X transfer framework: they freeze the original MMT parameters and train only a small number of additional parameters, the so-called adapter modules Rebuffi et al. (2017); Houlsby et al. (2019) via language-specific LM-fine-tuning. This approach effectively 1) increases the model capacity, and 2) captures language-specific knowledge in the adapters.In other words, there is a separate set of adapter parameters for each language with the added capacity reserved strictly for that language. The next step is then task-specific fine-tuning on the source-language task data with a fixed source-language adapter (see Figure 1): a separate set of task adapter parameters is optimized to keep task-specific knowledge separate from language-specific knowledge stored in language adapters. Source language adapters are then replaced with target language adapters (with task adapters kept fixed), to make task predictions in the target language. The MAD-X framework, however, does not provide any mechanism that would prevent language and task adapters from capturing redundant information, that is, from storing some knowledge already encoded in the MMT’s parameters.
In this work, we advance the idea of augmenting MMT’s knowledge through specialized adapter modules. We aim to maximize the injection of novel information into both language- and task-specific adapter parameters, that is, we enforce the adapters to encode the information that complements the knowledge encoded in MMT’s pretrained parameters. To achieve this, we propose to learn orthogonal adapters (or orthoadapters for short). We augment the training objectiveFor a language adapter, the training task is masked language modeling on the monolingual corpus of that language. with the orthogonality loss: it forces the representations produced by the adapters to be orthogonal to representations from the corresponding MMT layers.
Our proof-of-concept zero-shot transfer experiments, encompassing cross-lingual transfer for three tasks – POS-tagging, named entity recognition (NER), and natural language inference (XNLI) – and 10 typologically diverse languages, render language-specific and task-specific orthoadapters viable mechanisms for improving downstream language transfer performance. However, we demonstrate that the optimal use of orthogonality is also largely task-dependent. We hope that our study will inspire a wider investigation of applicability and usefulness of orthogonality constraints in the context of both language-specific and task-specific fine-tuning of pretrained transformers. We will make the orthoadapters code publicly available.
Orthogonal Adapters
Figure 1 provides an illustrative overview of our cross-lingual transfer framework for training and using language and task orthoadapters.
We first train orthoadapters for each language: we inject a language adapter into each transformer layer and train its parameters via masked language modeling (MLM) on the monolingual corpus of the respective language. We can augment the MLM objective with the orthogonality loss: it forces the adapter output to be orthogonal to the representations produced by the original transformer layer without the adapter (i.e., the multi-head attention (sub)layer coupled with the feed-forward layer). In the second step, we perform task-specific training in the source language. To this end, we inject the task adapter to store task-specific knowledge. We can again aim to acquire knowledge that is absent from both the original MMT parameters by augmenting the task-specific training objective with the orthogonality loss. Keeping the language knowledge separate from task knowledge, we hope to impede negative interference and forgetting Hashimoto et al. (2017); de Masson d’Autume et al. (2019); Wang et al. (2020b).Note that the task-specific fine-tuning procedure, in which we optimize task orthoadapters, is aware of the source language knowledge from the language orthoadapter. Finally, as in the MAD-X framework, we swap the source-language adapter with corresponding target-language adapters for the task-specific inference on target-language data.
Adapter Architecture. We adopt the well-performing and lightweight adapter configuration of Pfeiffer et al. (2020a), where only one adapter module is injected per transformer layer, after the feed-forward sublayer.Pfeiffer et al. (2020a) found this configuration to perform on a par with the configuration proposed by Houlsby et al. (2019), who inject two adapter modules per transformer layer (the other one after the multi-head attention sublayer), while being more efficient to train. The concrete adapter architecture we use is the variant of the so-called bottleneck adapter Pfeiffer et al. (2020a):
Orthogonality Loss. Adapters allow for computationally more efficient fine-tuning Houlsby et al. (2019), while simultaneously offering task performance comparable to that of standard full fine-tuning of all transformer parameters Pfeiffer et al. (2020a). However, there is currently no mechanism in adapter-based fine-tuning that would explicitly prevent adapter parameters from learning redundant information which is already captured by the pretrained model. Inspired by the idea of orthogonal text representations introduced in the context of multi-task learning with shared and task-specific (i.e., private) parameters Romera-Paredes et al. (2012); Liu et al. (2017), we introduce an additional orthogonality loss in adapter-based fine-tuning. It explicitly forces the adapters to dedicate their capacity to new knowledge, which should be complementary (i.e., non-redundant) to the knowledge already encoded in existing transformer’s parameters.
In particular, we can augment the training objective in each of the two training phases (i.e., when training both language adapters and task adapters) with an auxiliary orthogonality loss. Let denote the hidden representation in the -th layer of the MMT for the -th token in the sequence, input to the adapter. Let be the corresponding output of the same adapter for the same token, as given in Eq. (1). The orthogonality loss of the -th token in the -th MMT layer is then simply the square of the cosine similarity between and ; we then derive the overall orthogonality loss of the model by averaging token-level losses in each layer and then summing layer-level losses:
where is the maximal length of the input token sequence and is the number of MMT’s layers.
Two-Step Orthoadapter Training. In the first step (see part a of Figure 1), we train language orthoadapters, independently for each language. This way we aim to extend the language knowledge captured by the pretrained MMT. As in MAD-X Pfeiffer et al. (2020c), we use masked LM-ing (cross-entropy loss) as the main training objective . We then alternately update the parameters of language orthoadapters, first by minimizing and then by minimizing .We use two independent Adam optimizers Kingma and Ba (2015), one for each loss. We also experimented with minimizing the joint loss but this generally yielded poorer performance although we tried a wide range of values for the weight .
In the second step (part b in Figure 1), the goal is to maximize the amount of novel information useful for a concrete downstream task: we train the task orthoadapters on the task-specific training data (POS, NER, XNLI) by alternately minimizing 1) the task-specific objective Cross-entropy loss for the whole sequence for NLI; sum of token-level cross-entropy losses for POS and NER. and 2) the orthogonality loss . Note, however, that in this case is, in each transformer layer, first adapted by the source language adapter and then by the task adapter, and is the output of the task adapter.
Zero-Shot Cross-Lingual Transfer then proceeds in the same vein as in the MAD-X framework Pfeiffer et al. (2020c). The source-to-target transfer is conducted by simply replacing the source language (ortho)adapter with the target language (ortho)adapter while relying on exactly the same task adapter fine-tuned with the labeled source language data, stacked on top of the language adapters (see part c of Figure 1). We refer the reader to the original work Pfeiffer et al. (2020c, b) for further technical details.
Experimental Setup
Model Configurations. The decomposition into two adapter types in the two-step procedure (Figure 1) allows us 1) to use language orthoadapters (l-ort) instead of regular non-orthogonal language adapters (l-noo); and/or 2) to replace non-orthogonal task adapters (t-noo) with task orthoadapters (t-ort). These choices give rise to four different model variants, where the l-noo+t-noo variant is the baseline MAD-X variant.
We also test the usefulness of task orthoadapters in a standard “non-MAD-X” setup, i.e., without dedicated language adapters: t-ort variants are compared to t-noo variants, and also to standard full fine-tuning of the whole MMT (full-ft) with in-task labelled data, which is computationally more intensive than adapter-based fine-tuning.
Evaluation Tasks and Data. We evaluate all model variants on standard cross-lingual transfer tasks, relying on established evaluation benchmarks: 1) sentence-pair classification on XNLI Conneau et al. (2018)); 2) cross-lingual named entity recognition (NER) on the WikiANN dataset Pan et al. (2017);As prior work Pfeiffer et al. (2020c); Hu et al. (2020), we use the data splits provided by Rahimi et al. (2019). 3) part-of-speech tagging with universal POS tags from the Universal Depenedencies Nivre et al. (2018) (UD-POS).
In all experiments we rely on the pretrained multilingual XLM-R (Base) model Conneau et al. (2020), which showed state-of-the-art zero-shot performance in a recent comparative empirical study of Hu et al. (2020), and even stronger results when combined with the adapter-based MAD-X framework Pfeiffer et al. (2020c). We always treat English (en) as our (resource-rich) source language: the pretrained XLM-R model is fine-tuned on the English task data, and then evaluated in the zero-shot setting on different target languages. For completeness, we also report the results on the English test data, i.e., without any transfer.
Target Languages. The selection of target languages has been guided by several (sometimes clashing) criteria: C1) typological diversity; C2) availability in the standard evaluation benchmarks; C3) computational tractability; C4) evaluation also on truly low-resource languages. Given that the main computational bottleneck is MLM-ing for learning language adapters, we have started from the subset of languages represented in our evaluation datasets (C2) for which pretrained language adapters (regular, non-orthogonal) are already available online (C3) Pfeiffer et al. (2020b). The final list of target languages, available in Table 1 with their corresponding language codes, comprises 10 languages from 5 geographical macro-areas Dryer and Haspelmath (2013); Ponti et al. (2020a) representing 8 distinct language families and covering at least one language for each broad morphosyntactic language type (isolating, introflexive, fusional, agglutinative), satisfying C1. Finally, in NER evaluations we include three truly low-resource languages: Quechua, Ilocano, and Meadow Mari (C4).
Training and Evaluation: Technical Details. We rely on the AdapterHub library Pfeiffer et al. (2020b) built on top of the Transformers library Wolf et al. (2020) in all experiments based on the MAD-X framework. For training language and task adapters we follow the suggestions from prior work Houlsby et al. (2019); Pfeiffer et al. (2020c). For orthogonal language adapters, we conduct MLM-ing on the Wikipedia data of each language. For learning task orthoadapters, we rely on the standard training portions of our task data in English.Unlike prior work that effectively violates the zero-shot assumption by doing model selection using development data in the target language Conneau et al. (2020); Keung et al. (2020), we select task (ortho)adapters solely based on the performance on the source language (i.e, English) dev set. We provide the full details of our training and fine-tuning procedures (including the details on hyperparameter search), in the Appendix A.
Results and Discussion
The results of zero-shot transfer for the three evaluation tasks are summarized in Table 2 (for the XNLI task), Table 3 (UD-POS), and Table 4 (NER). For all three tasks, besides scores per language, we also include the average results of zero-shot transfer experiments, i.e., excluding English test scores as the non-transfer experiment (the AVGz column). A starting observation based on the comparison with the full fine-tuning variant (full-ft) confirms findings from previous work Pfeiffer et al. (2020a), further validating the use of the more efficient adapter-based fine-tuning: the scores with adapter-based variants are on a par with or even higher than the scores reported with full-ft across the board.
Regular vs Orthogonal Language Adapters. Our initial set of result indicates that the usefulness of our orthogonal language adapters (L-ORT variants) does depend on the task at hand and its complexity. As an encouraging finding, we observe consistent gains in cross-lingual NLI, which is arguably the most complex (reasoning) task in our evaluation and, unlike UD-POS and NER, requires successful modeling of complex semantic compositionality in a (target) language. The gains (at least +1 accuracy point) are observed for 4 out of 5 target languages when doing zero-shot transfer with the l-ort+t-noo variant. This variant also yields highest average transfer performance, and slight (but statistically insignificant) gains on English NLI. However, the picture is less clear for UD-POS and NER: l-ort+t-noo does have a slight edge over the baseline fully non-orthogonal variant (l-noo+t-noo) in UD-POS, but this seems to mostly stem from large gains in Chinese. In a similar vein, while the l-ort+t-noo variant is the best performing variant in NER experiments on average, the gains over the baseline l-noo+t-noo are slight, and inconsistent across languages (e.g., large gains on ilo, improvements on ar and sw, but some decrease on qu and mhr).
We speculate that this is mostly due to the nature and complexity of the task at hand. In order to perform cross-lingual transfer for language inference, the underlying MMT must capture and leverage more language-specific nuances than for sequence labeling tasks such as POS-tagging or NER. By enforcing the capture of non-redundant information in the additional language-specific adapter modules, we allow the model to store additional and, more importantly, novel target language information. While the same information is available also for NER and POS tagging, these tasks require ’shallower’ language-specific knowledge Lauscher et al. (2020) and this is why more complex target language-specific knowledge captured in orthoadapters (compared to regular non-orthogonal language adapters) does not make a difference.
Orthogonal Task Adapters, on the other hand, display a different behavior, but we can again largely relate it to the properties of the evaluation tasks. First, the use of task orthoadapters seems detrimental across the board in the MAD-X setup for XNLI (compare l-noo+t-ort vs. l-noo+t-noo as well as l-ort+t-ort vs. l-ort+t-noo in Table 2), and also yields no real benefits in the simpler setup with no language adapters (t-ort vs. t-noo). We hypothesize that the two main objectives – (i) MLM for the original MMT pretraining and language adapter training, and (ii) cross-entropy loss for the whole sequence for NLI – are structurally too different for the orthogonality loss to capture any additional task-related information. In fact, we speculate that the orthogonal loss might have emphasized this discrepancy between the objectives and actually even hurt the final performance.
However, task orthoadapters seem to be useful for UD-POS transfer, with substantial gains reported on 3/5 target languages – Arabic, Chinese, Hindi, all of which have non-Latin scripts, while there is no change in performance for Estonian and Turkish. Combining task orthoadapters with language orthoadapters, however, does deteriorate the performance. The overall trend is even more complex in NER transfer: while there are clear hints that using orthoadapters is useful for some languages and some model variants, there is still a substantial variance in the results. We partially attribute it to the documented volatility of the WikiAnn evaluation set, especially for low-resource languages Pfeiffer et al. (2020c).
In general, our results suggest that explicitly controlling for the information that gets captured in the adapter modules can have a positive impact on cross-lingual transfer via MMTs. The optimal use of the orthogonality loss, however, seems to be largely target-language- and task-dependent, warranting further investigations in future work.
Conclusion and Future Work
We have investigated how orthogonality constraints impact downstream performance of zero-shot cross-lingual transfer via massively multilingual transformers (e.g., XLM-R) for three standard tasks: NLI, POS, and NER. Relying on the standard adapter-based transfer techniques, we have introduced the idea of orthogonal language and task adapters (or orthoadapters): we explicitly enforce the information stored in the parameters of language and task adapters to be orthogonal to the information already stored in the pretrained MMT. With such an orthogonality mechanism in place we should be able to encode novel rather than redundant information. Our results have demonstrated the validity of orthoadapters, especially in the most complex XNLI task, although the optimal adapter setup seems language- and task-dependent.
Our work has pointed to the importance of enforcing and controlling what information gets stored in the adapter modules for improved cross-lingual transfer, but it has only scratched the surface. In future work, we will investigate more sophisticated orthogonality losses Bousmalis et al. (2016); Liu et al. (2017) and techniques such as gradual adapter unfreezing Howard and Ruder (2018); Peters et al. (2019). We will also explore the usefulness of orthoadapters in other tasks as well as for domain adaption Rücklé et al. (2020).
References
Appendix A Training Details
We processed the data for all tasks using the preprocessing pipeline provided with the XTREME benchmark Hu et al. (2020).https://github.com/google-research/xtreme
XNLI. In NLI training (i.e., for XNLI transfer) were trained for 30 epochs with the batch size of 32. Maximum sequence length was 128 input tokens. Gradient norms were clipped to 1.0.
UD-POS. We trained for 50 epochs with batch size of 16. Maximum sequence length was 128 input tokens. Gradient norms were clipped to 1.0.
NER. We trained for 100 epochs with the batch size of 16. Maximum sequence length was 128 input tokens. Gradient norms were clipped to 1.0.
A.2 Experimental Setup without Language Adapters.
For the “non-MAD-X” experimental setup (i.e., the setup without language adapters, see §3), we relied on our own implementation of the adapter module. The bottleneck size for the task adapter was set to . For (X)NLI, we searched the following learning rate grid: ; for UD-POS and WikiAnn the corresponding learning rate grid was . For task orthoadapters, we searched the following additional learning rate grid for the orthogonal loss optimizer: .
A.3 Full Experimental Setup
For the more complex multi-adapter setup based on the MAD-X framework (i.e., with both language and task adapters), we utilized the Adapter-Transformers library and the underlying AdapterHub service Pfeiffer et al. (2020b).
Task Adapters and Orthoadapters. We followed the recommendation from the original paper Pfeiffer et al. (2020c). We utilized the Pfeiffer configuration found in the Adapter-Transformers library with the adapter dimensionality of 48. Due to the computational constraints, the learning rate grid search took into account best settings observed in the baseline experiments. For XNLI our main learning rate was set at a well-performing . For UD-POS and WikiAnn, due to more instability, we tested the learning rate grid of . For the task orthoadapters, we used the same learning rate for the orthogonality loss as for the non-MAD-X setup: .
Regular Language Adapters. We utilized the pretrained language adapters readily available via the AdapterHub service Pfeiffer et al. (2020b). These language adapters have the dimensionality of 384. They were trained (while the rest of the model was frozen) by executing the MLM-ing for 250.000 iterations on the Wikipedia data in the target language.
Orthogonal Language Adapters. We started from the MLM-ing training script for training language adapters provided by the Adapter-Transformers library and trained language orthoadapters on the Wikipedia data, relying on the setup of AdapterHub’s regular language adapters (dimensionality 384, 250,000 iterations). Due to computational constraints we reduced the maximum sequence length of the input to 128 tokens, while the batch size was 8. Finally, for the main optimizer and orthogonality loss optimizer we used the learning rates of and , respectively.