Counter-Interference Adapter for Multilingual Machine Translation
Yaoming Zhu, Jiangtao Feng, Chengqi Zhao, Mingxuan Wang, Lei Li
Introduction
Machine translation (MT) is a core task in natural language processing. In recent years, neural machine translation (NMT) approaches have made tremendous progress and takes the lead in the field Bahdanau et al. (2015); Vaswani et al. (2017); Johnson et al. (2017). Conventionally, each NMT model only tackles a single language direction (e.g. English German). A commonly used model like Transformer has million parameters. Therefore, Translating language pairs requires training models separately for each direction, resulting in total parameters. The huge size of all models for every language direction can be too costly to deploy, considering more than 100 popular languages worldwide. Hence, developing a single unified yet parameter-efficient model for multilingual machine translation, enabling the translation of multiple directions, becomes crucially important. There is much effort towards multilingual machine translation. Johnson et al. (2017) first proposed training a single multilingual model via an additional target language tag, which became the paradigm for training multilingual MT models henceforth (Gu et al., 2018; Tan et al., 2018, 2019; Siddhant et al., 2020). Their approach is simple and parameter-efficient, however, it often lags behind separately trained bilingual models, especially on resource-rich language pairs, where the phenomenon is known as performance degradation (Aharoni et al., 2019).
This paper analyzes the performance degradation between multilingual MT models and separate bilingual ones. Prior research suggests that unified embedding on a joint vocabulary leads to meaning conflation deficiency in multilingual training (Camacho-Collados and Pilehvar, 2018). For example, words with identical spelling may have distinct meanings in different languages — bride in English refers to a woman soon to get married while in French it means horse bridle. When it comes to machine translation, the effect goes beyond word embedding. As a single model has bounded capacity, the multilingual learning may cause negative influences among shared parameters Liu et al. (2017); Zhang et al. (2020). We conjecture that such a performance degradation is due to the interference across languages brought by joint training on multiple language directions. Such interference affects both joint token embedding and representations from intermediate layers. We argue that resolving the interference is critical to improving multilingual translation performance.
Inspired by the insights above, we propose Counter-Interference Adapter for multilingual machine Translation (CIAT). The CIAT includes a major multilingual base model (i.e., a multilingual Transformer), which is universally pre-trained on multiple language directions, and two kinds of designed adapter modules, which are trained on specific language directions. Specifically, we propose embedding adapter and layer adapter to reduce multilingual interference among embedding and intermediate layers, respectively. We also seek a new parallel connection method for adapter units, which is more effective for multilingual MT than the previous series connection. We validate CIAT among three real-world datasets and observe that CIAT gains significant improvements in translation performance over all other multilingual baselines.
Our contributions are as follows: a) We analyze performance degradation in multilingual NMT and formulate it as two issues; b) We propose CIAT, an adapter-based framework to tackle the two issues and enhance the performance of the multilingual model with small amounts of extra parameters; c) We demonstrate the efficacy of CIAT through extensive experiments on IWSLT, OPUS-100, and WMT benchmark datasets, surpassing other multilingual models over most of the translation directions.
Related Work
The multilingual MT enjoys a rich research history, dating back to the age of statistical machine translation Gao et al. (2002); Haffari and Sarkar (2009); Seraj et al. (2015). In recent years, the prosperity of neural machine translation (NMT) has led to the growing prominence and popularity of multilingual MT systems. The encoder-decoder framework has made the de facto standard for NMT Bahdanau et al. (2015); Vaswani et al. (2017). Dong et al. (2015) did the pioneering work on extending conventional NMT to one-to-many translation, where the authors added a distinct decoder for each target language. Firat et al. (2016) further extended such framework into many-to-many settings by building exclusive encoders and decoders for each language. Those attempts still faced problems such as low parameter utilization. On the other hand, Lee et al. (2017) treated all sources as the same language by translating on a character level. However, it only meets the many-to-one scenarios. Johnson et al. (2017) managed to train a single model that applied to multiple translation directions. Their solution is relatively simple: they attached a dedicated token at the beginning of the source sentence to specify the target language, while the rest of the model was shared among all languages. The paper has set a milestone of multilingual MT and has become the basis for most subsequent work.
Recent studies paid more attention to the performance improvement of multilingual models based on Johnson et al. (2017)’s effort. Several improved the model with external knowledge from human or other models: Tan et al. (2018) boosted the multilingual model by knowledge distillation, Tan et al. (2019) pre-clustered languages to assist similar languages. Several studies enhance the model from data: Xia et al. (2019) and Siddhant et al. (2020) conducted data augmentation to low-resource languages via related high-resources or monolingual data. Taitelbaum et al. (2019) improved translation with relevant auxiliary languages. Some other studies enhanced the Transformer model by introducing language-aware modules and learning language-specific representation Wang et al. (2019); Zhu et al. (2020).
Our design derives from the residual adapters of the domain adaptation task. Concretely, Rebuffi et al. (2017) proposed the residual adapters in the computer vision area. They appended small networks (named adapters) to a pre-trained base network and only tuned the adapter on the specific task. Houlsby et al. (2019) adopted the idea into NLP domain adaptation tasks and designed the adapter for the Transformer, as shown in Fig. 2(a). Bapna and Firat (2019) further extended the model to MT domain adaptation, and they regarded multilingual MT as a domain adaptation task. Based on their design, Philip et al. (2020) proposed the monolingual adapter for easy extension to new pairs, and Zhang et al. (2021) introduced conditional language-specific routing strategy(CLSR) to enhance model capacity in language-specific representation. However, their adapter designs followed a serial connection manner, which might be limited for multilingual MT. We will discuss this in the following sections.
Challenges on Multilingual Machine Translation
Currently, Transformer Vaswani et al. (2017) gains popularity and becomes the paradigm for state-of-the-art NMT systems. Here, we follow the recent implementations Klein et al. (2017); Vaswani et al. (2018) of the pre-norm transformer, whose layer normalization is applied to the input of each sub-layer. The transformation of -th sub-layer taking as input can be formulated as:
The classic multilingual translation approach Johnson et al. (2017) takes sentences from all language pairs and results in a universal model that translates across various languages. However, as previous researches have discovered, the multilingual model yields an inferior performance on high-resource languages compared to the bilingual models under the same configuration. In this paper, we reconsider the limitations of multilingual models and attribute the performance degradation to the following two issues:
Camacho-Collados and Pilehvar (2018) addressed the meaning conflation deficiency problem of the word embedding as a single vector is limited for representing polysemy. We extend the meaning conflation deficiency into the multilingual scenario. Generally, words/tokens may have unrelated or even opposite meanings in different languages. For example, “娘” denotes mother in Chinese but daughter in Japanese. As the word embeddings are usually jointly trained on a multilingual corpus, representing a multilingual word with just one single vector may burden the model’s semantic representation. We refer to the problem as multilingual embedding deficiency.
Besides the word embedding, the insufficient capacity of a single NMT model also bottlenecks its performance on multilingual tasks Aharoni et al. (2019); Zhang et al. (2020). The parameter-sharing among different languages may be a potential cause of negative interference Liu et al. (2017); Wang et al. (2020). We here formulate this phenomenon as multilingual interference effects; that is, when a single model tries to learn multiple languages simultaneously, the extracted language features interfere with each other impose adverse effects upon overall performance. Accordingly, we regard the model trained on bilingual data offers the approximately optimal solution on language representation, compared to which the representation of the multilingual model is bias-influenced:
where denotes the parameters of bilingual baselines and indicates the multilingual model. is the interference noise in the -th layer.
Proposed Method
We propose a Counter-interference Adapter for Multilingual Machine Translation (CIAT) to address the two issues mentioned above with adapter-based architectures: embedding adapter and layer adapter. Figure 1 illustrates the overall architecture of CIAT, which we will describe in detail.
As discussed in section 3, a jointly trained multilingual word embedding could be problematic as a word may have different meanings among multiple languages. Empirically, we can fine-tune the whole embedding matrix for each language pair to address the multilingual embedding deficiency. However, tuning the whole matrix is quite expensive as the embedding matrix occupies a large part of model parameters, which also violets the advantage of parameter sharing in multilingual NMT. We hence introduce the embedding adapter to approximate the fine-tuned embedding matrix with much fewer parameters:
Fig. 1 shows the layout of the embedding adapter, including a layer normalization Ba et al. (2016) and a fully connected feed-forward neural network. Following the suggestions of Houlsby et al. (2019), we choose bottle-neck architecture for the adapter module to save the parameters: the first layer down-projects the embedding dimension to a smaller size , while the second layer projects it back to dimension. We select ReLU Nair and Hinton (2010) as the activation function for the middle layer while using no activation for the output layer.
2 Parallel De-noise Layer Adapter
As formulated in Eq. 2, the multilingual representation is regraded as a bias-influenced one compared to the bilingual representation . To alleviate the interference, we introduce the layer adapter to model the bias term :
Eq. 4 means that the layer adapter shares the same input as the sub-layer and de-noises the output . As a result, the output of -th sub-layer is adapted to .
We connect layer adapters parallel to sub-layers, as shown in Fig. 1. The architecture of the layer adapter is the same as the embedding adapter, except for the removed layer normalization. In practice, we find that the layer adapter can share the layer normalization structure with the corresponding sub-layer to achieve the best performance.
As the first study introduced adapter networks to machine translation, Bapna and Firat (2019) also conducted experiments on multilingual machine translation. They append adapters serial to the model architecture with a residual connection as illustrated in Fig. 2(a), while we argue that our parallel connection is more suitable for multilingual machine translation. Compared to Eq. 4, we can formulate the serial style adapters as:
where the adapter receives the bias-influenced hidden states other than the original input . However, the bias-influenced may not be distinguishable for training adapters, which is especially the case when the multilingual model is inferior, making fail to capture enough information of languages.
In contrast, our parallel design de-noise the bias-influenced term pre to the sub-layers. Corresponding to Eq. 4, the parallel layer adapters receive the same input as the sub-layers and de-noise directly to the output, which is a more intuitive and natural design for de-noising , since the adapter is independent of the sub-layer output. The parallel adapter can also be regarded as the low-rank “patch” for the corresponding sub-layer, which adjusts the parametrization of the high-rank sub-layer to the specific language pair and fix the multilingual interference.
3 Model Training
The training process of the whole model consists of two phases: the pre-training on the multilingual model and the learning of the adapter modules. First, we pre-train the standard Transformer on the entire corpus, making a universal multilingual model. The parameters of the Transformer are frozen once the model converges. Then, we “plug-in” the randomly initialized adapters for each specific language pair and only fine-tune the adapter parameters on this pair. Since the base Transformer model is frozen, each plugged adapter’s learning process is independent of other adapters. During the inference stage, we only apply the base model and the corresponding adapter to translate sentences into the target language. Note that Bapna and Firat (2019) only fine-tune their model on high-resource pairs, while we find CIAT can be applied to both high-resource and low-resource pairs. All the adapters are plug-able during the inference stage: disabling the adapter will degenerate the model into a basic multilingual translation model, and when the model is required to translate a specific pair, we just “plug-in” the specific adapter into the base model.
Theoretically, the well-trained adapted network guarantees a better performance compared to the multilingual base. When the layer adapter is disabled (i.e.. when all parameters are zero), the model is reduced to the multilingual base model. With proper training, the layer adapter should boost the model performance.
Experiments
We conducted experiments on three multilingual translation datasets to show the effectiveness of CIAT.
We focused on two mainstream multilingual cases: many-to-English and English-to-many, since the many-to-many case can be bridged via English as a pivot. We collected the following three datasets for our experiments:
IWSLT https://wit3.fbk.eu is a small dataset from TED talks, where we used 8 languages English from year 2014 to 2016 release.
OPUS-100 Zhang et al. (2020) http://opus.nlpl.eu/OPUS-100.php is an English-centric dataset covering 99 languages English pairs. We selected 20 language pairs, 17 of which have 1 million data samples while 3 language pairs are under low resource setting.
WMT Barrault et al. (2019) datasets are also involved, which contains five language pairs ranging from the year 2014 to 2019.
For simplicity, we use the ISO 639-1 code as the abbreviation for language names. The detailed data statistics are listed in the Appendix.
2 Implementation Details
For each dataset, we tokenize sentences using SentencePiece Kudo and Richardson (2018) jointly learned on the source and target side, and we set vocabulary size to 32,000. For model setup, we follow the same configuration as Tan et al. (2018) on IWSLT, including 2 layers for both encoder and decoder. The embedding dimension was 256, and the size of feed-forward hidden units was 1,024. The attention head was set to 4 for both self-attention and cross-attention. For OPUS-100 and WMT, we follow the standard Transformer-Big setting Vaswani et al. (2017), including 6 layers for encoder and decoder. The embedding dimension, feed-forward hidden size, and attention head were set to 1024, 4096, and 16, respectively. The hidden state’s dimension of the adapters’ inner layer are set to be half of the embedding size, i.e. 128 for IWSLT and 512 for OPUS-100 and WMT. We use Adam optimizer Kingma and Ba (2015) with the same schedule algorithm as Vaswani et al. (2017). During Inference, we use a beam width of 4 and length penalty of 0.6.
All our experiments are evaluated by tokenized BLEU Papineni et al. (2002) using multi-bleu.perl https://github.com/moses-smt/mosesdecoder/blob/master/scripts/generic/multi-bleu.perl. We implement our models via TensorFlow Abadi et al. (2016) and train models on NVIDIA Tesla V100 GPUs.
3 Main Results
We compared our model with several strong baselines and effective models:
Bilingual: The model is trained with only bilingual data with the same model configuration, which serves as the strong benchmark.
Multilingual Johnson et al. (2017): The data of all language pairs are mixed to train the model.
Knowledge Distillation (KD) Tan et al. (2018) https://github.com/RayeRen/multilingual-kd-pytorch. Note that we remove the lowercase option in their preprocessing script.: The bilingual models are first trained as the teachers, then the multilingual models are trained as students. We only conduct KD on IWSLT due to computational resource limitations.
Serial Bapna and Firat (2019): We re-implement the series adapter model as illustrated in Fig 2(a), which made the first attempt on applying adapters to machine translation.
To further analyze the impact of each component and compare the two designs of Serial and Parallel layer adapter under the same number of parameters, we introduce two model variants, named CIAT-layer and CIAT-block, and conduct the ablation study. CIAT-layer has the same number of parameters compared with Serial and CIAT-block removes embedding adapter. The ablation details are illustrated in the appendix.
We also reproduced a recently proposed adapter-based MT model, namely Mono Philip et al. (2020). The model also follow a serial adapter connection manner, which we replace with a parallel connection to validate the effectiveness of our parallel design. We show the comprehensive results of the ablation study and parallel variant of Mono in the Appendix. We here list the overall results of two model variants.
We present the BLEU score of three datasets on Table 2, 3 and 7 (Appendix) respectively, and summarize the overall results on Table 1. Since the number of parameters is quite different among distinct baseline and model variants, we also indicate parameters amount for each model with respect to the number of translation directions. We present our findings as follows:
On the IWSLT dataset, the bilingual baselines are significantly better than the multilingual models, which align with our expectations. To our surprise, we observe the opposite phenomenon on OPUS-100, which may be because (1) the domain of sentences in the OPUS-100 dataset is close. (2) Most languages selected have relatives from the same language family, which leads to promotion between the related languages Tan et al. (2019); while only zh has no similar language in the OPUS-100, and bilingual performs better in zh.
Among various models, the gaps in BLEU scores of anyen are smaller than that of enany. Meanwhile, anyen direction gain less improvement from CIAT and other baselines than enany direction. We attribute such phenomenon to the over-representation of English in the English-centric corpus Aharoni et al. (2019), so it is more difficult to improve the generation quality of English than other languages.
CIAT significantly outperforms other multilingual competitors among all datasets, and improve the BLEU score by at least 0.5 in 42 out of 64 language directions. On enany directions, the performance improvement is even more significant.
We also find that the performance of CIAT and baselines on different languages is also affected by language families and resource scarcity. According to Table 2, for languages of the same language family, CIAT can better improve performance (e.g., Spanish, Portuguese, French, and Italian get a greater improvement than Arabic and Persian). And for low-resource languages(e.g. Afrikaans, Belarusian), compared with Serial baseline, CIAT can bring more BLEU score improvement, especially when there are languages similar to these low-resources in the training set.
4 Discussion on Adapters
To obtain a comprehensive understanding of the adapters in CIAT, we conduct a series of analyses on embedding adapters and layer adapters respectively.
To study whether the embedding adapter alleviates multilingual embedding conflation, we calculate the Average Cosine Similarity(ACS) Lin et al. (2020) of words with the same meaning across different languages to verify if embedding adapter help to align cross-lingual synonyms. We select top frequent 1000 words of five language pairs from MUSE bilingual dictionarieshttps://github.com/facebookresearch/MUSE, and compare the ACS results between vanilla multilingual embedding and ones with the CIAT’s embedding adapter, where the models are trained on OPUS-100 dataset.
We plot the ACS results in Fig. 3 and observe two phenomena. Firstly, adding the embedding adapters increases ACS of the model among all selected pairs, which indicates that embedding adapters capture more semantic information between synonyms from different languages. The ACS improvement also suggests that embedding adapters indeed relieve the multilingual embedding conflation problem as the latent representation of synonyms between the two languages is brought closer with the auxiliary of the embedding adapter, which improves the overall performance. Secondly, compared to any en directions, en any pairs generally gain more improvement on ACS and BLEU scores, indicating embedding adapter is more effective on translating pivot language to other languages compared to the opposite direction.
4.2 The Influence of Layer Adapters
We also perform an extension experiment to discover the influence of layer adapters on the main model, taking enzh directions from OPUS-100 as the study cases.
We first plot the L2-norm ratio of the hidden state between the adapter and the base model across layers in Fig 4 to investigate how the “plug-in” adapter influence the multilingual model. In general, the adapters of decoder exerts a greater influence on the hidden states of the base model compared to the ones of the encoder, and the adapters provide a stronger signal in enzh compared to the opposite direction.
We further examine these adapters’ impact by re-evaluating the trained model with certain adapters from continuous layer spans removed, and we illustrate the BLEU score drop on Fig. 5. We find that removing the adapters on the decoder side raises a greater performance decline, consistent with the trend of the L2-norm ratio. Adapters of the enzh are more crucial to the multilingual models, which is in line with our experiments that adapters are more important for directions (Table 1). In addition, the decoder’s upper layers of the CIAT layer adapter have bigger impacts on the performance, consistent with Houlsby et al. (2019)’s findings on adapters for BERT model.
5 Parameter-Performance Trade-off
The bottle-neck adapter design utilizes a small middle layer to control parameter efficiency Houlsby et al. (2019), while empirically, a larger layer dimension improves the performance via increased capacity. We explore the parameter-performance trade-off by varying the dimension of adapters and illustrate the BLEU score over different layer adapter and embedding adapter size in Fig. 6 with different color, respectively. Here we discuss the trade-off on ende directions of IWSLT. The plot shows that the dimension of the layer adapters has a significant impact to the performance: as the dimension doubles, the BLEU increases 0.58 and 0.28 by an average in ende and deen directions respectively. In comparison, changing the embedding adapters’ dimension impacts less on the final performance. When the dimension is large already, expanding the size hardly increases the BLEU score. Considering that the parameter amounts of CIAT are small compared to the base model, and the final performance is sensitive to the layer adapter’s dimension, we regard expanding dimension to be regarded as a simple and effective way to improve the performance of CIAT.
Conclusion
This work analyzes the performance degradation problem in multilingual NMT systems and decomposes into multilingual embedding deficiency and multilingual interference effects. We then propose a novel framework to deal with degradation, named Counter-interference Adapter for Multilingual machine translation (CIAT). CIAT alleviates two issues above respectively by introducing two kinds of adapters.
We validate the effectiveness of CIAT on three multilingual translation datasets, where the results show that CIAT improves the performance of the multilingual NMT model on various translation directions. The experiments also demonstrate that CIAT variants outperform several strong baselines, approving our analysis and framework design. Furthermore, we investigate the behavior and utility of each component via empirical studies.
References
Appendix
We give the detailed statistics about the dataset used in Sec. 5
We almost follow the Tan et al. (2018)’s scripthttps://github.com/RayeRen/multilingual-kd-pytorch/blob/master/data/iwslt/raw/prepare-iwslt14.sh except that we removed their lowercase option. We collect training sets of 8 languages ranging from the year 2014 to 2016 and use the official valid/test set. We list the number of samples in training set in Table 4.
We collect data from Zhang et al. (2020)’s release https://object.pouta.csc.fi/OPUS-100/v1.0/opus-100-corpus-v1.0.tar.gz , and use its official valid/test set. Among the 20 language pairs selected, 17 have 1 million training samples while three language pairs are of low resources, which are be(67k), nb(142k), and af(275k).
We list the year of the training, valid and test set of each language in Table 5. Table 6 illustrate the number of samples in the training set.
2 Detailed Experiment Results of Ablation Study and Model Variants
To further study the efficacy of each component in CIAT, we propose two variants as ablation study:
CIAT-layer: This variant keeps only the layer adapter and removes all the embedding adapter.
CIAT-basic: Besides removing all embedding adapters, this variant introduces only one layer adapter for each attention block. We design this variant to compare with Serial Bapna and Firat (2019) under a similar amount of parameters to determine the effectiveness of our parallel connection further.
We present the detailed results on two CIAT variants on Table 7, 8 and 9 respectively. We give two major findings of ablation study: a) Compared with CIAT, CIAT-layer suffers from degradation in most language pairs (51 out of 66), especially in low-resource corpora(IWSLT). It further shows the effectiveness of the embedding adapter. b) With the same amount of parameters, CIAT-basic surpass Serial in both overall performance and the number of improved pairs among all datasets. It shows that parallel is a more suitable adapter connection schema for multilingual machine translation.
3 The Effectiveness of Parallel Connection on Other Adapter Model
As mentioned in Section 5, we also substitute another adapter-based model with parallel connection methods to illustrate our proposed parallel connection is more suitable for multilingual machine translation. We re-implement Philip et al. (2020)’s work, where they proposed monolingual adapter which is specific to the source/target language other than the translation pair. We present the results of their serial connection and our parallel variants in Tab. 7 as Mono-Serial and Mono-Parallel.
We find parallel connection boosts the Mono adapter in 15 out of 16 language pairs in the IWSLT data set, the improvement is even more prominent in the en any ones. The results further prove that our parallel layer adapter can provide improvement for all multilingual adapter models.
4 Case Study
We also conduct qualitative analysis by case study. We invite two German speakers to compare the translation of a news report from the WMT test set. The contestants are generated by CIAT, Serial and vanilla multilingual. We find German translation of CIAT are more favored by human annotators for the following reasons: a) The sentence tense and clause pattern fit the original sentence; b) CIATtend to use set phrases instead of simple expressions; c) CIATbetter captures the relationship between modifiers and the subjects.