Enhancing Cross-lingual Transfer by Manifold Mixup

Huiyun Yang, Huadong Chen, Hao Zhou, Lei Li

Introduction

Many natural language processing tasks have shown exciting progress utilizing deep neural models. However, these deep models often heavily rely on sufficient annotation data, which is not the case in the multilingual setting. The fact is that most of the annotation data are collected for popular languages like English and Spanish (Ponti et al., 2019; Joshi et al., 2020), while many long-tail languages could hardly obtain enough annotations for supervised training. As a result, cross-lingual transfer learning (Prettenhofer & Stein, 2011; Wan et al., 2011; Ruder et al., 2019) is crucial, transferring knowledge from the annotation-rich source language to low-resource or zero-resource target languages. In this paper, we focus on the zero-resource setting, where labeled data are only available in the source language.

Recently, multilingual pre-trained models (Conneau & Lample, 2019; Conneau et al., 2020a; Xue et al., 2020) offer an effective way for cross-lingual transfer, which yield a universal embedding space across various languages. Such universal representations make it possible to transfer knowledge from the source language to target languages through the embedding space, significantly improving the transfer learning performance (Chen et al., 2019; Zhou et al., 2019; Keung et al., 2019; Fang et al., 2020). Besides, Conneau et al. (2018) proposes translate-train, a simple yet effective cross-lingual data augmentation method, which constructs pseudo-training data for each target language via machine translation. Although these works have achieved impressive improvements in cross-lingual transfer (Hu et al., 2020; Ruder et al., 2021), significant performance gaps between the source language and target languages still remain (see Table 1). Hu et al. (2020) refers to the gap as the cross-lingual transfer gap, a difference between the performance on the source and target languages.

To investigate how the cross-lingual transfer gap emerges, we perform relevant analyses, demonstrating that transfer performance correlates well with the cross-lingual representation discrepancy (see Section 3 for details). Here the cross-lingual representation discrepancy means the degree of difference between the source and target language representations in the universal embedding space. As shown in Figure 1(a), in translate-train, the representation distribution of Spanish almost overlaps with English, while Arabic shows a certain representation discrepancy compared with English and Swahili performs larger discrepancy, where translate-train achieves 88.6 accuracy on English, 85.7 on Spanish, 82.2 on Arabic and 77.0 on Swahili. Intuitively, a larger representation discrepancy could lead to a worse cross-lingual transfer performance.

In this paper, we propose the Cross-Lingual Manifold Mixup (X-Mixup) approach to fill the cross-lingual transfer gap. Based on our analyses, reducing the cross-lingual representation discrepancy is a promising way to narrow the transfer gap. Given the cross-lingual representation discrepancy is hard to remove, X-Mixup directly faces the issue and explicitly accommodates the representation discrepancy in the neural networks, by mixing the representation of the source and target languages during training and inference. With X-Mixup, the model itself can learn how to escape the discrepancy, which adaptively calibrates the representation discrepancy and gives compromised representations for target languages to achieve better cross-lingual transfer performance. X-Mixup is motivated by robust deep learning (Vincent et al., 2008), while X-Mixup adopts the mixup (Zhang et al., 2018) idea to handle the cross-lingual discrepancy.

Specifically, X-Mixup is designed upon the translate-train approach, faced with the exposure bias (Ranzato et al., 2016) problem and data noise problem. During training, the source sequence is a real sentence and the target sequence is a translated one, while situations are opposite during inference. Besides, the translated text often introduces some noises due to imperfect machine translation systems. To address them, we further impose the Scheduled Sampling (Bengio et al., 2015) and Mixup Ratio in X-Mixup to handle the distribution shift problem and data noise problem, respectively.

We verify X-Mixup on the cross-lingual understanding benchmark XTREME (Hu et al., 2020), which includes several understanding tasks and covers 40 languages from diverse language families. Experimental results show X-Mixup achieves 1.8% performance gains across different tasks and languages, comparing with strong baselines. It also reduces the cross-lingual representation discrepancy significantly, as Figure 1(b) shows.

Related Work

Multilingual Representation Learning Recent studies have demonstrated the superiority of large-scale pre-trained multilingual representations on downstream tasks.

Multilingual BERT (mBERT; Devlin et al., 2019) is the first work to extend the monolingual pre-training to the multilingual setting. Then, several extensions achieve better cross-lingual performances by introducing more monolingual or parallel data and new pre-training tasks, such as Unicoder (Huang et al., 2019), XLM-R (Conneau et al., 2020a), ALM (Yang et al., 2020), MMTE (Siddhant et al., 2020), InfoXLM (Chi et al., 2020), HICTL (Wei et al., 2020), ERNIE-M (Ouyang et al., 2020), mT5 (Xue et al., 2020), nmT5 (Kale et al., 2021), AMBER (Hu et al., 2021) and VECO (Luo et al., 2021). They have been the standard backbones of current cross-lingual transfer methods.

Cross-lingual Transfer Learning Cross-lingual transfer learning (Prettenhofer & Stein, 2011; Wan et al., 2011; Ruder et al., 2019) aims to transfer knowledge learned from source languages to target languages. According to the type of transfer learning (Pan & Yang, 2010), previous cross-lingual transfer methods can be divided into three categories: instance transfer, parameter transfer, and feature transfer. The cross-lingual transferability improves a lot when engaged with the instance transfer by translation (i.e. translate-train, translate-test) or other cross-lingual data augmentation methods (Singh et al., 2019; Bornea et al., 2020; Qin et al., 2020; Zheng et al., 2021). Chen et al. (2019) and Zhou et al. (2019) focus on the parameter transfer to learn a share-private model architecture. Besides, other works implement the feature transfer to learn the language-invariant features by adversarial networks (Keung et al., 2019; Chen et al., 2019) or re-alignment (Libovický et al., 2020; Zhao et al., 2020). X-Mixup utilizes both the instance transfer and feature transfer, which is based on the translate-train data augmentation approach and implements the feature transfer by cross-lingual manifold mixup.

Mixup and Its Variants Mixup (Zhang et al., 2018) proposes to train models on the linear interpolation at both the input level and label level, which is effective to improve the model robustness and generalization. Generally, the interpolated pair is selected randomly. Manifold mixup (Verma et al., 2019) performs the interpolation in the latent space by conducting the linear combinations of hidden states. Previous mixup methods (Chen et al., 2020; Jindal et al., 2020) focus on the monolingual setting. However, X-Mixup focuses on the cross-lingual setting and faces many new challenges (see Section 4 for details). Besides, in contrast to previous mixup methods, X-Mixup mixes the parallel pairs, which share the same semantics across different languages. As a result, the choice of parallel pairs for interpolation can build a smart connection between the source and target languages.

Analyses of the Cross-lingual Transfer Performance

In this sectionIn our analyses, we take English as the source language, and the dissimilar language is the language which is dissimilar to English., we concentrate on the cross-lingual transfer performance and find it is strongly associated with the cross-lingual representation discrepancy. Firstly, we observe the cross-lingual transfer performance on different target languages and propose an assumption. Then we conduct qualitative and quantitative analyses to verify it.

Although previous studies (Hu et al., 2020; Ruder et al., 2021) have shown impressive improvements on cross-lingual transfer, the cross-lingual transfer gap is still pretty large, more than 16 points in Hu et al. (2020). Furthermore, results in Table 1 show the performance of low-resource languages and dissimilar languages fall far behind other languages in cross-lingual transfer tasks.

Compared with English, the representations of other languages, especially low-resource languages, are not well-trained (Lauscher et al., 2020; Wu & Dredze, 2020), because high-resource languages dominate the representation learning process, which results in the cross-lingual representation discrepancy. Besides, dissimilar languages often show differences in language characteristics (like vocabulary, word order), which also leads to the representation discrepancy. As a result, we assume that the cross-lingual transfer performance is closely related to the representation discrepancy between the source language and target languages.

Following Conneau et al. (2020b), we utilize the linear centered kernel alignment (CKA; Kornblith et al., 2019) score to indicate the cross-lingual representation discrepancy

where X and Y are parallel sequences from the source and target languages, respectively. A higher CKA score denotes a smaller cross-lingual representation discrepancy.

To verify our assumption, we perform qualitative and quantitative analyses on the relationship between the CKA score and cross-lingual transfer performance. Figure 3 in Appendix B indicates a higher CKA score tends to induce better cross-lingual transfer performance. We also calculate the Spearman’s rank correlation between the CKA score and the transfer performance in Table 2, which shows a strong correlation between them. Both the trend and correlation score confirm the cross-lingual transfer performance is highly related to the cross-lingual representation discrepancy.

Methodology: X-Mixup

Based on the aforementioned analyses, we believe that reducing the cross-lingual representation discrepancy is the key to filling the cross-lingual transfer gap. In this section, we propose X-Mixup to explicitly reduce the representation discrepancy by implementing the manifold mixup between the source language and target language. With X-Mixup, the model can adaptively calibrate the representation discrepancy and give compromised representations for target languages. This section will first introduce the overall architecture of X-Mixup and its details. After that, the training objectives and inference process will be shown.

Figure 2 illustrates the overall architecture of X-Mixup. Sequences from the source and target languages are first encoded separately. Then within the encoder, X-Mixup implements the manifold mixup between the paired sequences (original sequence and its translation) within a specific layer, where Mixup Ratio controls the degree of mixup and Scheduled Sampling schedules the data sampling process during training.

Basic Model We use mBERT (Devlin et al., 2019) or XLM-R (Conneau et al., 2020a) as the backbone model. Within each layer, there are two sub-layers: the multi-head attention layer and the feed-forward layerIn this section, the feed-forward layer is omitted for simplification., followed by the residual connection and layer norm. We use the same multi-head attention layer (see details in Appendix A.1) as BERT (Devlin et al., 2019), where inputs are query, key, and value respectively. In layer l+1l+1, the hidden states of the source sequence xS\bm{x}_{S} and target sequence xT\bm{x}_{T} are acquired by the multi-head attention

Manifold Mixup To reduce the cross-lingual representation discrepancy, a straightforward idea is to find compromised representations between the source and target languages. It’s difficult to find such representations because of varying degrees of differences across languages, like vocabulary and word order. However, manifold mixup provides an elegant way to get intermediate representations by conducting linear interpolation on hidden states.

To extract target-related information from the source hidden states, the target hidden states are used as the query, and the source hidden states are used as the key and value. This cross-attention process is computed as

which shares parameters with the multi-head attention. The manifold mixup process mixes the target hidden states hTl+1\bm{h}^{l+1}_{T} and the source-aware target hidden states hT∣Sl+1\bm{h}^{l+1}_{T|S} based on the mixup ratio λ\lambda

where λ\lambda is an instance-level parameter, ranging from 0 to 1, and indicates the degree of manifold mixup. LN denotes the layer norm operation.

Mixup Ratio The machine translation process may change the original semantics and introduce data noises in varying degrees (Castilho et al., 2017; Fomicheva et al., 2020). Thus we introduce the translation quality modeling in the mixup process to handle this problem. Following Fomicheva et al. (2020), we use the entropy of attention weights to measure the translation quality

II is the number of target tokens and JJ is the number of source tokens. Lower entropy implies better cross-lingual alignment and higher translation quality.

To introduce the translation quality modeling into the manifold mixup process, we compute the mixup ratio as λ=λ0⋅σ[(H(A)+H(A⊤))W+b]\lambda=\lambda_{0}\cdot\sigma[(\text{H}(\bm{A})+\text{H}(\bm{A}^{\top}))W+b], where σ\sigma is the sigmoid function, and WW, bb are trainable parameters. λ0\lambda_{0} is the max value of the mixup ratio, which is set to 0.5 in this paper. We consider two-way alignment in the translation quality modeling, i.e. H(A)\text{H}(\bm{A}) and H(A⊤)\text{H}(\bm{A}^{\top}).

where p∗p^{*} is decreasing during training to match the situation of inference. We utilize the inverse sigmoid decay (Bengio et al., 2015), which decreases p∗p^{*} as a function of the index of mini-batch.

2 Final Training Objective

The training loss is composed of two parts: the task loss and the consistency loss

where MSE(⋅)\text{MSE}(\cdot) is Mean Squared Error and KL(⋅)\text{KL}(\cdot) is Kullback-Leibler divergence. r∗\bm{r}_{*} is the sequence representationWe utilize the mean pooling of the last layer’s hidden states as the sequence representation, which is independent of the sequence length. and p∗\bm{p}_{*} is the predicted probability distribution of downstream tasks.

The task loss Ltask\mathcal{L}_{\text{task}} is the sum of the source language task loss LtaskS\mathcal{L}_{\text{task}}^{S} and target language one LtaskT\mathcal{L}_{\text{task}}^{T}, weighted by the hyper-parameter α\alpha, which is utilized to balance the training process

For classification, structured prediction, and span extraction tasks, the task loss is the cross-entropy loss (see details in Appendix A.2). For structured prediction tasks, it is non-trivial to implement the token-level label mapping across different languages. Thus we use the label probability distribution, predicted by the source language task model, as the pseudo-label for training, where tokens and labels are corresponding.

The consistency loss is composed of two parts: the representation consistency loss and the prediction consistency loss. The first loss is a regularization term and provides a way to align representations across different languages (Ruder et al., 2019). The second loss is to make better use of the supervision of downstream tasks. It only exists in the classification task, as the translation process does not change the label of this task, while in other tasks, it does.

3 Inference

Experiments

This section first introduces the cross-lingual understanding benchmark, XTREME. Then briefly introduces the configurations of downstream tasks and baselines. Finally, shows the main results of baselines and X-Mixup on XTREME.

Tasks In our experiments, we focus on three types of tasks in XTREME: (1) sentence pair classification task: XNLI (Conneau et al., 2018) and PAWS-X (Yang et al., 2019); (2) structured prediction task: POS (Nivre et al., 2018) and NER (Pan et al., 2017); (3) question answering task: XQuAD (Artetxe et al., 2020), MLQA (Lewis et al., 2020) and TyDiQA (Clark et al., 2020). The details of these datasets can refer to Hu et al. (2020). We utilize the translate-train and translate-test data from the XTREME repohttps://github.com/google-research/xtreme., which also provide the pseudo-label of translate-train data for classification tasks and question answering tasks. The rest translation data are from Google Translatehttps://translate.google.com/..

Models Experiments are based on two multilingual pre-trained models: mBERT and XLM-R. We use the pre-trained models of Huggingface TransformersWe use bert-base-multilingual-cased for mBERT and xlm-roberta-large for XLM-R. as the backbone model.

Hyper-parameters We select XNLI, POS, and MLQA as representative tasks to search for the best hyper-parameters. The final model is selected based on the averaged performance of all languages on the dev set. We perform grid search over the balance training parameter α\alpha and learning rate from [0.2, 0.4, 0.6, 0.8] and [3e-6, 5e-6, 2e-5, 3e-5]. We also search for the best manifold mixup layer from . In final results, we implement mixup in the first layer for classification tasks, 4th layer for structured prediction tasks. For QA tasks, we implement mixup in the 16th layer for large model, 8th layer for base model. Concrete details of experiments are presented in Appendix C.1.

2 Baselines

We conduct experiments on two strong multilingual pre-trained models to verify the generality of methods: (1) mBERT Multilingual BERT is a 12-layer transformer model pre-trained on the Wikipedias of 104 languages. (2) XLM-R XLM-R-large is a 24-layer transformer model pre-trained on 2.5T data extracted from Common Crawl covering 100 languages. Based on them, these are some strong baselines: (1) Trans-train Abbreviation for Translate-train. The training set of the source language is machine-translated to each target language and then the model is trained on the concatenation of all training sets. (2) Joint-Align Zhao et al. (2020) aligns the monolingual sub-spaces of the source and target language by minimizing the distances of embeddings for matched word pairs. (3) Filter Fang et al. (2020) splices the representation of the target sequence and its translation in intermediate layers to extract multilingual knowledge. (4) xTune Zheng et al. (2021) uses two types of consistency regularization based on four types of data augmentation.

3 Main Results

Results on the XTREME benchmark are shown in Table 3. Concrete results for each task are presented in Appendix C.2. Compared with strong baselines, X-Mixup shows its superiority across different backbones and tasks, which indicates its generality. X-Mixup outperforms Trans-train by 2.2% based on mBERT, and X-Mixup outperforms Filter by 1.5% based on XLM-R. The superiority of X-Mixup over Filter is that X-Mixup gives a calibrated representation for target languages, not just the concatenation of two representations. Besides, X-Mixup considers the noise of translation data and limits the noise propagation by introducing mixup ratio.

xTune achieves the best results on structured prediction tasks and the low-resource QA task TyDiQA (only 3.7k training data in English), but xTune uses three other data augmentation approaches in addition to machine translation. To make a fairer comparison, we conduct experiments under the same setting in Table 4, which indicates X-Mixup outperforms xTune on three types of tasks with only machine translation data augmentation. Besides, X-Mixup only needs one-stage training, while xTune implements a two-stage training algorithm. However, X-Mixup and xTune are complementary, where the former focuses on finding better representations for target languages while the latter concentrates on the cross-lingual data augmentation and consistency regularization.

Analysis and Discussion

To better understand X-Mixup and explore how X-Mixup influences the cross-lingual transfer performance, we conduct analysesIn this section, we utilize the XLM-R-large model as the backbone model. on several questions. Results show X-Mixup achieves performance improvements across different languages and it also reduces the cross-lingual representation discrepancy obviously. Table 6 in Appendix B verifies the effectiveness of X-Mixup on both seen and unseen languages. Besides, ablation results show the cross-lingual manifold mixup training contributes a lot to cross-lingual transfer.

(Q1) How X-Mixup influences the cross-lingual representation discrepancy? Language centroid (Rosenberg & Hirschberg, 2007) is the mean of the representations within each language. We plot the language centroid of different methods (see Figure 4 in Appendix B), which indicates X-Mixup brings closer language centroids significantly. We also calculate the CKA scores of the XNLI dataset (see Table 7 in Appendix B). Results show X-Mixup reduces the cross-lingual representation discrepancy evenly across different target languages, improving the CKA score by 10.4% on average. In conclusion, both the language centroids visualization and the CKA score improvement indicate X-Mixup reduces the cross-lingual representation discrepancy effectively.

(Q2) How X-Mixup influences the cross-lingual transfer gap? We compare the cross-lingual transfer gap in Appendix B Table 8. Compared with Trans-train, X-Mixup reduces the averaged gap by 39.8% and shows its superiority across three types of tasks. Compared with state-of-the-art methods, X-Mixup achieves the smallest cross-lingual transfer gap on four out of seven datasets, which suggests the effectiveness of X-Mixup on classification and QA tasks.

(Q3) What is the essential component of X-Mixup? There are five major components of X-Mixup: cross-lingual manifold mixup training, mixup inference, Mixup Ratio, Scheduled Sampling, and consistency loss. To better understand X-Mixup, we implement ablation studies in Table 5. Comparisons between X-Mixup and w/o mixup show the effectiveness of cross-lingual manifold mixup across different tasks, and even without mixup inference (translate-test data), the mixup training can also achieve 2.6% improvements on average. Besides, comparisons between X-Mixup and λ=λ0\lambda=\lambda_{0} show the effectiveness of introducing the translation quality modeling in the mixup process. Scheduled sampling achieves more performance improvements on the classification task, as the task shares labels across languages, and scheduled sampling can prevent the model from solely relying on the gold source sequence to make predictions. In addition, the consistency loss is also more effective on the classification task, because there is additional prediction consistency loss which can transfer the task capability from the source language to target languages. Detailed ablation results on the consistency loss are shown in Appendix B Table 9, which shows the KL consistency loss contributes more than the MSE consistency loss on the classification task.

(Q4) Which layer is the best to implement the manifold mixup? We implement the cross-lingual manifold mixup in different layers (see Figure 5 in Appendix B for details) and find different tasks prefer different mixup layers. Although different tasks have their own preferences, no matter which layer we mix, the cross-lingual transfer performance can be improved, except for mixing within a higher layer on classification tasks. The drop in classification task is mainly because the source and target sequences share the same task label. Performing mixup in a higher layer may make the model rely on the source sequence and ignore the target sequence. The structured prediction task is not sensitive to the mixup layer, mainly because this task relies on both the short and long dependence. For QA tasks, the cross-lingual transfer performance shows a trend from rise to decline as the mixup layer increases. The QA task needs higher-level understanding, but higher layers are more language-specific, where sequences from different languages have different gold answers.

Conclusion

This paper focuses on enhancing the cross-lingual transfer performance on understanding tasks. Considering the large cross-lingual transfer gap in recent works, this paper first analyses related factors and finds this gap is strongly associated with the cross-lingual representation discrepancy. Then X-Mixup is proposed to alleviate the discrepancy, which gives compromised representations for target languages by implementing the manifold mixup between the source and target languages. Empirical evaluations on XTREME verify the effectiveness of X-Mixup across different tasks and languages. Besides, both the visualization and quantitative analyses show X-Mixup reduces the cross-lingual representation discrepancy effectively. Furthermore, X-Mixup can also be applied to the multilingual pre-training process by implementing the cross-lingual manifold mixup on parallel data. Findings on the relationship between the cross-lingual transfer performance and representation discrepancy shed light on a promising way to boost cross-lingual transfer for future research.

References

Appendix A Method Details

In the multi-head attention layer, multiple attention heads are concatenated

and each head is the scaled dot-product attention

where WO\bm{W}^{O}, Wq\bm{W}^{q}, Wk\bm{W}^{k} and Wv\bm{W}^{v} are trainable parameters.

A.2 Training Objective

For classification tasks (e.g. NLI), the task loss is the cross-entropy loss

For structured prediction tasks (e.g. POS) and span extraction tasks (e.g. QA), the task loss is also the cross-entropy loss

Appendix B Analysis Results

Appendix C Experimental Details

For all tasks, we fine-tune on 8 Nvidia V100-32GB GPU cards with the batch size 64. For XQuAD and MLQA, we finetune 2 epochs. For other tasks, we finetune 4 epochs. There is no dev set in XQuAD, so we use the dev set of MLQA for the model selection. Table 10 shows hyper-parameters used for X-Mixup.

C.2 Detailed Results

Detailed results of each tasks and languages are shown below. Results of mBERT, XLM, MMTE and XLM-R are from XTREME (Hu et al., 2020). Results of Filter is the best results of Fang et al. (2020).