ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine Translation
Zhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao, Min Zhang
Introduction
Neural machine translation (NMT) has achieved success on the high-resource language pairs (Vaswani et al., 2017). However, it achieves an inadequate performance on the low-resource language pairs with merely bilingual data due to the data sparsity (Koehn and Knowles, 2017; Sennrich and Zhang, 2019). Transfer learning is a simple and powerful method that can be used to boost the model performance of low-resource NMT, which can transfer knowledge from an off-the-shelf high-resource parent model to the low-resource child model (Zoph et al., 2016).
The goal of transfer learning for low-resource NMT is to transfer knowledge sufficiently from the parent model. Prior works attempt to transfer more information from the embedding layer of the parent model, using methods such as building a joint dictionary (Kocmi and Bojar, 2018; Gheini and May, 2019), creating cross-lingual token mapping (Kim et al., 2019), and transferring the word of same form (Aji et al., 2020). These strategies achieve success in parameter transfer. The methods used in these works is summarized in Figure 1(a): the parent model transfers its parameters to the child model once and lacks continual guidance on the child training. This method can not utilize the parent model sufficiently, which potentially limits the learning of the child model.
This paper attempts to leverage a continual knowledge transfer from the parent model for the child training. The prediction distribution of the parent model is usually more informative since it has learned to predict various translations during training (Dreyer and Marcu, 2012; Khayrallah et al., 2020). Therefore, utilizing the prediction of the parent model to continuously guide the child training might be a nice plus.
To this end, we propose Consistency-based Transfer Learning (ConsistTL), a novel framework that can be used to model the consistency between the parent model and the child model during the child training. For each instance from the child data, ConsistTL constructs a semantically-equivalent pseudo source sentence for the parent by translating the child target sentence. Then ConsistTL encourages the child prediction to be consistent with the parent prediction since these two predictions are conditioned on the same target sentence and semantically-equivalent source sentences. ConsistTL is summarized in Figure 1(b): unlike in vanilla TL, the parent model releases continual guidance for the child model during training by receiving the transformation of the training data of the child model. With ConsistTL, the child model can acquire more knowledge from the parent predictions beyond the simple parameter transfer.
We conduct experiments on five low-resource NMT benchmarks. Experimental results show that ConsistTL achieves significant improvements over strong baselines of transfer learning. Combined with the representative data augmentation method of back-translation Sennrich et al. (2016a), our method can outperform the existing method up to 1.7 BLEU on the widely-used WMT17 Turkish-English benchmark. An extensive analysis reveals that our method, as an approach to model external consistency, is complementary to the method of modeling inner consistency.
We propose ConsistTL to model the consistency between the parent and child models in transfer learning for low-resource NMT. The cross-model consistency extends the knowledge transfer process to child training.
ConsistTL outperforms strong baselines significantly on five low-resource benchmarks, and is complementary to the other methods for low-resource NMT (e.g., back-translation and inner-model consistency learning).
We find that ConsistTL can improve the calibration of the child model by increasing translation accuracy and reducing overconfidence.
Related Work
Zoph et al. (2016) propose a parent-child framework to transfer the parameters from the parent model trained on a high-resource corpus to a low-resource child model with the same target language. The follow-up works aim to transfer the knowledge from the parent model parameters more sufficiently since Zoph et al. (2016) ignore the embedding layer mismatch caused by the vocabulary mismatch. These works include building the shared joint dictionary (Kocmi and Bojar, 2018; Gheini and May, 2019; Liu et al., 2019a, b), using an extra trained transformation to connect the embedding layers of two models (Kim et al., 2019; Sato et al., 2020; Liu et al., 2021a), or transferring partial word embedding from the parent model (Aji et al., 2020; Xu and Hong, 2022). In comparison, this presenting paper tries to transfer the knowledge contained in the parent prediction beyond parameters.
There are also some works that consider multilingualism to enhance transfer learning. Neubig and Hu (2018); Tan et al. (2019) propose to transfer knowledge from a multilingual parent model. Gu et al. (2018a) develop a universal architecture with shared lexical and sentence level representation to implement transfer learning. Gu et al. (2018b) propose a meta-learning approach to learn the initialization process using several high-resource language pairs. In this paper, we mainly focus on the bilingual parent model, since off-the-shelf bilingual models are easier to obtain.
Consistency Learning in NMT
Consistency learning aims to model consistency across different model predictions. The core idea is to craft the cross-view supervision and enhance cross-view consistency, which enriches the supervision signals. In the early studies, consistency learning methods focus on semi-supervised learning since consistency signals enable models to learn from data without manual labels (Xie et al., 2020; Miyato et al., 2019; Lowell et al., 2021a). The semi-supervised consistency learning also enhances NMT models with the combination of forward-translation (Clark et al., 2018). Recent studies of consistency learning gradually consider supervised learning areas (Chen et al., 2021; Lowell et al., 2021b). R-Drop is one of the simple and generalized methods (Liang et al., 2021), which enhances the consistency between multiple model structures produced by dropout, and improves NMT performance.
There are some consistency learning methods specially designed for supervised NMT based on data augmentation. Shen et al. (2020); Guo et al. (2022); Gao et al. (2022) leverage token-level perturbations to craft different predictions. Kambhatla et al. (2022) explore the insight of cryptology to create a ciphertext with the same semantic meaning for each training instance. Xie et al. (2021) develop a scheme to augment target side input by the output prediction from the first forward. Wang et al. (2022) enhance the Chinese representation by modeling the consistency between stroke representation and its corresponding encryption. These consistency modeling works in NMT mainly exploit the inner model consistency, i.e., modeling consistency within the same model. Compared with these works, we exploit transferable information from the cross-model consistency in transfer learning for low-resource NMT. Chen et al. (2017) also exploit the consistency between two models and the differences from us are: 1) we construct the parent data instead of the child data; 2) we enhance the transfer learning framework for low-resource NMT while their method relies on a pivot language to improve zero-resource NMT.
ConsistTL
Previous works concentrate on knowledge transfer from the parent model parameters (Kocmi and Bojar, 2018; Kim et al., 2019; Aji et al., 2020). This paper argues that continual knowledge transfer from the predictions of the parent model is also beneficial to transfer learning. Since the parent model is trained on a large number of diverse translations (Dreyer and Marcu, 2012; Khayrallah et al., 2020), the prediction of the parent model might be informative for child training.
In this paper, we propose ConsistTL, a framework that can be used to jointly transfer the knowledge from the parent parameters and the prediction distribution of the parent model. ConsistTL can enhance the consistency between the distribution from the parent model (e.g., German-English) and the child model (e.g., Turkish-English) conditioned on the semantically-equivalent sentences. ConsistTL is composed of two parts: semantically-equivalent parent data construction and parent-child cross-model consistency learning. The overall framework is illustrated in Figure 2.
2 Semantically-Equivalent Parent Data Construction
To implement the proposed ConsistTL, we need to model the cross-model consistency between the parent-side source sentence and child-side source sentence. As the parent model and child model are trained for different language pairs, for each instance in the child data, we need a semantically-equivalent sentence written in the source language of the parent language pair.
3 Parent-Child Cross-Model Consistency Learning
To make the distribution of the parent model and child model consistent, we need to measure and minimize the gap between these distributions. The proposed framework minimizes the distribution gap with the following cross-model loss:
where denotes the logits output before softmax is computed, and represents the size of the target-side vocabulary. We can smooth the distribution of the acquired prediction with .
Final Child Training Objective Function
We interpolate the weighted cross-model consistency loss and negative log-likelihood loss to create the final child training objective function. The final child training objective is formulated as:
In this manner, the child model can learn each instance under the guidance of the parent model.
Experiments
This paper follows Aji et al. (2020) to transfer knowledge from the German-English (De-En) task. Our parent model is trained on the WMT17 De-En training set and validated on newstest2013. The training set consists of 5.8M sentences. The vocabulary is implemented using joint source-target BPE with 40k merge operations (Sennrich et al., 2016b).
Child Language Pairs
We conduct transfer learning experiments on five low-resource translation benchmarks. There are four translation benchmarks from Global Voices (Tiedemann, 2012; Khayrallah et al., 2020): Polish (Pl), Hungarian (Hu), Indonesian (Id), and Catalan (Ca). The sentences of corresponding languages are paired with English sentences. The subset splits follow Khayrallah et al. (2020). For the fifth language pair, we adopt WMT17 Turkish (Tr)-English benchmark and use newstest2016 as the validation set. Before subword segmentation, we apply normalization and tokenization to all datasets using the Moses (Koehn et al., 2007). The sentences with more than 60 words and a length ratio of over 1.5 in training data would be filtered for the Tr-En language pair. After data filtering, we then apply joint source-target BPE to the child language pairs with fewer merge operations than were used for the parent language pair. We display detailed statistics of the preprocessed datasets in Table 1.
2 Model Configuration
We use the Transformer-base architecture (Vaswani et al., 2017) to implement NMT models in our experiments from the toolkit Fairseq (Ott et al., 2019). For all the low-resource translation models, we tie the input embedding layers of the decoder and the output projections (Press and Wolf, 2017; Inan et al., 2017). For the high-resource parent model, we tie all embedding layers.
We implement the transformer models training from scratch as baselines. These models share the vocabularies of source and target languages.
TL
Zoph et al. (2016) propose a transfer learning method to randomly assign a parent word embedding to each child word. The rest of the parameters are copied from the parent model. We implement this method as a baseline in our experiment. This vanilla transfer learning method is marked as “TL”.
TM-TL
Aji et al. (2020) propose a method named “token matching” (TM), which is simple and effective. We implement this transfer learning method as a baseline in our experiments marked as “TM-TL”. Some tokens are common to both parent model and child model vocabularies for child language. TM-TL assigns the embedding parameters of such common tokens in the source embedding layer of the child model. For the other parameters of the source embedding layer, TM-TL initializes them randomly following the vanilla NMT. The rest settings of TM-TL follow TL.
ConsistTL
Since our method , ConsistTL , transfers knowledge from the prediction distribution of the parent model, it is orthogonal to the parameter transfer methods. We adopt TM to implement ConsistTL.
3 Settings
We train all the NMT models using the Adam optimizer (Kingma and Ba, 2015) with . We also use the inverse square root schedule to control the variation of the learning rate during training. For the parent model, the value of linear warmup steps is set to and the peak learning rate is set to . The parent model is trained for 200 epochs with 460K () tokens per batch and a low dropout rate of 0.1. We use 4,000/2,000 max tokens per batch for Tr/Pl-En tasks, and 1,000 for the other tasks. For the the training of the low-resource translation model from scratch, we set the number of linear warm-up steps to and the peak learning rate is set to . For the low-resource translation models with transfer learning, as the inner layers are adequately trained on the parent language pair, we set narrow warm-up steps to . We set a lower peak learning rate to to prevent over-fitting. All the low-resource translation models are trained for 200 epochs. To prevent the models from over-fitting on low-resource language pairs, we set the dropout rate to 0.3, both the attention dropout rate and activation dropout rate to 0.1. For checkpoint selection, the checkpoints with the best validation BLEU would be selected as the final model checkpoints. We train an En-De model on the same parent data as the reversed parent model. We set to 1.0 and tune on .
Evaluation
We used beam search with a beam width of 5 and a length penalty of 1 to evaluate all the NMT models. And we also use such decoding setting to generate pseudo parent source sentences. The translation quality is evaluated using SacreBLEU (Post, 2018)Signature: nrefs:1 + case:mixed + eff:no + tok:13a + smooth:exp + version:2.0.0 and the BERTScore (Zhang et al., 2020)https://github.com/Tiiiger/bert_score, which evaluate the surface-form correctness and semantic correctness, respectively.
4 Main Results
Table 2 displays the results of the five aforementioned translation tasks in terms of the BLEU score and the BERTScore. We report the existing results from Khayrallah et al. (2020) and Baziotis et al. (2020). ConsistTL outperforms the existing works on the same test sets in terms of the BLEU score.
The data scale of the four language pairs (Id-En, Ca-En, Hu-En, Pl-En) collected from Global Voices is extremely lower than that of WMT17 Tr-En. The parameter transfer methods TL and TM-TL are strong since they outperform vanilla NMT in terms of both BLEU and BERTScore. Our method ConsistTL achieves a significant average improvement of 1.1 BLEU and 2.0 BERTScore over the TM-TL.
WMT17 Tr-En
Different from the four aforementioned language pairs, the data scale of Tr-En is larger. Compared with the aforementioned extremely low-resource language pairs, the transfer learning baseline methods have lower performance gains (TM-TL) or even no performance gains (TL) over Vanilla NMT on this benchmark with larger data. Although baseline methods are less informative on the Tr-En, our method still achieves a significant improvement over the strong baseline TM-TL with 0.7 BLEU and 2.0 BERTScore. The results of BERTScore on five translation benchmarks indicate that ConsistTL provides consistent and significant improvements of the semantic correctness. Our follow-up analysis is based on the WMT17 Tr-En benchmark.
Analysis
There are many ways to measure the difference between two distributions. Different measure functions would result in different consistency loss types. This part compares two typical choices Jensen–Shannon divergence (Lin, 1991) and Kullback–Leibler divergence (Kullback and Leibler, 1951) to study the effect of the chosen loss type on ConsistTL. To make the most of their capabilities, we tune the hyperparameters for each loss type. The final hyperparameters of the JS loss and the KL loss are and respectively. As seen in Table 3, the model with JS achieves the best improvements, while the model with KL loss realizes limited improvements. Perhaps minimizing the KL divergence is more difficult than minimizing the JS divergence since we observe that lowering the KL loss through smoothing () is helpful.
2 Comparison of Learning Curves
The learning curve is one way to describe the temporal effect of the proposed method during training , which has also been adopted in previous consistency modeling works (Liang et al., 2021; Kambhatla et al., 2022; Xie et al., 2021). This part compares the proposed method , ConsistTL , with the baseline methods TL and TM-TL in terms of their learning curves described by validation set BLEU. Figure 3 displays the learning curves of these three transfer learning methods. During the child training, the ConsistTL and TM-TL grow into larger validation BLEU values with higher convergence speeds. The difference between ConsistTL and TM-TL is invisible in the early learning phase and gradually changes into a stable and significant value. Finally, ConsistTL achieves the largest upper bound among these three learning curves, which indicates that ConsistTL is informative to child training.
3 Effect of Pseudo Source Generation
Although beam search is a de facto decoding method that is used to generate sentences in the NMT task, there are also other decoding methods that can be used to generate synthetic data (Ott et al., 2018; Edunov et al., 2018), including greedy search, and greedy sampling. To study the effect of decoding methods in parent-side synthetic data generation, we compare the results produced by beam search, greedy search and greedy sampling. As shown in Table 4, we find that greedy sampling produces lower improvement in contrast to the beam search. This can be interpreted to mean that greedy sampling introduces low-quality translations due to the randomness during decoding (Ippolito et al., 2019). Though greedy search performs worse than beam search in usual model testing, the final results in Table 4 empirically shows that there is no significant difference in final results between beam search and greedy search. That means the translation quality of greedy search meets the ConsistTL requirement. Therefore, we can adopt the greedy search to generate pseudo parent source sentences for future usage.
4 Effect of Data Scale
To verify the robustness of the improvements brought by ConsistTL across data scales, we compare our method and the baselines methods, TM-TL and TL, on the subsets with different scales. The training subsets are randomly sampled from the full Tr-En training data according to the specific ratios. Figure 4 demonstrates the result of data ablation. The performances of ConsistTL and TM-TL are more stable than TL on different data scales, since TL drops more performance on smaller-scale data. ConsistTL achieves significant improvements consistently in contrast to TM-TL across different data scales. This result indicates the robust effectiveness of ConsistTL with data scales varying.
5 Effect of Back-Translation
Back-translation (Sennrich et al., 2016a; Edunov et al., 2018) is an essential method for NMT that is used to augment training corpus by generating synthetic sentences from target-side monolingual data. In this part, to study the complementarity between this work and back-translation, we compare the proposed method and the baseline TM-TL on the training data augmented by back-translation. We sample 200k English monolingual data from News Crawl 2015 for back-translation, following the approximate 1:1 ratio according to Aji et al. (2020). In Aji et al. (2020), the parent model for Tr-En is also a Transformer-base model trained on WMT17 De-En data. Table 5 displays the results reported from Aji et al. (2020) and the results of our implementation. Our proposed method ConsistTL outperforms the existing baselines with 1.7 BLEU. In combination with back-translation, our method yields a performance gain with the significance of 95% () on the test set, which means ConsistTL is complementary to back-translation. This result confirms the generality of ConsistTL since ConsistTL can be generalized to a scenario of in which training data are augmented by back-translation.
6 Combination with Inner Consistency
As ConsistTL learns the cross-model consistency, here we investigate the effectiveness of such cross-model consistency under inner consistency modeling. R-Drop is a typical method to model inner consistency (Liang et al., 2021). We compare the transfer learning methods combined with inner consistency implemented by R-Drop. Table 6 shows that our method achieves the best performance and consistent improvements across the validation set and test set (). This result validates that the cross-model consistency is indeed effective since ConsistTL still improves performance when combined with R-Drop.
7 Model Calibration
As this paper proposes leverag ing the prediction distribution of parent model, the following questions arises: how does the parent prediction help shape the child prediction? The inference calibration can be used to describe the prediction distribution for NMT models, which requires mitigating the gap between the prediction confidence and translation accuracy during inference (Wang et al., 2020). We exploit the gap between the averaged confidence and translation accuracy to evaluate inference calibrationhttps://github.com/shuo-git/InfECE. The average accuracy illustrates the correctness of generated tokens, which is annotated by the translation error rate (TER) (Snover et al., 2006). The averaged confidence is computed on the estimated probabilities of generated tokens. More details can refer to Wang et al. (2020). Figure 5 shows the averaged accuracy and the averaged confidence of child models implemented by ConsistTL, TM-TL, and TL. The visualization results show that the averaged confidences are higher than the corresponding averaged accuracy, namely overconfidence of the models. Compared with other methods, ConsistTL improves the translation accuracy and reduces overconfidence. In this way, ConsistTL narrows the gap between the confidence and accuracy, which indicates that ConsistTL can be beneficial to child model inference calibration by introducing the prediction distribution of the parent model. This provides a reasonable explanation of why ConsistTL can give a performance boost to the child model.
Conclusion and Future Work
We introduce ConsistTL for transfer learning in low-resource NMT, which can continuously transfer the knowledge from parent prediction during child training. In order to transfer knowledge from parent prediction, we propose to model cross-model consistency between parent model and child model with the child source sentences and pseudo parent source sentences. Experimental results on five translation benchmarks verify the effectiveness of ConsistTL, which can significantly improve the translation performance. ConsistTL is complementary to other powerful methods for low-resource NMT (e.g., back-translation). Further analysis reveals that introducing parent prediction is helpful to shape child model prediction distribution, resulting in better inference calibration.
In the future, we would like to apply curriculum learning (Liu et al., 2020; Zhan et al., 2021) to better organize the learning of the child model. It is also worthwhile to enhance the parent and child models by utilizing pre-trained knowledge learned from unlabeled data Liu et al. (2021b, c).
Limitations
Our proposed framework would occupy extra computation resources compared with baseline methods. In the part of semantically-equivalent parent data construction, we need to train an additional reversed parent model to back-translate the target sentences of child data. In the part of parent-child cross-model consistency learning, at each training step of child model, the parent model needs an additional forward pass to generate prediction distribution for guidance.
Acknowledgments
This work was supported in part by the National Natural Science Foundation of China (Grant No. 62206076), the Science and Technology Development Fund, Macau SAR (Grant No. 0101/2019/A2), Shenzhen College Stability Support Plan (Grant No. GXWD20220811173340003 and GXWD20220817123150002) and the Multi-year Research Grant from the University of Macau (Grant No. MYRG2020-00054-FST). We would like to thank the anonymous reviewers and meta-reviewer for their insightful suggestions.