Towards Neural Machine Translation for African Languages
Jade Z. Abbott, Laura Martinus
Introduction
Given that South Africa’s education system is in crisis , strategies for sustainability within education must be explored. One suggested strategy to improve youth education would be to augment the learning process with online content .
The internet comprises of 53.5% English content, while the other 10 official South African languages comprise of less than 0.1% of the languages spoken on the internet . According to the 2011 South African census, only 9.8% of South Africans speak English as a primary language . Similar statistics exist for many other African countries .
For the South African youth to benefit from the massive amount of educational online content, the translation of online resources into the many low-resourced African languages is sorely needed. Unfortunately, machine translation of low-resourced languages has proven difficult with both conventional statistical machine translation (SMT) and the more recent neural machine translation (NMT) methods .
The convolutional sequence-to-sequence (ConvS2S) architecture improved translation results on multiple languages, including low-resourced languages . Additionally, by pre-training the Transformer architecture on many languages and then specialising to a single language, Gu et al. were able to improve performance on low-resourced languages .
This paper aims to serve as the initial work towards using modern neural machine translation (NMT) techniques to improve machine translation on African languages, and invigorate future research into using such techniques. NMT techniques are often overlooked in favour of conventional phrase-based translation systems. This is due to vanilla NMT’s reputation for comparatively under-performing on low-resource languages . This paper shows the performance of training convolutional sequence-to-sequence learning and the Transformer architecture on the Autshumato English-Setswana Parallel Corpora .
Related Work
Kato et al. used statistical phrase-based translation, based on Moses, in order to perform English-to-Setswana translation . They achieve a BLEU score of 32.71 on a dataset that is not publically-available and so was excluded from the comparison. Wilken et al. used a similar technique as , but focused on linguistically-motivated pre- and post-processing of the corpus in order to improve translation with phrase-based techniques . Wilken et al. was trained on the same Autshumato dataset used in this paper, and also used an additional monolongual dataset for language modelling.
Methodology
Section 3.1 describes the parallel dataset used, while the selected models and their hyperparameters for training are described in Section 3.2.
The publically-available Autshumato English-Setswana Parallel Corpora is an aligned corpus of South African governmental data which was created for the use in machine translation systems. The dataset consists of three smaller parallel corpora obtained from different sources which were combined to form a single corpus. The combined corpus was sharded into 111 300 sentences as training data, 44 700 as validation data, and 3 000 sentences set aside as test data for evaluation. This dataset is available for download at the South African Centre for Digital Language Resources website.Available online at: https://rma.nwu.ac.za/index.php/resource-catalogue/autshumato-english-setswana-multi-bilingual-corpus.html
2 Models
Limited work has been done using NMT techniques for African languages. In fact, as far as the authors can tell, this is the first work using modern NMT techniques for translation for South African official languages. We thus selected two recent NMT architectures, convolutional sequence-to-sequence and Transformer, to compare to existing research. The existing research uses phrase-based SMT for English-to-Setswana translation.
The Fairseq(-py) tookit was used to model the convolutional sequence to sequence model . The model used was one of Fairseq’s named architectures “fconv”. The learning rate was set to 0.25, a dropout of 0.2, and the maximum tokens for each mini-batch was set to 4000. The dataset was preprocessed using Fairseq’s preprocess script to build the vocabularies and to binarize the dataset. To decode the test data, beam search was used, with a beam width of 5.
The Tensor2Tensor implementation of Transformer was used . The model was trained for 125K steps . The learning rate was set to 0.4, with a batch size of 1024, and a learning rate warm-up of 45000 steps. The dataset was encoded using the Tensor2Tensor data generation algorithm which invertibly encodes a native string as a sequence of subtokens . Beam search was used to decode the test data, with a beam width of 4.
Training took less than 12 hours for both algorithms on a NVIDIA K80 GPU.
Results
Section 4.1 describes the quantitative performance of the models by comparing BLEU scores, while a brief qualitative analysis is performed in in Section 4.2.
According to the BLEU scores reported in Table 1, ConvS2S achieved a BLEU score of 27.77, which is 1.02 BLEU points below the performance of the phrase-based English-Setswana system from . The Transformer model significantly outperformed the phrase-based English-Setswana system, by 5.33 BLEU points. Despite the fact neither ConvS2S nor Transformer had information from additional language models or linguistic pre-processing, the architectures performed extremely competitively, with Transformer achieving a new state-of-the-art of 33.12 for English-to-Setswana translation.
2 Qualitative
Table 2 shows qualitative results for specific sentences from our test set. In order to understand the feasibility of using such models, a Setswana speaker translated the English-to-Setswana translations generated by our model back to English. Although not perfect, the translations capture much of the meaning from the original sentence. Impressively, the translations also use synonyms for other concepts: for example, Transformer translated "40% of people in a community are unemployed" to the Setswana equivalent of "40% of people are not working".
In conjunction with the quantitative results, these results confirm our hypothesis that the use of NMT systems, in particular the Transformer model, can improve the state of the art in English to Setswana translation.
Conclusion
Due to the rising need for African translations of online educational resources, the development of accurate machine translation systems for low-resourced languages has become an issue of importance.
We showed that state-of-the-art NMT architectures can significantly outperform existing SMT architectures for translation from English to Setswana with minimal hyperparameter optimization, and only a small amount of training time. This result suggests the promise of using the Transformer architecture to train models to translate other African languages. Future work includes training the Transformer architecture on multiple African languages at once, and then specialising on a specific language, as is done for Romanian by Gu et al .
The source code and the data used are available at https://github.com/LauraMartinus/ukuxhumana.
Acknowledgements
We would like to thank the organisers of the Deep Learning Indaba. Without the Indaba we would never have met, nor would we have had the resources and confidence to pursue and submit such research. Thank you to Guy Bosa for aiding us with our qualitative translations.