Towards a parallel corpus of Portuguese and the Bantu language Emakhuwa of Mozambique
Felermino D. M. A. Ali, Andrew Caines, Jaimito L. A. Malavi
Introduction
Machine translation (MT) is the process by which computational models are trained to transform a source language text into a target language text, and is a technology which in recent years has seen great improvements in performance thanks to the development of MT with neural networks (NMT) Kalchbrenner and Blunsom (2013); Sutskever et al. (2014); Bahdanau et al. (2015). NMT typically depends on supervised machine learning from large parallel corpora of texts in the source and target language pair. The creation of such corpora usually depends on extrinsic motivation for the existence of abundant parallel text: for instance, government activity in multilingual settings such as the European and Canadian Parliaments. In the context of NMT, this fact has led to a dichotomous situation of high and low-resource language pairings. For the academic community, English-German or English-French are prototypical high-resource translation pairs whereas most other language pairs are low-resource in comparison.
We present a corpus for a low-resource language pair: Portuguese and Emakhuwa (alternatively, ‘Makhuwa’ or ‘Makua’), the official and the most widely spoken languages of Mozambique, respectively. Emakhuwa is spoken in all the provinces of northern Mozambique, namely Niassa, Cabo Delgado and Nampula and also in Zambezia in central northern Mozambique. It is estimated that approximately 25% of the country’s population of 30 million people make use of the language on a daily basis as an alternative to Portuguese de Paula and Duarte (2016). Emakhuwa is also spoken in some of the neighbouring countries to the north of Mozambique – namely, Tanzania and Malawi Ngunga (2012) but speaker populations in these countries are relatively small.
There are 8 variants of Emakhuwa Ngunga and Faquir (2012): Elomwe (ISO-639 code: ngl), Esankaci, Esaaka (ISO-639 code: xsq), Echirima (ISO-639 code: mhm), Emarevoni (ISO-639 code: xmc), Emeeto (ISO-639 code: mgh), Enahara and Central (ISO-639 code: vmw). The variants are distributed across different districts of northern and central Mozambique (see Figure 1) and differ slightly in terms of accents and lexicon.
Like many languages spoken on the African continent, Emakhuwa has limited resources for computational linguistics compared with English, French, Portuguese, and so on. Thus we are collecting a dataset for MT in which Portuguese texts are paired with Emakhuwa translations. The 47,415 sentence pairs we have collected contain 699,976 word tokens of Emakhuwa and 877,595 word tokens in Portuguese. The corpus will be made available for research use when final normalization and data collection has been completedplaceholder.com. In addition we will seek out new data sources and continue to expand the corpus.
Related Work
In general, African languages have been relatively little explored by the NLP research community. However, lately this scenario seems to be improving as more data are being made available for research use, including translation work. OPUS Tiedemann (2012), the 1000Langs corpus crawled from Bible corpora Asgari and Schütze (2017), and the JW300 project Agić and Vulić (2019) are examples of this trend, as they made available parallel corpora of over 300 languages including Emakhuwa-Portuguese. The Crúbadán Projecthttp://crubadan.org/languages/vmw also used Jehovah’s Witness resources to create a corpus of Emakhuwa. However, the corpus is monolingual and relatively small with around 44,071 tokens extracted from 4 documents.
The Masakhane project works mainly on low-resource MT for African languages, and is an excellent source of African language parallel corpora Nekoto et al. (2020). Currently, they have a growing online community of researchers who collaborate, contribute and share advancements in NLP for African languages. They have an open repository https://github.com/masakhane-io/masakhane-mt for sharing data as well as code and tools for building and facilitating NLP research in African languages. Nevertheless, the Emakhuwa language has not been explored yet by the community.
Data
The dataset contains Portuguese-Emakhuwa parallel sentences, collected from crawling the Jehovah’s Witness (JW) websitehttps://www.jw.org and the African Story Book websitehttps://www.africanstorybook.org. The African Story Book contains 17 short stories in Emakhuwa with Portuguese translations. The Jehovah’s Witness website contains The Watchtower—Study Edition magazine as well as The Watchtower and Awaque! from 2016 to 2021. Also, contains 27 books of the New Testament Bible.
The corpus also contains text from various institutional sources. These were digitized by performing optical character recognition (OCR) on a set of PDF files from Centro Catequético Paulo VI (Paul VI Catechetical Centre), the Universal Declaration of Human Rights, and Mozambican Land Law. The majority of texts come from the Centro Catequético Paulo VI which has published a series of three catechism books: Yesu Mopoli Ahu – Jesus Nosso Salvador (‘Jesus Our Saviour’). These books, besides being religious, also touch upon a wide range of topics.
The dataset contains texts written in the central Emakhuwa variant only, because this variant has relatively more resources than others. The sentences come from 125 pair documents in total, and amount to 699,976 word tokens of Emakhuwa. Table 1 shows examples from the dataset, with Portuguese and Emakhuwa parallel sentences and English glosses.
Processing of the OCR obtained from these documents involved the removal of some inconsistencies. For instance, the letter “ì” was wrongly recognized as “T”, and the letter “á” was recognized as “6”, to give just two examples of OCR errors. Processing also involved dividing longer sentences into shorter and concise sentences as well as discarding sentences wrongly recognized. Then followed the alignment of Portuguese and Emakhuwa sentences which has been manually carried out. More processing needs to be done, as it transpires that spelling standards vary between the data sources. As an example, there are some lexical conflicts in the corpus, particularly regarding the use of the letters “j” and “c” along with the suffix or infix “erya”. JW texts tend to use the letter “cerya”, while others use “jerya”. So for example the English word “begin” would be written as “opacerya” in JW texts, whereas in others resources it would be written as “opajerya”. Another example is the English word “kiss”: its translation to Emakhuwa in the African Story Book is “epeeco”, whereas in other resources is translated as “epexo”.
Word borrowing is another issue to take into account, since Emakhuwa takes many words from Portuguese, but without a consensus on spelling conventions. As an example, take the word “bible” which in Portuguese is ”bíblia”. In Emakhuwa it is sometimes written as “biblia” and at other times written as “biibiliya”. This inconsistency also exists within Emakhuwa language resources itself, as the alphabet and spelling standards went through several revisions – the latest in 2012 Ngunga and Faquir (2012). Therefore, some future processing is necessary to normalize the word standards in the dataset. The corpus will be available for research use, and updates may be obtained from our project websiteplaceholder.com. Table 2 summarizes the corpus in terms of sentences and the number of tokens coming from each data source.
Baseline translation model
We train an initial NMT model with the default OpenNMT configurationhttps://github.com/OpenNMT/OpenNMT-py/blob/master/docs/source/quickstart.md, which consists of a 2-layer LSTM with 500 hidden units on both the encoder and decoder Klein et al. (2017). Sentences from the corpus are randomly assigned to training and test sets at a ratio of 9:1. These data splits will be released with the corpus so that others may compare the performance of their models against this one. We evaluate with BLEU Papineni et al. (2002) as shown in Table 3. This presents a baseline level of performance to improve upon in future work.
Conclusion
In this paper we describe the creation of a new parallel corpus of Emakhuwa and Portuguese. The preparation of the corpus is on-going as the texts require more processing and normalization, but it will be made freely available for research use, most probably for machine translation. Currently, the dataset is made up of mostly religious and legal content but in future our objective is to diversify the range of sources and topics covered. This is important in order that Portuguese-Emakhuwa translation models work well across various domains.
Acknowledgements
The authors wish to acknowledge the Centro Catequético Paulo VI for allowing to use their materials for research purpose. The second author is supported by Research England via the University of Cambridge Global Challenges Research Fund. He wishes to thank Paula Buttery, Tanya Hall, Carol Nightingale & Sara Serradas Duarte for their help and guidance.