Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation

Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Hassan Muhammad, Ibrahim Sa'id Ahmad, Subhadarshi Panda, Ondřej Bojar, Bashir Shehu Galadanci, Bello Shehu Bello

Introduction

Machine translation is the use of a computer to automatically generate the equivalent of a given source text in a language that is different from the original language. While Neural Machine Translation (NMT) [Bahdanau et al. (2015, Vaswani et al. (2017, Gehring et al. (2017] has revolutionized automatic translation, the absence of sufficient training data in many languages has limited the benefits of such systems to a few languages rich in resources, although at least some treatment is possible even for low-resource languages [Sennrich and Zhang (2019].

Multi-modal Machine Translation (MMT) enables the use of visual information to improve the translation quality, supplementing the missing context and providing cues to the machine translation system for better disambiguation. Despite the increasing popularity of multi-modal techniques, sufficiently large and clean datasets are scarce to fully benefit from the potential. For languages with such data, various approaches have been proposed, demonstrating their usability in improving translation quality, e.g., see [Krishna et al. (2017, Lin et al. (2020, Long et al. (2021, Liu et al. (2021].

The images in Figure 1 present some examples where the absence of context allows to consider two different translations, where each is correct in a different setting. In the first one, the English word “court” is translated as kotu, which is the Hausa word for a legal court. But the image illustrates that the men were standing on a [tennis] field. The absence of the word “tennis” misled the standard machine translation system and even many human translators into thinking that the former translation is required. The second example mentions a “story” of a two-story house. The MT system translated the description as labarin, meaning story (narrative), instead of the correct bene (house story/storey). Without the picture, even human translators may make the same error given the very short and not quite correct English source.

Hausa is a Chadic language and a member of the Afro-Asiatic language family. Hausa is the most-spoken language in this family, with an estimate of about 100 to 150 million first-language and second-language speakers. https://www.herald.ng/full-list-hausa/ The majority of these speakers are concentrated in the Northern part of Nigeria in cities such as Kano, Daura, Sokoto, Zaria, etc., and the Southern Niger Republic. The language is written in Arabic or Latin characters. The Arabic script is known as the Ajami and was mostly used in the pre-colonial era, dating back to the 17th century [Jaggar (2006]. The language is nowadays written in the Latin script known as boko.

Despite a large number of speakers and many written books, e.g., ?), ?), ?), Hausa is considered a low resource language in NLP. This is due to the absence of enough publicly available resources to implement most of the tasks in NLP. While some datasets exist, they are either scarce, machine-generated, or in the religious domain. This limits diversity, restricting the usage of trained models to very few domains. For tasks such as multi-modal translation, and image-to-text translation (image captioning), among others, there exist no training or evaluation data. For translation in the news domain, only an evaluation dataset exist[Goyal et al. (2021]. Therefore, there is a need to create training and evaluation datasets for building machine learning models to help reduce the research gap between the low-resourced Hausa language and other languages.

This work, therefore, presents the Hausa Visual Genome (HaVG), a dataset that contains the description of an image or a section within the image in English and its equivalent in Hausa. The dataset was prepared by automatically translating the English description of the images in the Hindi Visual Genome (HVG) [Parida et al. (2019]. The data is made of 32,923 images and their descriptions that are divided into training, development, test, and challenge test set. The machine-generated Hausa descriptions were then carefully post-edited taking into account the corresponding images. The HaVG is the first dataset of its kind in Hausa and can be used for Hausa-English machine translation, multi-modal research, and image description, among various other natural language processing and generation tasks.

To describe the process of building the multimodal dataset for the Hausa language suitable for English-to-Hausa machine translation, image captioning, and multimodal research.

To demonstrate some sample use cases of the newly created multimodal dataset: HaVG.

The rest of the paper is arranged as follows: Section 2 presents the available datasets for NLP in the Hausa language. Section 3 presents the processes of data collection and labeling. In Section 4, we present some experiments and results on the application of the HaVG data. Finally, we conclude the work and provide directions for the future in Section 5.

Related Work

While the Hausa language does not have any dataset for multimodal tasks, a few others have been created for other NLP tasks. ?) produced sentiment annotations of tweets and used them in their work. ?) provided a pseudo-parallel corpus for machine translation. The Tanzil dataset https://opus.nlpl.eu/Tanzil.php [Tiedemann (2012], a translation of the Quran in many languages including Hausa, and the JW300 https://opus.nlpl.eu/JW300.php [Agić and Vulić (2019] are available for machine translation tasks. All of these data, though, are either not natural or strictly in the religious domain, limiting the accuracy or general applicability of the translation models trained on them. Apart from the FLoRes evaluation dataset [Goyal et al. (2021] for machine translation tasks, which is not reflective of the domain of available training data, there exists no standard benchmark evaluation (test) sets that truly indicate the performance of natural language processing models to the best of our knowledge.

Resources for other NLP tasks in the language are also scarce. ?) provided two sets of word embeddings in Hausa for NLP. ?) and ?) built a collection of transcribed speech resources for automatic speech recognition (ASR) and similar tasks in the language. ?) trained a part-of-speech tagger for Hausa.

Initiatives such as Masakhane https://www.masakhane.io/ and HausaNLP https://www.hausanlp.org/ have started creating these data for Hausa and other African languages, most of which are considered low-resource and these will help in future NLP research and application in such languages.

Training and Evaluation Data

The HaVG training and evaluation (development test and challenge test) data were produced by automatically translating the Hindi Visual Genome (HVG) and revising it as described below.

The HVG training, evaluation, and test dataset consist of randomly selected images and their descriptions from the Visual Genome (VG) corpus [Krishna et al. (2017]. The HVG challenge test set was specifically sampled so that each sentence contains an English word that is lexically ambiguous when translated into Hindi. While the VG data contains multiple captions in English, with each caption representing a particular region in an image, the HVG data contains only a single random caption of a section in each image.

2. Annotation

To prepare the HaVG data, therefore, we implemented the following steps:

We use Google Translate https://translate.google.com/ to translate all the available 32,923 HVG English captions into Hausa.

We developed a web-based annotation tool https://github.com/abumafrim/visual-genome-dataset-creation-tool and hosted it locally to help with the post-editing of these translations. The web interface enables the annotator to edit the generated translations by showing them the image and the original caption side-by-side. See the illustration in Figure 2.

We gave the machine translations of the captions to Hausa volunteers for post-editing. The translations of many of the unambiguous sentences were mostly found to be correct.

For a secondary check, we sampled 3,500 of the post-edited captions (representing about 10% of the whole dataset) for manual verification. It was found that a small number of the sentences were found unedited even though there were obvious translation errors. The errors in these sentences were corrected by the verifiers.

Some statistics in the annotated HaVG dataset are provided in Table 1. We used the NLTK punkt tokenizer [Bird et al. (2009] to estimate the statistics. The Hausa sentences of the HaVG were found to have 36 and 1 word in the longest and shortest sentences, respectively. The average sentence length ranges from 4.99 to 6.80 words per sentence, with the challenge test set statistically having longer sentences. The training set has a low type-token ratio (TTR) – a measure of vocabulary variation or lexical richness of a text – of 0.05. This is reflective of the restricted domain of the data as most of the sentences are in the sports domain, mainly tennis.

Sample Applications of HaVG

We used the Transformer model [Vaswani et al. (2018] as implemented in OpenNMT-py [Klein et al. (2017]. http://opennmt.net/OpenNMT-py/quickstart.html Subword units were constructed using the word pieces algorithm [Johnson et al. (2017]. Tokenization is handled automatically as part of the pre-processing pipeline of word pieces.

We generated a vocabulary of 32k subword types jointly for both the source and target languages, sharing it between the encoder and decoder. We used the Transformer base model [Vaswani et al. (2018]. We trained the model on a single GPU and followed the standard ‘‘Noam’’ learning rate decay, https://nvidia.github.io/OpenSeq2Seq/html/api-docs/optimizers.html see ?) or ?) for more details. Our starting learning rate was 0.2 and we used 8000 warm-up steps. The text-to-text translation results for the development (D-Test), dev (D-Test), test (E-Test), and challenge test (C-Test) are shown in Table 2.

In Table 3, we present some examples where the text-only translation system was able to generate correct translations, although not the exact wording of the reference translations. The system translated “stand” as “tsayuwa” whereas the most appropriate translation should have been “mazauni” (with mazauni meaning a place where something is kept while tsayuwa means something/someone is in a standing position). The system also translated “block stone” as “dutse (stone)”, omitting “block”.

2. Multimodal Translation

Multimodal translation involves utilizing the image modality in addition to the English text for translation to Hausa. We take the multimodal neural machine translation approach using object tags derived from the image [Parida et al. (2021a]. We first extract the list of (English) object tags for a given image using the pre-trained Faster R-CNN [Ren et al. (2015] with ResNet101-C4 [He et al. (2016] backbone. We pick the top 10 object tags based on their confidence scores. In cases where less than 10 object tags are detected, we consider all tags.

Next, the object tags are concatenated to the English sentence which needs to be translated to Hausa. The concatenation is done using the special token ‘##’ as the separator. The separator is followed by comma-separated object tags. Adding object labels enables the otherwise text-based model to utilize visual concepts which may not be readily available in the original sentence. The English sentences along with the object tags are fed to the encoder of a text-to-text Transformer model. The decoder generates the Hausa translations auto-regressively. We generated a vocabulary of 50k subwords for both source and target languages. Then we trained the Transformer base model using the “Noam” learning rate decay. We used an initial learning rate of 2, dropout of 0.1, and 8000 warm-up steps. The results of the multimodal translation are shown in Table 2.

The automatic evaluation indicates that the text-only translation performs better on both the evaluation and challenge test sets when compared to the multimodal translation. However, upon manual inspection of the outputs, we observed instances where the multimodal system was able to resolve ambiguity and generate a more appropriate translation of the given source sentence, see Table 3 for some examples. The performance is strikingly lower on the challenge test set compared to the evaluation set in both setups. We performed a manual evaluation on a sample of this data to investigate the reason for this low performance.

About 10% of the translations of the challenge test set by both the text-only and multimodal systems were sampled and manually evaluated to assess the quality of the generated sentences. We categorized these sentences as either correct, partially correct, or incorrect. We also checked instances where the multimodal system is not only correct (or partially correct) but was also able to resolve ambiguity. Lastly, we checked whether the sentences generated by the multimodal system are reasonable or not, i.e. whether they generally capture the original meaning. The results of this evaluation are provided in Table 4.

While the multimodal system was found to be half as accurate compared to the text-only model, it was able to resolve ambiguity in about 10% of the sampled data. Finally, we observe that the annotation for “reasonable” translations (i.e. whether the meaning is “generally captured” is apparently much more permissive that the annotation for correctness: a substantial amount of the generated text (74 items, i.e. 53%) was found to be reasonable even though only about 37% of the sentences are either correct or partially correct translations of the source sentences. This detailed analysis nevertheless confirmed that the multi-modal system produces overall worse translations, perhaps confused by the automatic object captions.

3. Image Caption Generation

To generate the Hausa captions, we followed ?) who proposed a region-specific image captioning method through the fusion of the encoded features of the region and the complete image. The model consists of three modules – an encoder, fusion, and decoder – as shown in Figure 3.

In the proposed approach, the features of the entire image, as well as features of the sub-region, are considered to train the model. The features from the corresponding regions are extracted through Region of Interest (RoI) pooling [Girshick (2015]. Specifically, the feature vector is the output of the fourth block of ResNet-50 in our experiments. It is a 2048-dimensional vector for both the image and the sub-region. We keep the image encoder module non-trainable. In other words, it is used as a feature extractor.

While the region-level features capture details of the region (objects) to be described, the image-level features provide an overall context. To generate a meaningful caption, both need to be fused appropriately. We obtain the final feature vector by simple concatenation of features from the region and features from the entire image. The concatenation resulted in a 4096-dimensional vector.

The concatenated feature vector is passed through a linear layer to project it into a 128-dimensional vector which is then fed as input to an LSTM decoder as the first time step. The decoder generates the tokens of the caption autoregressively using a greedy search approach. A single-layer LSTM is used and its hidden size is set to 256. The dropout is set to 0.3. While the image encoder module is non-trainable, the LSTM decoder module is trainable. During training, the cross-entropy loss is minimized, which is computed using the output logits and the tokens in the gold caption. Weights are optimized using the Adam optimizer [Kingma and Ba (2014] with an initial learning rate of 0.0001. Training is halted when the validation loss does not improve for 10 consecutive epochs.

The results of the image captioning in terms of BLEU scores are shown in Table 5. We observe that the BLEU scores of the generated image captions are much lower than the translation-based captions.

This is not very surprising because automatic captioning is free to choose a very different aspect of the image or use wording very different from the reference caption. BLEU only checks for n-gram overlap between the caption and the reference. Therefore, we perform a manual evaluation to further analyze the performance of the image caption generation model.

3.1. Manual Evaluation

A sample of about 10% of the generated captions was manually evaluated and categorized into the following classes:

for captions that describe the object of interest provided in the reference caption, exactly or closely.

for captions that describe a different object within the region of interest.

for captions that describe an object in the image that is outside the region of interest.

for captions that do not describe any object in the associated image.

Figure 4 presents the result of the manual evaluation of the sampled machine-generated captions. From the evaluated sample, it was observed that about 68% of the generated captions correctly describe an object in the image. Of this number, about 54% of the captions describe an object in the region of interest. However, most of the descriptions, although correct, do not match the description given in the reference caption (our evaluation does not quantify this aspect.)

This explains the low BLEU scores reported in Table 2. A more appropriate metric may be needed, therefore, to correctly measure the performance of such systems.

In Table 6, we provide examples of each of these manual evaluation classes.

Conclusion and Future Work

We present the HaVG, the multimodal dataset suitable for English→\rightarrowHausa machine translation, image captioning, and multimodal research.

The dataset is freely available for research and non-commercial usage under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License https://creativecommons.org/licenses/by-nc-sa/4.0/ at: http://hdl.handle.net/11234/1-4749.

In future versions of the HaVG, we plan to create the dataset from scratch without relying on an initial MT system and post-editing. Other future works include i) organizing a shared task using the HaVG, ii) extending the HaVG corpus for Visual Question Answering (VQA).

Acknowledgements

This work has received funding from the grant 19-26934X (NEUREM3) of the Czech Science Foundation, and has also been supported by the Ministry of Education, Youth and Sports of the Czech Republic, Project No. LM2018101 LINDAT/CLARIAH-CZ. This work is also financed by National Funds through the Portuguese funding agency, FCT - Fundação para a Ciência e a Tecnologia, within project LA/P/0063/2020.

Bibliographical References

References