Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, Ben Kao
Introduction
Multimodal Machine Translation (MMT) aims at designing better translation systems by extending conventional text-only translation systems to take into account multimodal information, especially from visual modality (Specia et al., 2016; Wang et al., 2019). Despite many previous success in MMT that report improvements when models are equipped with visual information (Calixto et al., 2017; Helcl et al., 2018; Ive et al., 2019; Lin et al., 2020; Yin et al., 2020), there have been continuing debates on the need for visual context in MMT.
In particular, Specia et al. (2016); Elliott et al. (2017); Barrault et al. (2018) argue that visual context does not seem to help translation reliably, at least as measured by automatic metrics. Elliott (2018); Grönroos et al. (2018a) provide further evidence by showing that MMT models are, in fact, insensitive to visual input and can translate without significant performance losses even in the presence of features derived from unrelated images. A more recent study (Caglayan et al., 2019), however, shows that under limited textual context (e.g., noun words are masked), models can leverage visual input to generate better translations. But it remains unclear where the gains of MMT methods come from, when the textual context is complete.
The main tool utilized in prior discussion is adversarial model comparison — explaining the behavior of complex and black-box MMT models by comparing performance changes when given adversarial input (e.g., random images). Although such an opaque tool is an acceptable beginning to investigate the need for visual context in MMT, they provide rather indirect evidence Hessel and Lee (2020). This is because performance differences can often be attributed to factors unrelated to visual input, such as regularization Kukačka et al. (2017), data bias Jabri et al. (2016), and some others Dodge et al. (2019).
From these perspectives, we revisit the need for visual context in MMT by designing two interpretable models. Instead of directly infusing visual features into the model, we design learnable components, which allow the model to voluntarily decide the usefulness of the visual features and reinforce their effects when they are helpful. To our surprise, while our models are shown to be effective on Multi30k (Elliott et al., 2016) and VaTex (Wang et al., 2019) datasets, they learn to ignore the multimodal information. Our further analysis suggests that under sufficient textual context, the improvements come from a regularization effect that is similar to random noise injection (Bishop, 1995) and weight decay (Hanson and Pratt, 1989). The additional visual information is treated as noise signals that can be used to enhance model training and lead to a more robust network with lower generalization error (Salamon and Bello, 2017). Repeating the evaluation under limited textual context further substantiates our findings and complements previous analysis (Caglayan et al., 2019).
Our contributions are twofold. First, we revisit the need for visual context in the popular task of multimodal machine translation and find that: (1) under sufficient textual context, the MMT models’ improvements over text-only counterparts result from the regularization effect (Section 5.2). (2) under limited textual context, MMT models can leverage visual context to help translation (Section 5.3). Our findings highlight the importance of MMT models’ interpretability and the need for a new benchmark to advance the community.
Second, for the MMT task, we provide a strong text-only baseline implementation and two models with interpretable components that replicate similar gains as reported in previous works. Different from adversarial model comparison methods, our models are interpretable due to the specifically designed model structure and can serve as standard baselines for future interpretable MMT studies. Our code is available at https://github.com/LividWo/Revisit-MMT.
Background
One can broadly categorize MMT systems into two types: (1) Conventional MMT, where there is gold alignment between the source (target) sentence pair and a relevant image and (2) Retrieval-based MMT, where systems retrieve relevant images from an image corpus as additional clues to assist translation.
Most MMT systems require datasets consist of images with bilingual annotations for both training and inference. Many early attempts use a pre-trained model (e.g., ResNet (He et al., 2016)) to encode images into feature vectors. This visual representation can be used to initialize the encoder/decoder’s hidden vectors (Elliott et al., 2015; Libovický and Helcl, 2017; Calixto et al., 2016). It can also be appended/prepended to word embeddings as additional input tokens (Huang et al., 2016; Calixto and Liu, 2017). Recent works (Libovický et al., 2018; Zhou et al., 2018; Ive et al., 2019; Lin et al., 2020) employ attention mechanism to generate a visual-aware representation for the decoder. For instance, Doubly-ATT (Calixto et al., 2017; Helcl et al., 2018; Arslan et al., 2018) insert an extra visual attention sub-layer between the decoder’s source-target attention sub-layer and feed-forward sub-layer. While there are more works on engineering decoders, encoder-based approaches are relatively less explored. To this end, Yao and Wan (2020) and Yin et al. (2020) replace the vanilla Transformer encoder with a multi-modal encoder.
Besides the exploration on network structure, researchers also propose to leverage the benefits of multi-tasking to improve MMT Elliott and Kádár (2017); Zhou et al. (2018). The Imagination architecture Elliott and Kádár (2017); Helcl et al. (2018) decomposes multimodal translation into two sub-tasks: translation task and an auxiliary visual reconstruction task, which encourages the model to learn a visually grounded source sentence representation.
Retrieval-based MMT
The effectiveness of conventional MMT heavily relies on the availability of images with bilingual annotations. This could restrict its wide applicability. To address this issue, Zhang et al. (2020) propose UVR-NMT that integrates a retrieval component into MMT. They use TF-IDF to build a token-to-image lookup table, based on which images sharing similar topics with a source sentence are retrieved as relevant images. This creates image-bilingual-annotation instances for training. Retrieval-based models have been shown to improve performance across a variety of NLP tasks besides MMT, such as question answering (Guu et al., 2020), dialogue (Weston et al., 2018), language modeling (Khandelwal et al., 2019), question generation (Lewis et al., 2020), and translation (Gu et al., 2018).
Method
In this section we introduce two interpretable MMT models: (1) Gated Fusion for conventional MMT and (2) Dense-Retrieval-augmented MMT (RMMT) for retrieval-based MMT. Our design philosophy is that models should learn, in an interpretable manner, to which degree multimodal information is used. Following this principle, we focus on the component that integrates multimodal information. In particular, we use a gating matrix Yin et al. (2020); Zhang et al. (2020) to control the amount of visual information to be blended into the textual representation. Such a matrix facilitates interpreting the fusion process: a larger gating value indicates that the model exploits more visual context in translation, and vice versa.
Given a source sentence of length and an associated image , we compute the probability of generating target sentence of length by:
We next generate a gating matrix to control the fusion of and :
where and are model parameters. Note that this gating mechanism has been a building block for many recent MMT systems (Zhang et al., 2020; Lin et al., 2020; Yin et al., 2020). We are, however, the first to focus on its interpretability. Finally, we generate the output vector by:
is then fed into the decoder directly for translation as in vanilla Transformer.
2 Retrieval-Augmented MMT (RMMT)
RMMT consists of two sequential components: (1) an image retriever that takes as input and returns Top- most relevant images from an image database; (2) a multi-modal translator that generates each conditioned on the input sentence , the image set returned by the retriever, and the previously generated tokens .
Based on the TF-IDF model, searching in existing retrieval-based MMT Zhang et al. (2020) ignores the context information of a given query, which could lead to poor performance. To improve the recall of our image retriever, we compute the similarity between a sentence and an image with inner product:
where and are -dimensional representations of and , respectively. We then retrieve top- images that are closest to . For , we compute it by Eq. 2. For , we implement it using BERT (Devlin et al., 2019):
Multimodal Translator
Here, denotes a max-pooling layer with window size .
Experiment
In this section, we evaluate our models on the Multi30k and VaTex benchmark.
We perform experiments on the widely-used MMT datasets: Multi30k. We follow a standard split of 29,000 instances for training, 1,014 for validation and 1,000 for testing (Test2016). We also report results on the 2017 test set (Test2017) with extra 1,000 instances and the MSCOCO test set that includes 461 more challenging out-of-domain instances with ambiguous verbs. We merge the source and target sentences in the officially preprocessed version of Multi30khttps://github.com/multi30k/dataset to build a joint vocabulary. We then apply the byte pair encoding (BPE) algorithm (Sennrich et al., 2016) with 10,000 merging operations to segment words into subwords, which generates a vocabulary of 9,712 (9,544) tokens for En-De (En-Fr).
Retriever pre-training. We pre-train the retriever on a subset of the Flickr30k dataset (Plummer et al., 2015) that has overlapping instances with Multi30k removed. We use Multi30k’s validation set to evaluate the retriever. We measure the performance by recall-at- (), which is defined as the fraction of queries whose closest images retrieved contain the correct images. The pre-trained retriever achieves of and of .
2 Setup
We experiment with different model sizes (Base, Small, and Tiny, see Appendix A for details). Base is a widely-used model configuration for Transformer in both text-only translation (Vaswani et al., 2017) and MMT (Grönroos et al., 2018b; Ive et al., 2019). However, for small datasets like Multi30k, training such a large model (about 50 million parameters) could cause overfitting. In our preliminary study, we found that even a Small configuration, which is commonly used for low-resourced translation (Zhu et al., 2019), can still overfit on Multi30k. We therefore perform grid search on the EnDe validation set in Multi30k and obtain a Tiny configuration that works surprisingly well.
We use Adam with , for model optimization. We start training with a warm-up phase (2,000 steps) where we linearly increase the learning rate from to 0.005. Thereafter we decay the learning rate proportional to the number of updates. Each training batch contains at most 4,096 source/target tokens. We set label smoothing weight to 0.1, dropout to 0.3. We follow Zhang et al. (2020) to early-stop the training if validation loss does not improve for ten epochs. We average the last ten checkpoints for inference as in Vaswani et al. (2017) and Wu et al. (2018). We perform beam search with beam size set to 5. We report 4-gram BLEU and METEOR scores for all test sets. All models are trained and evaluated on one single machine with two Titan P100 GPUs.
3 Baselines
Our baselines can be categorized into three types:
The conventional MMT models: Doubly-ATT and Imagination;
Details of these methods can be found in Section 2. For fairness, all the baselines are implemented by ourselves based on FairSeq Ott et al. (2019). We use top-5 retrieved images for both UVR-NMT and our RMMT. We also consider two more recent state-of-the-art conventional methods for reference: GMNMT (Yin et al., 2020) and DCCN (Lin et al., 2020), whose results are reported as in their papers.
Note that most MMT methods are difficult (or even impossible) to interpret. While there exist some interpretable methods (e.g., UVR-NMT) that contain gated fusion layers similar to ours, they perform sophisticated transformations on visual representation before fusion, which lowers the interpretability of the gating matrix. For example, in the gated fusion layer of UVR-NMT, we observe that the visual vector is order-of-magnitude smaller than the textual vector. As a result, interpreting gating weight is meaningless because visual vector has negligible influence on the fused vector.
4 Results
Table 1 shows the BLEU scores of these methods on the Multi30k dataset. From the table, we see that although we can replicate similar BLEU scores of Transformer-Base as reported in Grönroos et al. (2018b); Ive et al. (2019), these scores (Row 1) are significantly outperformed by Transformer-Small and Transformer-Tiny, which have fewer parameters. This shows that Transformer-Base could overfit the Multi30k dataset. Transformer-Tiny, whose number of parameters is about times smaller than that of Transformer-Base, is more robust and efficient in our test cases. We therefore use it as the base model for all our MMT systems in the following discussion.
Based on the Transformer-tiny model, both our proposed models (Gated Fusion and RMMT) and baseline MMT models (Doubly-ATT, Imagination and UVR-NMT) significantly outperform the state-of-the-arts (GMNMT and DCCN) on EnDe translation. However, the improvement of all these methods (Rows 4-10) over the base Transformer-Tiny model (Row 3) is very marginal. This shows that visual context might not be as important as we expected for translation, at least on datasets we explored.
We further evaluate all the methods on the METEOR scores (see Appendix C). We also run experiments on the VaTex dataset (see Appendix B). Similar results are observed as Table 1. Although various MMT systems have been proposed recently, a well-tuned model that uses text only remain competitive. This motivates us to revisit the importance of visual context for translation in MMT models.
Model Analysis
Taking a closer look at the results given in the previous section, we are surprised by the observation that our models learn to ignore visual context when translating (Sec 5.1). This motivates us to revisit the contribution of visual context in MMT systems (Sec 5.2). Our adversarial evaluation shows that adding model regularization achieves comparable results as incorporating visual context. Finally, we discuss when visual context is needed (Sec 5.3) and how these findings could benefit future research.
To explore the need for visual context in our models, we focus on the interpretable component: the gated fusion layer (see Equation 3 and 5). Intuitively, a larger gating weight indicates the model learns to depend more on visual context to perform better translation. We quantify the degree to which visual context is used by the micro-averaged gating weight . Here , are the total number of sentences and words in the corpus, respectively. add up all elements in a given matrix, and is a scalar value ranges from 0 to 1. A larger implies more usage of the visual context.
We first study models’ behavior after convergence. From Table 2, we observe that is negligibly small, suggesting that both models learn to discard visual context. In other words, visual context may not be as important for translation as previously thought. Since is insensitive to outliers (e.g., large gating weight at few dimensions), we further compute -: percentage of gating weight entries in that are larger than -. With no surprise, we find that on all test splits - are always zero, which again shows that visual input is not used by the model in inference.
The Gated Fusion’s training process also shed some light on how the model accommodates the visual information during training. Figure 1 (a) and (b) shows how changes during training, from the first epoch. We find that, Gated Fusion starts with a relatively high (0.5), but quickly decreases to after the first epoch. As the training continues, gradually decreases to roughly zero. In the early stages, the model relies heavily on images, possibly because they could provide meaningful features extracted from a pre-trained ResNet-50 CNN, while the textual encoder is randomly initialized. Compared with text-only NMT, utilizing visual features lowers MMT models’ trust in the hidden representations generated from the textual encoders. As the training continues, the textual encoder learns to represent source text better and the importance of visual context gradually decreases. In the end, the textual encoder carries sufficient context for translation and supersedes the contributions from the visual features. Nevertheless, this doesn’t explain the superior performance of the multimodal systems (Table 1). We speculate that visual context is acting as regularization that helps model training in the early stages. We further explore this hypothesis in the next section.
2 Revisit need for visual context in MMT
In the previous section, we hypothesize that the gains of MMT systems come from some regularization effects. To verify our hypothesis, we conduct experiments based on two widely used regularization techniques: random noise injection (Bishop, 1995) and weight decay (Hanson and Pratt, 1989). The former simulates the effects of assumably uninformative visual representations and the later is a more principled way of regularization that does not get enough attention in the current hyperparameter tuning stage. Inspecting the results, we find that applying these regularization techniques achieves similar gains over the text-only baseline as incorporating multimodal information does.
For random noise injection, we keep all hyper-parameters unchanged but replace visual features extracted using ResNet with randomly initialized vectors, which are noise drawn from a standard Gaussian distribution. A MMT model equipped with ResNet features is denoted as a ResNet-based model, while the same model with random initialization is denoted as a noise-based model. We run each experiment three times and report the averaged results. Note that values in parentheses indicate the performance gap between the ResNet-based model and its noise-based adversary.
Table 3 shows BLEU scores on the Multi30k dataset. Each column in the table corresponds to a test set “contest”. From the table, we observe that, among 18 (3 methods 3 test sets 2 tasks) contests with the Transformer model (row 1), noise-based models (rows 2-4) achieve better performance 13 times, while ResNet-based models win 14 cases. This shows that noise-based models perform comparably with ResNet-based models. A further comparison between noise-based models and ResNet-based models shows that they are compatible after 18 contests, in which the former wins 8 times and the latter wins 10 times.
Further, we regularize the models with weight decay. We consider three models: the text-only Transformer, the representative existing MMT method Doubly-ATT, and our Gated Fusion method. Figure 2 and 3 (in Appendix C) show the BLEU and METEOR scores of these methods on En→De translation as weight decay rate changes, respectively. We see that the best results of the text-only Transformer model with fine-tuned weight decay are comparable or even better than that of the MMT models Doubly-ATT and Gated Fusion that utilize visual context. This again shows that visual context is not as useful as we expected and it essentially plays the role of regularization.
3 When is visual context needed in MMT
Despite the less importance of visual information we showed in previous sections, there also exist works that support its usefulness. For example, Caglayan et al. (2019) experimentally show that, with limited textual context (e.g., masking some input tokens), MMT models will utilize the visual input for translation. This further motivates us to investigate when visual context is needed in MMT models. We conduct experiment with a new masking strategy that does not need any entity linking annotations as in Caglayan et al. (2019). Specifically, we follow Tan and Bansal (2020) to collect a list of visually grounded tokens. A visually grounded token is the one that has more than 30 occurrences in the Multi30k dataset with stop words removed. Masking all visually grounded tokens will affect around 45% of tokens in Multi30k.
Table 4 shows the adversarial study with visually grounded tokens masked. In particular, we select Transformer, Gated Fusion and RMMT as representative methods. From the table, we see that random noise injection (row 5,6) and weight decay (row 2) can only bring marginal improvement over the text-only Transformer model. However, ResNet-based models that utilize visual context significantly improve the translation results. For example, RMMT achieves almost 50% gain over the Transformer on the BLEU score. Moreover, both Gated Fusion and RMMT using ResNet features lead to a larger value than that when textual context is sufficient as shown in Table 2. Those results further suggest that visual context is needed when textual context is insufficient. In addition to token masking, sentences with incorrect, ambiguous and gender-neutral words Frank et al. (2018) might also need visual context to help translation. Therefore, to fully exert the power of MMT systems, we emphasize the need for a new MMT benchmark, in which visual context is deemed necessary to generate correct translation.
Interestingly, even with ResNet features, we observe a significant drop in both BLEU and METEOR scores compared with those in Table 1 and 8, similar to that reported in (Chowdhury and Elliott, 2019). The reason could be two-fold. On the one hand, there are many words that can not be visualized. For example, in Table 5 (a), although Gated Fusion can successfully identify the main objects in the image (“little boys pose with a puppy”), it fails to generate the more abstract concept “family picture”. On the other hand, when translating different words, it is difficult to capture correct regions in images. For example, in Table 5 (b), we see that Gated Fusion incorrectly generates the word frauen (women) because it captures the woman at the top-right corner of the image.
4 Discussion
Finally, we discuss how our findings might benefit future MMT research. First, a benchmark that requires more visual information than Multi30k to solve is desired. As shown in Section 5.2, sentences in Multi30k are rather simple and easy-to-understand. Thus textual context could provide sufficient information for correct translation, making visual modules relatively redundant in these systems. While the MSCOCO test set in Multi30k contains ambiguous verbs and encourages models to use image sources for disambiguation, we still lack a corresponding training set.
Second, our methods can serve as a verification tool to investigate whether visual grounding is needed in translation for a new benchmark.
Third, we find that visual feature selection is also critical for MMT’s performance. While most methods employ the attention mechanism to learn to attend relevant regions in an image, the shortage of annotated data could impair the attention module (see Table 5 (b)). Some recent efforts Yin et al. (2020); Lin et al. (2020); Caglayan et al. (2020) address the issue by feeding models with pre-extracted visual objects instead of the whole image. However, these methods are easily affected by the quality of the extracted objects. Therefore, a more effective end-to-end visual feature selection technique is needed, which can be further integrated into MMT systems to improve performance.
Conclusion
In this paper we devise two interpretable models that exhibit state-of-the-art performance on the widely adopted MMT datasets — Multi30k and the new video-based dataset — VaTex. Our analysis on the proposed models, as well as on other existing MMT systems, suggests that visual context helps MMT in the similar vein as regularization methods (e.g., weight decay), under sufficient textual context. Those empirical findings, however, should not be understood as us downplaying the importance existing datasets and models; we believe that sophisticated MMT models are necessary for effective grounding of visual context into translation. Our goal, rather, is to (1) provide additional clarity on the remaining shortcomings of current dataset and stress the need for new datasets to move the field forward; (2) emphasise the importance of interpretability in MMT research.
Acknowledgement
Zhiyong Wu is partially supported by a research grant from the HKU-TCL Joint Research Centre for Artificial Intelligence.
References
Appendix A Training Settings
Table 6 shows the configuration of different model sizes.
Appendix B Results on VaTex
The results are shown in Table 7. We observe that although most MMT systems show improvement over the Transformer baseline, the gains are quite marginal. Indicating that although image-based MMT models can be directly applied to video-based MMT, there is still room for improvement due to the challenge of video understanding. We also note that (a) regularize the text-only Transformer with weight decay demonstrates similar gains as injecting video information into the models; (b) replacing video features with random noise replicate comparable performance, which further supports our findings in Section 5.2.
Appendix C Results on METEOR
We also report our results based on METEOR (Banerjee and Lavie, 2005), which consistently demonstrates higher correlation with human judgments than BLEU does in independent evaluations such as in EMNLP WMT 2011 http://statmt.org/wmt11/papers.html. From Table 8, we can see that on En-Fr translation, MMT systems demonstrate similar improvements over text-only baselines in both METEOR and BLEU(see Table 1). On En-De translation, however, MMT systems are mostly on-par with Transformer-tiny on METEOR and do not show consistent gains as BLEU. We hypothesis the reason being that En-De sets are created in a image-blind fashion, in which the crowd-sourcing workers produce translations without seeing the images (Frank et al., 2018). Such that source sentence can already provide sufficient context for translation. When creating the En-Fr corpus, the image-blind issue is fixed (Elliott et al., 2017), thus images are perceived as “needed” in the translation for whatever reason. Although BLEU is unable to elicit this difference, evaluation based on METEOR captured it and confirmed previous research. We also compute METEOR scores for our experiments that regularize models with random noise (see Table 9) and weight decay (see Figure 3). The results are consistent with those evaluated using BLEU and further complement our early findings.
Appendix D Results on IWSLT’14
We also evaluate the retrieval-based model RMMT on text-only corpus — IWSLT’14. The IWSLT’14 dataset contains 160k bilingual sentence pairs for En-De translation task. Following the common practice, we lowercase all words, split 7k sentence pairs from the training dataset for validation and concatenate dev2010, dev2012, tst2010, tst2011, tst2012 as the test set. The number of BPE operations is set to 20,000. We use the Small configuration in all our experiments. The dropout and label smoothing rate are set to 0.3 and 0.1, respectively. Since there is no images associated with IWSLT, we follow Zhang et al. (2020) and retrieve top-5 images from Multi30K corpus.
From Table 10, we see that Transformer without weight decay is marginally outperformed by RMMT, but achieves slightly higher BLEU scores when trained with a 0.0001 weight decay. Our discussion in Section 5.2 sheds light on why visual context is helpful on non-grounded low-resourced datasets like IWSLT’14 — for low-resourced dataset like IWSLT’14, injecting visual context help regularize model training and avoid overfitting.