Dynamic Context-guided Capsule Network for Multimodal Machine Translation
Huan Lin, Fandong Meng, Jinsong Su, Yongjing Yin, Zhengyuan Yang, Yubin Ge, Jie Zhou, Jiebo Luo
Introduction
MMT significantly extends the conventional text-based machine translation by taking corresponding images as additional inputs (Specia et al., 2016; Elliott et al., 2017; Barrault et al., 2018). The assumption behind this is that the translation is expected to be more accurate compared to purely text-based translation, since the visual context helps to solve data sparsity and ambiguity problems (Ive et al., 2019). As shown in Figure 1, with the help of the image, MMT model is able to correctly translate “bank” to “erdwall”. Overall, the research on MMT is of great significance. On the one hand, similar to other multimodal tasks such as image captioning (Chen et al., 2017a, b; Song et al., 2019; Guo et al., 2019) and visual question answering (Zhou et al., 2017; Fang et al., 2018; Peng et al., 2019; Liu et al., 2019), MMT involves computer vision and natural language processing (NLP) and proves the effectiveness of visual features in translation tasks. In other words, it not only requires an algorithm with in-depth understanding of visual contexts, but also connects its interpretation with a language model to create a natural sentence. On the other hand, MMT has wide applications, such as translating multimedia news, product information and movie subtitles (Zhou et al., 2018). Therefore, MMT has become an attractive but challenging multimodal task.
Very importantly, one of the key issues in MMT is how to effectively utilize visual features during the process of translation. To achieve this goal, three categories of methods have been investigated: (1) exploiting visual features as global visual context (Huang et al., 2016; Calixto and Liu, 2017; Grönroos et al., 2018); (2) applying attention mechanism to extract visual context, where one common approach is to employ a timestep-specific attention mechanism to extract visual context (Calixto et al., 2017; Delbrouck and Dupont, 2017c; Helcl et al., 2018) and another way is to use source hidden states to consider visual features and then use the obtained invariant visual context as a complement to source hidden states (Delbrouck and Dupont, 2017b; Arslan et al., 2018); (3) learning multimodal joint representations (Elliott and Kádár, 2017; Zhou et al., 2018; Calixto et al., 2019). Despite their successes, these approaches still have various shortcomings. First, global visual context and learning multimodal joint representations can not encode the observed variability when generating translation. Second, according to previous studies (Lu et al., 2016; Wu et al., 2018), extracting visual context is beyond the capacity of a single-step attention due to the complexity of multimodal tasks. Although multiple attention layers can refine the extraction of visual context, the improvement is still limited. One possible reason is that too many parameters of multiple attention layers make the model vulnerable to over-fitting, especially when limited training examples are given in MMT. Moreover, the above methods only use visual features at global or regional level, which is unable to offer sufficient visual guidance. As a result, visual features are not fully utilized, limiting the potential of MMT models.
To overcome these issues, in this paper, we propose a novel Dynamic Context-guided Capsule Network (DCCN) for MMT. At each timestep of decoding, we first employ the standard source-target attention to produce a timestep-specific source-side context vector. Next, DCCN takes this context vector as input and uses it to guide the iterative extraction of related visual context during the dynamic routing process, where a multimodal context vector is updated simultaneously. In particular, to fully exploit image information, we employ DCCN to extract visual features at two complementary granularities: global visual features and regional visual features, respectively. In this way, we can obtain two multimodal context vectors that are then fused for the prediction of the current target word. Compared with previous studies, DCCN is able to dynamically extract visual context without introducing a large number of parameters, which is suitable to model such kind of variability observed in machine translation. Potentially, DCCN learns a better multimodal joint representation for MMT. Therefore, it is also applicable to other related tasks that require a joint representation of two different modalities, such as visual question answering. In summary, the major contributions of our work are three-fold:
We introduce a capsule network to effectively capture visual features at different granularities for MMT, which has advantage of effectively capturing visual features without explosive parameters. To the best of our knowledge, our work is the first attempt to explore a capsule network to extract visual features for MMT.
We propose a novel context-guided dynamic routing for the capsule network, which uses the timestep-specific source-side context vector as the guiding signal to dynamically produce a multimodal context vector for MMT.
We conduct experiments on the Multi30K English-German and English-French datasets. Experimental results show that our model significantly outperforms several competitive MMT models.
Related Work
The related work mainly includes the studies of multimodal context modeling in MMT and capsule networks.
Multimodal context modeling in MMT. How to fully exploit context for neural machine translation has always been a hot research topic (Zhang et al., 2016; Wang et al., 2017; Zhang et al., 2018a; Su et al., 2018b, a; Zhang et al., 2018b; Zeng et al., 2018; Su et al., 2019b, a), which is the same for MMT. The commonly-used approaches to extract multimodal context for MMT can be classified into three categories: (1) Learning global visual features for MMT. For instance, Huang et al. (2016) concatenated global and regional visual features with source sequences. Calixto and Liu (2017) utilized global visual features as additional tokens in the source sequence, to initialize the encoder hidden states or initialize the first decoder hidden state. (2) Leveraging attention mechanism to exploit visual features. In this aspect, Caglayan et al. (2016) and Calixto et al.(2017) incorporated spatial visual features into the MMT model via an independent attention mechanism. Furthermore, Delbrouck and Dupont (2017c) employed Compact Bilinear Pooling to fuse the attention-based context vectors of two modalities. Meanwhile, Libovický and Helcl (2017) explored flat and hierarchical combinations to fuse the attention-based context vectors of two modalities. Instead of using attention mechanism, Grönroos et al. (2018) introduced a gating layer to modify the prediction distribution on both visual features and decoder states. Unlike previous studies, Delbrouck and Dupont (2017b) utilized the attention mechanism on visual inputs for the source hidden states. Along this line, Arslan et al. (2018) extended this approach into Transformer, and Helcl et al. (2018) used timestep-specific source-side context vector as attention query to dynamically produce the visual context vectors. (3) Applying multi-task learning to jointly model translation task with other visual related tasks. For example, Elliott and Kádár (2017) decomposed multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. Zhou et al. (2018) optimized the learning of a shared visual language embedding and a multimodal attention-based translator. Calixto et al. (2019) introduced a continuous latent variable for MMT, which contains the underlying semantic information extracted from texts and images. Recently, Yin et al. (2020) uses a unified multi-modal graph to capture various semantic relationships between multi-modal semantic units.
Capsule Network. Recently, capsule network has been widely used in computer vision (Xinyi and Chen, 2019; Li et al., 2018; Singh et al., 2019; Xiang et al., 2018; Jaiswal et al., 2018; McIntosh et al., 2018) and NLP (Xinyi and Chen, 2019; Chen and Qian, 2019; Yang et al., 2018; Aly et al., 2019) tasks. Specific to machine translation, Wang et al. (2019) employed dynamic routing algorithm to model child-parent relationships between lower and higher encoder layers. Yang et al. (2019) proposed a query-guided capsule networks to cluster context information into different perspectives from which the target translation may concern. Zheng et al. (2019) separated translated and untranslated source words into different groups of capsules.
To the best of our knowledge, our work is the first attempt to introduce the capsule network into MMT. Furthermore, we use the timestep-specific source-side context vector rather than static source hidden states to guide the routing procedure. Therefore, we can dynamically produce visual context for MMT.
Our Model
As shown in Figure 2, our model is based on Transformer (Vaswani et al., 2017). The most important feature of our model is that two DCCNs are introduced to dynamically learn multimodal context vectors for generating translations.
Given the source sentence , we represent each source word as the sum of its word embedding and positional encoding. Next, we follow Vaswani et al. (2017) to use a stack of identical layers to encode , where each layer consists of two sub-layers. Note that we also introduce residual connection and layer normalization to each sub-layer, of which the descriptions are omitted.
Specifically, at the -th layer (), the first sub-layer is a multi-head self-attention:
where and are learnable parameter matrices.
The second sub-layer is a position-wise fully connected feed-forward network. It is applied to each position separately and identically, forming the representation of the source sentence as
2. Decoder
As shown in Figure 2, our decoder is an extension of the Transformer decoder (Vaswani et al., 2017). It takes the already generated sequence as inputs and uses a stack of identical layers to produce target-side hidden states. Similar to the standard Transformer decoder, each layer of our decoder contains three sub-layers. The only difference is that at the last decoder layer, two DCCNs are equipped to produce timestep-specific multimodal context vectors for MMT.
Specifically, the first sub-layer is also a multi-head self-attention:
where denotes the temporary decoder hidden states, produced by a multi-head self-attention mechanism fed with the target-side hidden states at the previous layer.
Typically, at the second sub-layer, a multi-head source-target attention mechanism is used to dynamically produce the timestep-specific source-side context vectors :
The third sub-layer is a position-wise fully connected feed-forward neural network, of which the definition depends on the decoder layer. At the first layers, this sub-layer produces the target-side hidden states as follows:
Very importantly, at the last (-th) decoder layer, we introduce two DCCNs between the second and third sub-layers to learn multimodal context representations at the -th timestep as follows:
where and are learnable parameters. Correspondingly, the Eq. 6 at the last decoder layer becomes
Finally, with the target-side hidden states generated by Eq. 11, our decoder adopts a Softmax layer to generate the probability distribution of the current target word :
To fully exploit visual information for MMT, we investigate two kinds of visual features to enhance text-based translation: (1) global visual features, which represent an input image with high-level concepts. Here we use the res4f layer activations of pre-trained 50-layer Residual Network (ResNet-50)(He et al., 2016) as global visual features. These spatial features encode an image in a 14×14 grid, where each grid is represented by a 1,024D feature vector, only encoding the information about the specific part of the image. Before fed into DCCN, we first follow Calixto et al. (2017) to transform global visual features into a 196×256 matrix where each of the 196 rows consists of a 256D feature vector; and (2) regional visual features that illustrate class annotations of each region (e.g. cat, arms, peak). Following Anderson et al. (2018), we employ the R-CNN based bottom-up attention to identify the regions with class annotations. For each region, we generate the corresponding prediction probability distribution over 1,600 classes from Visual Genome. To represent each region as a vector, we project its class annotations into word embeddings and define the region vector as the weighted sum of its class annotation embeddings. Finally, all region vectors are concatenated to represent the semantics of input image. In practice, we keep the number of predicted regions up to 10 so as to reduce negative effects of abundant regions, therefore the regional visual features can be represented as a 10×256 matrix , where each of the 10 rows consists of a 256D feature vector.
2.2. Dynamic Context-guided Capsule Network
Our DCCN is a significant extension of the conventional capsule network. Thus, it retains the advantages on iterative feature extraction of capsule network, which has shown effective in many computer vision (Xinyi and Chen, 2019; Li et al., 2018; Singh et al., 2019; Xiang et al., 2018; Jaiswal et al., 2018; McIntosh et al., 2018) and NLP (Xinyi and Chen, 2019; Chen and Qian, 2019; Yang et al., 2018; Aly et al., 2019; Yang et al., 2019; Zheng et al., 2019; Wang et al., 2019) tasks. More importantly, unlike the conventional capsule network that only captures static visual context, DCCN introduces the timestep-specific source-side context vector to guide the extraction of multimodal context, which can model the observed variability during translation.
Figure 3 shows the architecture of DCCN. Similar to the conventional capsule network, DCCN consists of (1) low-level capsules (the first column from the right) encoding the visual features of the input image and (2) high-level capsules (the second column from the left) encoding the related visual features. Besides, it includes multimodal context capsules (the first column from the left), where indicates a temporary multimodal context vector iteratively updated with and is used as query to guide the extraction of related visual features.
Next, we introduce the coefficient to measure the cross-modal correlation between and the multimodal context vector (Line 13), which can be subsequently used to generate high-level capsules and update , benefiting the extraction of related visual features. Formally, the correlation function is defined as
where indicates the Pearson Correlation Coefficients, is a parameter matrix that maps to the same semantic space of , is the covariance and is the standard deviation. When is close to +1, the visual features encoded by is closely related to , otherwise it indicates negative correlation.
Then, we conduct iterations of routing to capture related visual features at current timestep (Line 16 to Line 32). At each iteration, we generate high-level capsules from low-level capsules. To do this, we employ a Softmax function along columns with the logits initialized as 0 to calculate the coupling coefficient (Line 19). Afterwards, we generate the high-level capsule to represent visual context as the weighted sum of according to their corresponding and the cross-modal correlation coefficient (Line 23). Note that unlike the conventional dynamic routing algorithm (Sabour et al., 2017) where only depends on and , we further introduce that enables the most relevant visual features to be iteratively clustered into high-level capsules. Note that the conventional capsule network (Sabour et al., 2017) uses the norm of high-level capsule to represent prediction probability, thus the norm of is adjusted to using the “squashing” function. Different from that, represents visual context in DCCN. Therefore we do not apply “squashing” function in our algorithm. Further, we introduce a transformation matrix to map visual context into the semantic space of multimodal context and follow Wu and Mooney (2019) to update with the captured visual context (Line 24). By doing so, we expect the updated multimodal context capsule can be better exploited to guide the routing procedure at the next iteration.
Finally, we update (Line 28), and then use it to guide the updating of (Line 29). Different from conventional dynamic routing algorithm, where only depends on the cumulative “agreement” between and , we control the updating range of according to the cross-modal correlation .
Through iterations of routing, we obtain multimodal context capsules , fused by a linear transformation to produce the final multimodal context vector (Line 33).
Experiments
To investigate the effectiveness of our proposed model, we conduct experiments on the Multi30K dataset (Elliott et al., 2016), which is an extended version of the Flickr30K Entities and has been widely used in MMT (Barrault et al., 2018; Elliott et al., 2017; Specia et al., 2016). For each image, one of the English (EN) descriptions was selected and manually translated into German (DE) and French (FR) by professional translators (Specia et al., 2016). The dataset contains 29,000 instances for training, 1,024 for validation and 1,000 for testing. We also report results on the WMT2017 test set with 1,000 instances and the MSCOCO test set containing 461 out-of-domain instances with ambiguous verbs. Besides, as mentioned in subsection 3.2.1, we represent the input image with visual features in two granularities.
We apply the MOSES scriptshttp://www.statmt.org/moses/ to preprocess datasets. We then employ the Byte Pair Encoding (Sennrich et al., 2016) with 10,000 merging operations to convert tokens into subwords.
2. Setup
We develop our proposed model based on OpenNMT Transformer (Klein et al., 2017). Since the size of training corpus is small and the trained model tends to be over-fitting, we first perform a small grid search to obtain a set of hyper-parameters on the ENDE validation set. Specifically, the layer numbers of both encoder and decoder are set to 4 and the number of attention heads is set to 8. Both hidden size and embedding size are set to 256.
As implemented in (Vaswani et al., 2017), we use the Adam optimizer with and scheduled learning rate to optimize various models. The learning rate is initialized as 1. During training, each batch consists of approximately 3,700 source and target tokens. Besides, we employ the dropout strategy (Srivastava et al., 2014) with rate 0.5 to enhance the robustness of our model. Finally, we adopt the MultEval scripts (Clark et al., 2011) to evaluate the translation quality in terms of BLEU (Papineni et al., 2002) and METEOR (Denkowski and Lavie, 2011). Particularly, we run all models three times for each experiment and report the average results.
The context-guided dynamic routing is important for the generation of multimodal context. Therefore, we investigate the impacts of its hyper parameters on the routing mechanism: high-level capsule number and routing iteration number . To this end, we try different numbers of high-level capsules and routing iteration numbers to train our model: from 1 to 3, from 1 to 4 on the validation set. We observe that larger than 1 and larger than 3 do not lead to significant improvements and increase the GPU memory requirement. Hence, we use =1 and =3 in all subsequent experiments.
3. Baselines
We directly refer to our MMT model as DCCN and compare it with the following commonly-used MMT baselines:
Transformer (Vaswani et al., 2017) A text-only machine translation model.
Encoder-attention (Delbrouck and Dupont, 2017b). It incorporates an encoder-based visual attention mechanism into Transformer, which uses source hidden states to consider visual features and then augment each source hidden state with its corresponding visual context. Please note that (Delbrouck and Dupont, 2017b) is based on RNN and we implement it on Transformer for comparability.
Doubly-attention (Helcl et al., 2018). A doubly attentive Transformer that introduces an additional visual attention sub-layer to exploit visual features. Specifically, this sub-layer is inserted between the source-target attention and feed-forward sub-layer. For visual attention, the context vectors from the source-target attention are used as queries, and the context vectors of visual attention are fed into the feed-forward sub-layer.
We also display the performance of several dominant MMT models on the same datasets. Stochastic attention (Delbrouck and Dupont, 2017a) is a stochastic and sampling-based attention mechanism, which focuses on only one spatial location of the image at every timestep. It is also the model of best performance in (Delbrouck and Dupont, 2017a). Imagination (Elliott and Kádár, 2017) employs multitask learning to jointly two sub-tasks: translating and visually grounded representation prediction. Fusion-conv (Caglayan et al., 2017) employs a single feed-forward network to establish the attention alignment between visual features and target-side hidden states at each timestep, where all spatial locations of image are considered to derive the context vector. Trg-mul (Caglayan et al., 2017) modulates each target word embedding with visual features using element-wise multiplication. Latent Variable MMT (Calixto et al., 2019) exploits the interactions between visual and textual features for MMT through a latent variable, which can be seen as a multimodal stochastic embedding of an image and its target language description. Deliberation Network (Ive et al., 2019) is based on a translate-and-refine strategy, where visual features are only used by the decoder at the second stage.
4. Results on the EN⇒⇒\RightarrowDE Translation Task.
Parameters Introducing visual feature features into Transformer model brings more parameters. As shown in Table 1, Transformer (Row 7) has 16.1M parameters. Encoder-attention (Row 8) adds a layer normalization, a fully connected layer and two multi-head attention layers, introducing 1.1M parameters, and Doubly attention (Row 9) increases 4.0M parameters by adding two layer normalization and two multi-head attention layers. By contrast, DCCN (Row 10) only introduces 1.0M extra parameters. Thus, DCCN introduces a small number of extra parameters compared to Transformer, and requires smaller parameters than two multimodal baselines.
Model Performance Table 1 shows the translation quality on ENDE translation task. It is obvious that DCCN outperforms most of the existing models and all baselines, except Fusion-conv (Row 3) and Trg-mul (Row 4) on METEOR. Note that these two systems are the state-of-the-arts on WMT 2017, with parameter selection based on METEOR. Moreover, we draw two interesting conclusions:
First, DCCN model outperforms Encoder-attention, which uses static source hidden states to attend to visual features. The underlying reasons consist of two aspects: (1) encoder-attention depends on static source representations to extract visual context. By contrast, DCCN utilizes the timestep-specific source-side context vector to extract visual context; and (2) the context-guided dynamic routing mechanism exploits the interactions between different modalities to produce better multimodal context vector.
Second, although Doubly-attention also uses timestep-specific source-side context vector, DCCN model still achieves a significant improvement over it. This demonstrates again the advantage of modeling the semantic interactions between different modalities on learning multimodal context vectors.
Finally, following Bahdanau et al. (2015), we divide our test sets into different groups based on the lengths of source sentences, and then compare different models in each group. Figure 4 reports the BLEU scores of two test sets on different groups. Overall, our model still consistently achieves the best performance in most groups. Thus, we confirm again the effectiveness and generality of DCCN.
5. Ablation Study
To explore the effectiveness of different components in DCCN, we further compare our model with the following variants in Table 2:
(1) dynamic routing (global). To build this variant, we only use global visual features in our model. The result in Row 2 indicates that removing the regional visual features lead to performance drop. This result suggests that regional visual features are indeed useful for multimodal representation learning.
(2) dynamic routing (regional). Unlike the above variant, we only use regional visual features to represent the input image in this variant. According to the result shown in Row 3, we observe this change results in a significant performance decline, demonstrating that global visual features also bring useful visual information to our model.
(3) dynamic routing attention. Apparently, one advantage of our model lies in leveraging context-guided dynamic routing to exploit the semantic interactions between different modalities for learning multimodal representation. Here we separately replace the context-guided dynamic routing with the conventional attention mechanism to exploit global visual features, regional visual features, and both of them. Then we investigate the change of model performance. To facilitate the following descriptions, we refer to these three variants as dynamic routing (global) + attention (regional), dynamic routing (regional) + attention (global), and attention (global) + attention (regional), respectively. From Row 4 to Row 6 of Table 2, we observe that dynamic routing (global) + attention (regional) slightly outperforms dynamic routing (global) while is inferior to DCCN. Similarly, the performance of dynamic routing (regional) + attention (global) is between dynamic routing (regional) and DCCN. Moreover, when we use attention mechanism rather than dynamic routing to extract two kinds of visual features, the performance of our model degrades most. Based on these experimental results, we can draw the conclusion that context-guided dynamic routing is able to better extract two types of visual features than conventional attention mechanism.
(4)w/o context guidance in dynamic routing. By removing the context guidance from DCCN, we adopt the standard capsule network to extract visual features. As shown in Row 7, the model performance drops drastically in this case. This result is consistent with our intuition that the ideal visual features required for translation should be dynamically captured at different timesteps.
6. Case Study
When encountering ambiguous source words or complicated sentences, it is difficult for MMT models to translate correctly without corresponding visual features. To further demonstrate the effectiveness of DCCN, we display the 1-best translations of the four cases generated by different models, as shown in Figure 5.
When encountering ambiguous nouns, the regional visual features are more helpful. For example, in case (a), both Transformer and Encoder-attention miss the translation of source word “cliff”, while Doubly-attention and DCCN translate it correctly according to the detected object “rock”. In case (b), we can find that the ambiguous source word “student” is translated to “étudiants (college students)” by all baselines, while only DCCN correctly translates it with “élèves (generally refers to all students)” with the help of the detected object “kid”.
When the model has difficulty in translating words out of the predicted objects such as verbs and adjectives, the global visual features are more helpful. In case (c), the source word “rides” is not associated with any object and thus all baselines choose “fährt (drive)”, while only DCCN translates it correctly. In case (d), three baselines translate “harvest rice” to “circulant (flow)”, “travaillent (work)” and “pagaient (paddle)”, respectively. By contrast, only DCCN can produce the correct translation with the help of image information.
These cases reveal that DCCN can fully utilize complementary visual information to learn more accurate representations and disambiguate during translation in different cases.
7. Results on the EN⇒⇒\RightarrowFR Translation Task
To investigate the generality of our proposed model, we also conduct experiments on the ENFR translation task. Table 3 reports the final experimental results. Likewise, no matter which evaluation metric is used, our model still achieves better performance than all baselines. This result strongly demonstrates again that DCCN is effective and general to different language pairs in MMT.
Conclusion
In this paper, we have proposed a novel context-guided capsule network (DCCN) for MMT. As a significant extension of the conventional capsule network, DCCN utilizes the timestep-specific source-side context vector to dynamically guide the extraction of visual features at different timesteps, where the semantic interactions between modalities can be fully exploited for MMT via context-guided dynamic routing mechanism. Moreover, we employ DCCN to extract visual features in two complementary granularities: global visual features and regional visual features, respectively. Experimental results on English-to-German and English-to-French MMT tasks strongly demonstrate the effectiveness of our model. In the future, we plan to apply DCCN to other multimodal tasks such as visual question answering and multimodal text summarization.
Acknowledgments
This work was supported by the Beijing Advanced Innovation Center for Language Resources (No. TYR17002), the National Natural Science Foundation of China (No. 61672440), and the Scientific Research Project of National Language Committee of China (No. YB135-49).