Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations
Po-Yao Huang, Xiaojun Chang, Alexander Hauptmann
Introduction
Joint visual-semantic embeddings (VSE) are central to the success of many vision-language tasks, including cross-modal search and retrieval Kiros et al. (2014); Karpathy and Fei-Fei (2015); Gu et al. (2018), visual question answering Antol et al. (2015); Goyal et al. (2017), multimodal machine translation Huang et al. (2016); Elliott and Kádár (2017), etc. Learning VSE requires extensive understanding of the content in individual modalities and an in-depth alignment strategy to associate the complementary information from multiple views.
With the availability of large-scale parallel English-Image corpora Lin et al. (2014); Young et al. (2014), a rich line of research has advanced learning VSE under the monolingual setup. Most recent works Kiros et al. (2014); Vendrov et al. (2015); Karpathy and Fei-Fei (2015); Klein et al. (2015); Wang et al. (2016, 2018); Huang et al. (2019b) leverage triplet ranking losses to align English sentences and images in the joint embedding space. In VSE++ Faghri et al. (2018), Faghri et al. improve VSE by emphasizing hard negative samples. Recent advancement in VSE models explores methods to enrich the English-Image corpora. Shi et al. (2018) propose to augment dataset with textual contrastive adversarial samples to combat adversarial attacks. Recently, Huang et al. (2019a) utilize textual semantics of regional objects and adversarial domain adaptation for learning VSE under low-resource constraints.
An emerging trend generalizes learning VSE in the multilingual scenario. Rajendran et al. (2016) learn -view representations when parallel data is available only between one pivot view and the rest of views. PIVOT Gella et al. (2017) extends the work from Calixto et al. (2017) to use images as the pivot view for learning multilingual multimodal representations. Kádár et al. (2018) further confirms the benefits of multilingual training.
Our work is motivated by Gella et al. (2017) but has important differences. First, to disentangle the alignments in the joint embedding space, we employ visual object detection and multi-head attention to selectively align salient visual objects with textual phrases, resulting in visually-grounded multilingual multimodal representations. Second, as multi-head attention Vaswani et al. (2017) is appealing for its efficiency and ability to jointly attend to information form different perspectives, we propose to further encourage the diversity among attention heads to learn an improved visual-semantic embedding space. Figure 1 illustrates our gradient updates promoting diversity. The proposed model achieves state-of-the-art performance in the multilingual sentence-image matching tasks in Multi30K Elliott et al. (2016) and the semantic textual similarity task Agirre et al. (2012, 2014, 2015).
Related Works
Classic attention mechanisms have been addressed for learning VSE. These mechanisms can be broadly categorized by the types of {Query, Key, Value} as discussed in Vaswani et al. (2017). For intra-modal attention, {Query, Key, Value} are within the same modality. In DAN Nam et al. (2017), the content in each modalities is iteratively attended through multiple steps with intra-modal attention. In SCAN Lee et al. (2018), inter-modal attention is performed between regional visual features from Anderson et al. (2018) and text semantics. The inference time complexity is (for generating query representations for a size datatset). In contrast to the prior works, we leverage intra-modal multi-head attention, which can be easily parallelized compared to DAN and is with a preferred inference time complexity compared to SCAN.
Inspired by the idea of attention regularization in Li et al. (2018), for learning VSE, we propose a new margin-based diversity loss to encourage a margin between attended outputs over multiple attention heads. Multi-head attention diversity within the same modality and across modalities are jointly considered in our model.
The Proposed Model
Figure 1 illustrates the overview of the proposed model. Given a set of images as the pivoting points with the associate English and GermanFor clarity in notation, we discuss only two languages. The proposed model can be intuitively generalized to more languages by summing additional terms in Eq. 4 and Eq. 6-7. descriptions or captions, the proposed VSE model aims to learn a multilingual multimodal embedding space in which the encoded representations of a paired instance are closely aligned to each other than non-paired ones.
Multi-head attention with diversity: We employ -head attention networks to attend to the visual objects in an image as well as the textual semantics in a sentence then generate fixed-length image/sentence representations for alignment. Specifically, the -th attended German sentence representation is computed by:
With where as the set of attended fixed-length image and sentence representations in a sampled batch, we use the widely-used hinge-based triplet ranking loss with hard negative mining Faghri et al. (2018) to align instances in the visual-semantic embedding space. Taking Image-English instances as an example, we leverage the triplet correlation loss defined as:
where is the correlation margin between positive and negative pairs, is the hinge function, and is the cosine similarity. and are the indexes of the images and sentences in the batch. and are the hard negatives. When the triplet loss decreases, the paired images and German sentences are drawn closer down to a margin than the hardest non-paired ones. Our model aligns , and in the joint embedding space for learning multilingual multimodal representations with the sampled batch. We formulate the overall triplet loss as:
Note that the hyper-parameter controls the contribution of since (, ) may not be a translation pair even though (, ) and (, ) are image-caption pairs.
One of the desired properties of multi-head attention is its ability to jointly attend to and encode different information in the embedding space. However, there is no mechanism to support that these attention heads indeed capture diverse information. To encourage the diversity among attention heads for instances within and across modalities, we propose a new simple yet effective margin-based diversity loss. As an example, the multi-head attention diversity loss between the sampled images and the English sentences (i.e. diversity across-modalities) is defined as:
As illustrated with the red arrows for update in Figure 1, the merit behind this diversity objective is to increase the distance (up to a diversity margin ) between attended embeddings from different attention heads for an instance itself or its cross-modal parallel instances. As a result, the diversity objective explicitly encourage multi-head attention to concentrate on different aspects of information sparsely located in the joint embedding space to promote fine-grained alignments between multilingual textual semantics and visual objects. With the fact that the shared embedding space is multilingual and multimodal, for improving both intra-modal/lingual and inter-modal/lingual diversity, we model the overall diversity loss as:
where the first three terms are intra-modal/lingual and the rest are cross-modal/lingual. With Eq. 4 and Eq. 6, we formalize the final model loss as:
where is the weighting parameter which balances the diversity loss and the triple ranking loss. We train the model by minimizing .
Experiments
Following Gella et al. (2017), we evaluate on the multilingual sentence-image matching tasks in Multi30K Elliott et al. (2016) and the semantic textual similarity task Agirre et al. (2012, 2015).
We use the model in Anderson et al. (2018) which is a Faster-RCNN Ren et al. (2015) network pre-trained on the MS-COCO Lin et al. (2014) dataset and fine-tuned on the Visual Genome Krishna et al. (2016) dataset to detect salient visual objects and extract their corresponding features. 1,600 types of objects are detectable. We then pack and represent each image as a feature matrix where 36 is the maximum amount of salient visual objects in an image and 2048 is the dimension of the flattened last pooling layer in the ResNet He et al. (2016) backbone of Faster-RCNN.
For the text processing, we lower-case, tokenize, and then truncate the maximum sentence length to 100. We use 300-dim word embedding matrices initialized either randomly or with pre-trained multilingual embeddings. (We use the multilingual version of FastText Mikolov et al. (2018)). We also experiment incorporating the last layer of contextualized multilingual BERT embeddings Devlin et al. (2018) to replace the word embedding matrices as the textual input features for the bi-directional LSTMs.
For training, we sample batches of size 128 and train 20 epochs on the training set of Multi30K. We use the Adam Kingma and Ba (2014) optimizer with learning rate then after 15-th epoch. Models with the best summation of validation R@1,5,10 are selected to generate image and sentence embeddings for testing. Weight decay is set to and gradients larger than are clipped. We use 3-head attention () and the embedding dimension . The same dimension is shared by all the context vectors in the attention modules. Other hyper-parameters are set as follows: and .
2 Multilingual Sentence-Image Matching
We evaluate the proposed model in the multilingual sentence-image matching (retrieval) tasks on Multi30K: (i) Searching images with text queries (Sentence-to-Image). (ii) Ranking descriptions with image queries (Image-to-Sentence). English and German are considered.
Multi30K Elliott et al. (2016) is the multilingual extension of Flickr30K Young et al. (2014). The training, validation, and testing split contain 29K, 1K, and 1K images respectively. Two types of annotations are available in Multi30K: (i) One parallel English-German translation for each image and (ii) five independently collected English and five German descriptions/captions for each image. We use the later. Note that the German and English descriptions are not translations of each other and may describe an image differently.
As other prior works, we use recall at (R@) to measure the standard ranking-based retrieval performance. Given a query, R@ calculates the percentage of test instances for which the correct one can be found in the top- retrieved instances. Higher R@ is preferred.
Table 1 presents the results on the Multi30K testing set. The VSE baselines in the first five rows are trained with English and German descriptions independently. In contrast, PIVOT Gella et al. (2017) and the proposed model are capable of handling multilingual input queries with single model. For a fair comparison with PIVOT, we also report the result of swapping Faster-RCNN with VGG as the visual feature encoder in our model.
As can be seen, the proposed models successfully obtain state-of-the-art results, outperforming other baselines by a significant margin. German-Image matching benefit more from joint training with English-Image pairs. The models with pre-trained multilingual embeddings and contextualized embeddings achieve better performance in comparison to randomly initialized word embeddings, especially for German. One explanation is that the degradation from German singletons is alleviated by the multi-task training with English and the pre-trained embeddings. While the model with BERT performs better in English, FastText is preferred for German-Image matching.
3 Semantic Textual Similarity Results
For semantic textual similarity (STS) tasks, we evaluate on the video task from STS-2012 Agirre et al. (2012) and the image tasks from STS-2014-15 Agirre et al. (2014, 2015). The video descriptions are from the MSR video description corpus Chen and Dolan (2011) and the image descriptions are from the PASCAL dataset Rashtchian et al. (2010). In STS, a system takes two input sentences and output a semantic similarity ranging from . Following Gella et al. (2017), we directly use the model trained on Multi30K to generate sentence embeddings then scaled the cosine similarity between the two embeddings as the prediction.
Table 2 lists the standard Pearson correlation coefficients between the system predictions and the STS gold-standard scores. We report the best scores achieved by paraphrastic embeddings Wieting et al. (2017) (text only) and the VSE models in the previous section. Note that the compared VSE models are all with RNN as the text encoder and no STS data is used for training. Our models achieve the best performance and the pre-trained word embeddings are preferred.
4 Qualitative Results and Grounding
In Figure 2 we samples some qualitative multilingual text-to-image matching results. In most cases our model successfully retrieve the one and only one correct image. Figure 3 depicts the t-SNE visualization of the learned visually grounded multilingual embeddings of the pairs pivoted on in the Multi30K testing set. As evidenced, although the English and German sentences describe different aspects of the image, our model correctly aligns the shared semantics (e.g. (“man”, “mann”), (“hat”, “wollmütze”)) in the embedding space. Notably, the embeddings are visually-grounded as our model associate the multilingual phrases with exact visual objects (e.g. glasses and ears). We consider learning grounded multilingual multimodal dictionary as the promising next step.
As limitations we notice that actions and small objects are harder to align. Additionally, the alignments tends to be noun-phrase/object-based whereas spatial relationships (e.g. “on”, “over”) and quantifiers remain not well-aligned. Resolving these limitations will be our future work.
Conclusion
We have presented a novel VSE model facilitating multi-head attention with diversity to align different types of textual semantics and visual objects for learning grounded multilingual multimodal representations. The proposed model obtains state-of-the-art results in the multilingual sentence-image matching task and the semantic textual similarity task on two benchmark datasets.
Acknowledgement
This research is supported in part by the DARPA grants FA8750-18-2-0018 and FA8750-19-2-0501 under AIDA and LwLL program. It is also supported by the IARPA grant via DOI/IBC number D17PC00340. We would like to thank the anonymous reviewers for their constructive suggestions.