Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
Yale Song, Mohammad Soleymani
Introduction
Visual-semantic embedding frome-nips13; karpathy-cvpr15 aims to find a joint mapping of instances from visual and textual domains to a shared embedding space so that related instances from source domains are mapped to nearby places in the target space. This has a variety of downstream applications in computer vision including tagging frome-nips13, retrieval gong-ijcv14, captioning karpathy-cvpr15, visual question answering jang-cvpr17.
where is a certain distance measure, such as Euclidean and cosine distance. This simple and intuitive setup, which we refer to as injective instance embedding, is currently the most popular approach in the literature wang-arxiv16.
Unfortunately, injective embedding can suffer when there is ambiguity in individual instances. Consider an ambiguous instance with multiple meanings/senses, e.g., polysemy words and images containing multiple objects. Even though each of the meanings/senses can map to different points in the embedding space, injective embedding is always forced to find a single point, which could be an (inaccurate) weighted geometric mean of all the desirable points. The issue gets intensified for videos and sentences because the ambiguity in individual images and words can aggregate and get compounded, severely limiting its use in real-world applications such as text-to-video retrieval.
Another case where injective embedding could be problematic is partial cross-domain association, a characteristic commonly observed in the real-world datasets. For instance, a text sentence may describe only certain regions of an image while ignoring other parts xu-icml15, and a video may contain extra frames not described by its associated sentence li-cvpr16. These associations are implicit/hidden, making it unclear which part(s) of the image/video the text description refers to. This is especially problematic for injective embedding because information about any ignored parts will be lost in the mapped point and, once mapped, there is no way to recover from the information loss.
In this work, we address the above issues by (1) formulating instance embedding as a one-to-many mapping task and (2) optimizing the mapping functions to be robust to ambiguous instances and partial cross-modal associations.
To address the issues with ambiguous instances, we propose a novel one-to-many instance embedding model, Polysemous Instance Embedding Network (PIE-Net), which extracts embeddings of each instance by combining global and local information of its input. Specifically, we obtain locally-guided representations by attending to different parts of an input instance (e.g., regions, frames, words) using a multi-head self-attention module lin-iclr17; vaswani-nips17. We then combine each of such local representation with global representation via residual learning he-cvpr16 to avoid learning redundant information. Furthermore, to prevent the embeddings from collapsing into the mode (or the mean) of all the desirable embeddings, we regularize the locally-guided representations to be diverse. To our knowledge, we are the first to apply multi-head self-attention with residual learning for the application of instance embedding.
To address the partial association issue, we tie-up two PIE-Nets and train our model in the multiple-instance learning (MIL) framework dietterich-ai97. We call this approach Polysemous Visual-Semantic Embedding (PVSE). Our intuition is: when two instances are only partially associated, the learning constraint of Equation (1) will unnecessarily penalize embedding mismatches because it expects two instances to be perfectly associated. Capitalizing on our one-to-many instance embedding, our MIL objective relaxes the constraint of Equation (1) so that only one of embedding pairs is well-aligned, making our model more robust to partial cross-domain association. We illustrate this intuition in Figure 2. This relaxation, however, could cause a discrepancy between two embedding distributions because embedding pairs are left unconstrained. We thus regularize the learned embedding space by minimizing the discrepancy using the Maximum Mean Discrepancy (MMD) gretton-nips07, a popular technique for determining whether two sets of data are from the same probability distribution.
We demonstrate our approach on two cross-modal retrieval scenarios: image-text and video-text. For image-text retrieval, we evaluate on the MS-COCO dataset lin-eccv14; for video-text retrieval, we evaluate on the TGIF dataset li-cvpr16 as well as our new MRW (my reaction when) dataset, which we collected to promote further research in cross-modal video-text retrieval under ambiguity and partial association. The dataset contains 50K video-sentence pairs collected from social media, where the videos depict physical or emotional reactions to certain situations described in text. We compare our method with well-established baselines and carefully conduct an ablation study to justify various design choices. We report strong performance on all three datasets, and achieve the state-of-the-art result on image-to-text retrieval task on the MS-COCO dataset.
Related Work
Here we briefly review some of the most relevant work on instance embedding for cross-modal retrieval.
Correlation maximization: Most existing methods are based on one-to-one mapping of instances into a shared embedding space. One popular approach is maximizing correlation between related instances in the embedding space. Rasiwasia et al. rasiwasia-mm10 use canonical correlation analysis (CCA) to maximize correlation between images and text, while Gong et al. gong-ijcv14 extend CCA to a triplet scenario, e.g., images, tags, and their semantic concepts. Most recent methods incorporate deep neural networks to learn their embedding models in an end-to-end fashion. Andrew et al. andrew-icml13 propose deep CCA (DCCA), and Yan et al. yan-cvpr15 apply it to image-to-sentence and sentence-to-image retrieval.
Triplet ranking: Another popular approach is based on triplet ranking frome-nips13; kiros-arxiv14; wang-cvpr16; ye-eccv18, which encourages the distance between positive pairs (e.g., ground-truth pairs) to be closer than negative pairs (e.g., randomly selected pairs). Frome et al. frome-nips13 propose a deep visual-semantic embedding (DeViSE) model, using a hinge loss to implement triplet ranking. Faghri et al. faghri-bmvc17 extend this with the idea of hard negative mining, which focuses on maximum violating negative pairs, and report improved convergence rates.
Learning with auxiliary tasks: Several methods learn the embeddings in conjunction by solving auxiliary tasks, e.g., signal reconstruction feng-mm14; eisenschtat-cvpr17; tsai-iccv17, semantic concept categorization rasiwasia-mm10; huang-cvpr18, and minimizing the divergence between embedding distributions induced by different modalities tsai-iccv17; zhang-eccv18. Adversarial training goodfellow-nips14 is also used by many: Wang et al. wang-mm17 encourage the embeddings from different modalities to be indistinguishable using a domain discriminator, while Gu et al. gu-cvpr18 learn the embeddings with image-to-text and text-to-image synthesis tasks in the adversarial learning framework.
Attention-based embedding: All the above approaches are based on one-to-one mapping and thus could suffer from polysemous instances. To alleviate this, recent methods incorporate cross-attention mechanisms to selectively attend to local parts of an instance given the context of a conditioning instance from another modality huang-cvpr17; lee-eccv18, e.g., attend to different image regions given different text queries. Intuitively, this can resolve the issues with ambiguous instances and their partial associations because the same instance can be mapped to different points depending on the presence of the conditioning instance. However, such approach comes with computational overhead at inference time because each query instance needs to be encoded as many times as the number of references instances in the database; this severely limits its use in real-world applications. Different from previous approaches, our method is based on multi-head self-attention lin-iclr17; vaswani-nips17 which does not require a conditioning instance when encoding, and therefore each instance is encoded only once, significantly reducing computational overhead at inference time.
Beyond injective embedding: Similar to our motivation, some attempts have been made to go beyond the injective mapping. One approach is to design the embedding function to be stochastic and map an instance to a certain probability distribution (e.g., Gaussian) instead of a single point ren-mm16; mukherjee-emnlp16; oh-arxiv18. However, learning distributions is typically difficult/expensive and often lead to approximate solutions such as Monte Carlo sampling.
The work most similar to ours is by Ren et al. ren-bmvc17, where they compute multiple representations of an image by extracting local features using the region proposal method girshick-iccv15; text instances are still represented by a single embedding vector. Different from theirs, our method computes multiple and diverse representations from both modalities, where each representation is a combination of global context and locally-guided features, instead of just a local feature. Song et al. song-arxiv18, a prequel to this work, also compute multiple representations of each instance using multi-head self-attention. We extend their approach by combining global and locally-guided features via residual learning. We also extend the preliminary version of the MRW dataset with an increased number of sample pairs. Lastly, we report more comprehensive experimental results, adding results on the MS-COCO lin-eccv14 dataset for image-text cross-retrieval.
Approach
Our Polysemous Visual-Semantic Embedding (PVSE) model, shown in Figure 3, is composed of modality-specific feature extractors followed by two sub-networks with an identical architecture; we call the sub-network Polysemous Instance Embedding Network (PIE-Net). The two PIE-Nets are independent of each other and do not share the weights.
The PIE-Net takes as input a global context vector and multiple local feature vectors (Section 3.1), computes locally-guided features using the local feature transformer (Section 3.2), and outputs embeddings by combining the global context vector with locally-guided features (Section 3.3). We train the PVSE model in the Multiple Instance Learning (MIL) dietterich-ai97 framework. We explain how we make our model robust to ambiguous instances and partial cross-modal associations via our loss functions (Section 3.4) and finish with implementation details (Section 3.5).
2 Local Feature Transformer
The local feature transformer takes local features and transforms them into locally-guided representations . Our intuition is that different combinations of local information could yield diverse and refined representations of an instance. We implement this intuition by employing a multi-head self-attention module to obtain attention maps, prepare combinations of local features by attending to different parts of an instance, and apply non-linear transformations to obtain locally-guided representations.
3 Feature Fusion With Residual Learning
The fusion block combines global features and locally-guided features to obtain the final embedding output. We note that there is an inherent information overlap between the two features (both are derived from the same instance). To prevent from becoming redundant with and encourage it to learn only locally-specific information, we cast the feature fusion as a residual learning task. Specifically, we consider as input to the residual block and as residuals with its own parameters to optimize (). As shown in he-cvpr16, this residual mapping makes it easier to optimize the parameters associated with , helping us find meaningful locally-specific information; in the extreme case, if global features were the optimal, the residuals will be pushed to zero and the approach will fall back to the standard injective embedding.
4 Optimization and Inference
Given a dataset with instance pairs ( are either images or videos, are sentences), we optimize our PVSE model to minimize a learning objective:
where and are scalar weights that balance the influence of the loss terms. We describe each loss term below.
MIL Loss: We train our model in the Multiple Instance Learning (MIL) framework dietterich-ai97, designing a learning constraint for the cross-modal retrieval scenario:
where and are the PIE-Net embeddings of and , respectively, and . We use the cosine distance as our distance metric, ).
Making an analogy to the MIL for binary classification amores-ai13, the left side of the constraint is the “positive” bag where at least one of embedding pairs is assumed to be positive (match), while the right side is the “negative” bag containing only negative (mismatch) pairs. Optimizing under this constraint allows our model to be robust to partial cross-modal association because it can ignore mismatching embedding pairs of partially associated instances.
We implement the above constraint by designing our MIL loss function to be:
where is a margin parameter. Notice that we have the min operator for , similar to ren-bmvc17; this can be seen as a form of hard negative mining, which we found to be effective and accelerate the convergence.
Diversity Loss: To ensure that our PIE-Net produces diverse representations of an instance, we design a diversity loss that penalizes the redundancy among locally-guided features. To measure the redundancy, we compute a Gram matrix of (and of ) that encodes the correlations between all combinations of locally-guided features, i.e., . We normalize each prior to the computation so that they are on an ball.
The diagonal entries in are always one (they are on a unit ball); the off-diagonals are zero iff two locally-guided features are orthogonal to each other. Therefore, the sum of off-diagonal entries in indicates the redundancy among locally-guided features. Based on this, we define our diversity loss as:
Note that we do not compute the diversity loss on the final embedding representations and because they already have global information baked in, making the orthogonality constraint invalid. This also ensures that the loss gets back-propagated through appropriate parts in the computational graph, and does not affect the global feature encoders, i.e., the FC layer for the image encoder, and the bi-GRUs for the video and sentence encoders.
Domain Discrepancy Loss: Optimizing our model under the MIL loss has one drawback: two distributions induced by and , which we denote by and , respectively, may diverge quickly because we only consider the minimum distance pair, , in loss computation and let the other pairs left to be unconstrained. It is therefore necessary to regularize the discrepancy between the two distributions.
One popular way to measure the discrepancy between two probability distributions is the Maximum Mean Discrepancy (MMD) gretton-nips07. The MMD between two distributions and over a function space is
where the summation in each term is taken over all pairs of embeddings . We use a radial basis function (RBF) kernel as our kernel function.
Inference: At test time, we assume a database of instances (e.g., videos) and their embedding vectors. Given a query instance (e.g., a sentence), we compute embedding vectors and find the best matching instance in the database by comparing the cosine distances between all combinations of embeddings.
5 Implementation Details
We subsample frames at 8 FPS and store them in a binary storage format. https://github.com/TwentyBN/GulpIO We set the maximum length of video to be 8 frames; for videos longer than 8 frames we select random subsequences during training, while during inference we sample 8 frames evenly spread across each video. We do not limit the sentence length as it has a minimal effect on the GPU memory footprint. We cross-validate the optimal hyper-parameter settings, varying , . We use the AMSGRAD optimizer reddi-iclr18 with an initial learning rate of 2e-4 and reduce it by half when the loss stagnates. We train our model end-to-end, except for the pretrained CNN weights, for 50 epochs with a batch of 128 samples. We then finetune the whole model (including the CNN weights) for another 50 epochs.
MRW Dataset
To promote future research in video-text cross-modal retrieval, especially with ambiguous instances and their partial cross-domain association, we release a new dataset of 50K video-sentence pairs collected from social media; we call our dataset MRW (my reaction when).
Table 1 provides descriptive statistics of several video-sentence datasets. Most existing datasets are designed for video captioning rohrbach-ijcv17; xu-cvpr16; li-cvpr16, with sentences providing textual descriptions of visual content in videos (video text relationship). Our dataset is unique in that it provides videos that display physical or emotional reactions to the given sentences (text video relationship); these are called reaction GIFs. According to a subreddit r/reactiongif https://www.reddit.com/r/reactiongifs:
A reaction GIF is a physical or emotional response that is captured in an animated GIF which you can link in response to someone or something on the Internet. The reaction must not be in response to something that happens within the GIF, or it is considered a “scene”.
This definition clearly differentiates ours from existing datasets: There is an inherently weaker association of concepts between video and text; see Figure 4. This introduces several additional challenges to cross-modal retrieval, part of which are the focus of this work, i.e., dealing with ambiguous instances and partial cross-domain association. We provide detailed data analyses and compare it with existing video captioning datasets in the supplementary material.
Experiments
We evaluate our approach on image-text and video-text cross-modal retrieval scenarios. For image-text cross-retrieval, we evaluate on the MS-COCO dataset lin-eccv14; for video-text we use the TGIF li-cvpr16 and our MRW datasets.
For MS-COCO we use the data split of kiros-arxiv14, which provides 113,287 training, 5K validation and 5K test samples; each image comes with 5 captions. We report results on both 1K unique test images (averaged over 5 folds) and the full 5K test images. For TGIF we use the original data split li-cvpr16 with 80K training, 10,708 validation and 34,101 test samples; since most test videos come with 3 captions, we report results on 11,360 unique test videos. For MRW, we use a data split of 44,107 training, 1K validation and 5K test samples; all the videos come with one caption.
Following the convention in cross-modal retrieval, we report results using Recall@ (R@) at , which measures the the fraction of queries for which the correct item is retrieved among the top results. We also report the median rank (Med R) of the closest ground truth result in the list, as well as the normalized median rank (nMR) that divides the median rank by the number of total items. For cross-validation, we select the best model that achieves the highest in both directions (visual-to-text and text-to-visual) on a validation set.
While we report quantitative results in the main paper, our supplementary material contains qualitative results with visualizations of multi-head self-attention maps.
Table 2 shows the results on MS-COCO. To facilitate comprehensive comparisons, we provide previously reported results on this dataset. We omit results from cross-attention models huang-cvpr17; lee-eccv18 that require a pair of instances (e.g., image and text) when encoding each instance. Our approach outperforms most of the baselines, and achieves the new state-of-the-art on the image-to-text task on the 5K test set. We note that both GXN gu-cvpr18 and SCO huang-cvpr18 are trained with multiple objectives; in addition to solving the ranking task, GXN performs image-text cross-modal synthesis as part of training, while SCO performs classification of semantic concepts and their orders as part of training. Compared to the two methods, our model is trained with a single objective (ranking) and thus could be considered as a simpler model.
The most direct comparison to ours would be with VSE++ faghri-bmvc17. Both our model and VSE++ share the same image and sentence encoders. When we let our PIE-Net to produce single embeddings for input instances (K=1), the only difference becomes that VSE++ directly uses our global features as their embedding representations, while we use the output from our PIE-Nets. The performance gap between ours (K=1) and VSE++ shows the effectiveness of our PIE-Net, which combines global context with locally-guided features produced by our local feature transformer.
2 Video-Text Retrieval Results
Table 3 and Table 4 show the results on TGIF and MRW datasets. Because there is no previously reported results on these datasets for the cross-model retrieval scenario, we run the baseline models and report their results. We can see that our method show strong performance compared to all the baselines. We provide implementation details of the baseline models in the supplementary material.
We notice is that the overall performance is much lower than the results from MS-COCO. This shows how challenging video-text retrieval is (and video understanding in a broader context), and calls for further research in this task. We can also see that there is a large performance gap between the two datasets. This suggests the two datasets have significantly different characteristics: the TGIF contains sentences describing visual content in videos, while our MRW dataset contains videos showing one of possible reactions to certain situations described in sentences. This makes the association between video and text modalities much weaker for the MRW than for the TGIF.
3 Ablation Results
The number of embeddings : Tables 2, 3, 4 show that computing multiple embeddings per instance improves performance compared to just a single embedding (see the last two rows in each table). To better understand the effect of , we vary it from 1 to 8, and also compare with , a baseline where we bypass our Local Feature Transformer and simply use the global feature as the final embedding representation. Figure 5 shows the performance on all three datasets based on the rsum metric (R@1 + R@5 + R@10 for image/video-to-text and back). The results are from the models before fine-tuning the ResNet-152 weights. We can see that there is a significant improvement from to ; this shows the effectiveness of our Local Feature Transformer. We can make an interesting observation by comparing the optimal settings across different datasets: for COCO and TGIF, and for MRW. While this cannot be used as strong evidence, we believe this shows the level of ambiguity is higher on MRW than the other two datasets.
Global vs. locally-guided features: We analyze the importance of global and locally-guided features, as well as different strategies to combine them. Figure 6 shows results on several ablative settings: No Global is when we use locally-guided features alone (discard global features); No Residual is when we simply concatenate global and locally-guided features, instead of combining them via residual learning. We report results on both MS-COCO and MRW because the two datasets exhibit the biggest difference in the level of ambiguity.
We notice that the performance drops significantly on both datasets when we discard global features. Together with results in Figure 5 (discard locally-guided features), this shows the importance of balancing global and local information in the final embedding. We also see that simply concatenating the two features (no residual learning) hurts the performance, and the drop is more significant on the MRW dataset. This suggests our residual learning setup is especially crucial for highly ambiguous data.
MIL objective: Figure 6 also shows the result of No MIL, which is when we concatenate the embeddings and optimize the standard triplet ranking objective frome-nips13; kiros-arxiv14; faghri-bmvc17, i.e., the “Conventional” setup in Figure 2. While the differences are relatively smaller than with the other ablative settings, there are statistically significant differences between the two results on both datasets ( on MS-COCO and on MRW). We also see that the difference between No MIL and Ours on MRW is more pronounced than on MS-COCO. This suggests thed MIL objective is especially effective for highly ambiguous data.
Sensitivity analysis on different loss weights: Figure 7 shows the sensitivity of our approach when we vary the relative loss weights, i.e., and in Equation (5). Note that the weights are relative, not absolute, e.g., instead of directly multiplying to , we first scale it to and then multiply it to . The results show that both loss terms are important in our model. We can see, in particular, that plays an important role in our model. Without it, the two embedding spaces induced by different modalities may diverge quickly due to the MIL objective, which may result in a poor convergence rate. Overall, our results suggests that the model is not much sensitive to the two relative weight terms.
Conclusion
Ambiguous instances and their partial associations pose significant challenges to cross-modal retrieval. Unlike the traditional approaches that use injective embedding to compute a single representation per instance, we propose a Polysemous Instance Embedding Network (PIE-Net) that computes multiple and diverse representations per instance. To obtain visual-semantic embedding that is robust to partial cross-modal association, we tie-up two PIE-Nets, one per modality, and jointly train them using the Multiple Instance Learning objective. We demonstrate our approach on the image-text and video-text cross-modal retrieval scenarios and report strong results compared to several baselines.
Part of our contribution is also in the newly collected MRW dataset. Unlike existing video-sentence datasets that contain sentences describing visual content in videos, ours contain videos illustrating one of possible reactions to certain situations described in sentences, which makes video-sentence association somewhat ambiguous. This poses new challenges to cross-modal retrieval; we hope there will be further progress on this challenging new dataset.
Appendix A MRW Dataset
Our dataset consists of 50,107 video-sentence pairs collected from popular social media websites including reddit, Imgur, and Tumblr. We crawled the data using the GIPHY API https://developers.giphy.com with query terms mrw, mfw, hifw, reaction, and reactiongif; we crawled the data from August 2016 to March 2019. Table 5 shows the descriptive statistics of our dataset. We are continuously crawling the data, and plan to release updated versions in the future.
Note that most of the videos in our dataset have the animated GIF format. Technically speaking, animated GIFs and videos have different formats; the former is lossless, palette-based, and has no audio. In this paper, however, we use the two terms interchangeably because the distinction is unnecessary in our method. Below, to provide the context for our work, we briefly review previous work that focused on animated GIF.
There is increasing interest in conducting research around animated GIFs. Bakhshi et al. bakhshi-chi16 studied what makes animated GIFs engaging on social networks and identified a number of factors that contribute to it: the animation, lack of sound, immediacy of consumption, low bandwidth and minimal time demands, the storytelling capabilities and utility for expressing emotions. Previous work in the computer vision and multimedia communities used animated GIFs for various tasks in video understanding. Jou et al. jou-mm14 propose a method to predict viewer perceived emotions for animated GIFs. Gygli et al. gygli-cvpr16 propose the Video2GIF dataset for video highlighting, and further extended it to emotion recognition gygli-mm16. Chen et al. chen-acii17 propose the GIFGIF+ dataset for emotion recognition. Zhou et al. zhou-wacv18 propose the Image2GIF dataset for video prediction, along with a method to generate cinemagraphs from a single image by predicting future frames.
Recent work use animated GIFs to tackle the vision & language problems. Li et al. li-cvpr16 propose the TGIF dataset for video captioning; Jang et al. jang-cvpr17 propose the TGIF-QA dataset for video visual question answering. Similar to the TGIF dataset li-cvpr16, our dataset includes video-sentence pairs. However, our sentences are created by real users from Internet communities rather than study participants, thus posing real-world challenges. More importantly, our dataset has implicit concept association between videos and sentences (videos contain physical or emotional reactions to sentences), while the TGIF dataset has explicit concept association (sentences describe visual content in videos).
A.2 Analysis of Facial Expressions
Facial expression plays an important role in our dataset: 6,380 samples contain the hashtag MFW (my face when), indicating that those GIFs contain emotional reactions manifested by facial expressions. To better understand the landscape of our dataset, we analyze the types of facial expressions contained in our dataset by leverage automatic tools.
First, we count the number of faces appearing in the animated GIFs. To do this, we applied the dlib CNN face detector dlib09 on five frames sampled from each animated GIF with an equal interval. The results show that there are, on average, faces in a given frame of an animated GIF. Also, 34,052 animated GIFs contain at least one face. This means that 72% of our videos contain faces, which is quite significant. This suggests that employing techniques tailored specifically for face understanding could potentially improve performance on our dataset.
Next, we use the Affectiva Affdex McDuff2016Affdex to analyze facial expressions depicted in the animated GIFs, detecting the intensity of expressions from two frames per second in each animated GIF. We looked at six expressions of basic emotions ekman1992argument, namely, joy, fear, sadness, disgust, surprise and anger. We analyzed only the frames that contain a face with its bounding box region larger than 15% of the image. Figure 8 shows the results. Overall, joy with average intensity of 9.1% and disgust (7.4%) are the most common facial expressions in our dataset.
A.3 Comparison to the TGIF Dataset
Image and video captioning often involves describing objects and actions depicted explicitly in visual content lin-eccv14; li-cvpr16. For reaction GIFs, however, visual-textual association is not always explicit. For example, as is the case in our dataset, objects and actions depicted in visual content might be a physical or emotional reaction to the scenario posed in the sentence.
In this section, we qualitatively compare our dataset with the TGIF dataset li-cvpr16, which contains 120K video-sentence pairs for video captioning. We chose the dataset because both datasets contain animated GIFs collected from social media, and thus contain similar visual content.
We first compare words appearing in both datasets. Figure 9 shows word clouds of nouns and verbs extracted from our MRW dataset and the TGIF dataset li-cvpr16. Sentences in the TGIF dataset are constructed by crowdworkers to describe the visual content explicitly displayed in animated GIFs. Therefore, its nouns and verbs mainly describe physical objects, people and actions that can be visualized, e.g., cat, shirt, stand, dance. In contrast, MRW sentences are constructed by the Internet users, typically from subcommunities in social networks that focus on reaction GIFs. As can be seen from Figure 9, verbs and nouns in our MRW dataset additionally include abstract terms that cannot necessarily be visualized, e.g., time, day, realize, think. This shows that our dataset contains ambiguous terms and their associations, which pose significant challenges to cross-modal retrieval.
Next, we compare whether video-sentence associations are explicit/implicit in both datasets. To this end, we conducted a user study in which we asked six participants to verify the association between sentences and animated GIFs. We randomly sampled 100 animated GIFs from the test sets of both our dataset and TGIF dataset li-cvpr16. We paired each animated GIF with both its associated sentence and a randomly selected sentence from the corresponding dataset, resulting in 200 GIF-sentence pairs per dataset.
The results show that, in case of our dataset (MRW), 80.4% of the associated pairs are positively marked as being relevant, suggesting humans are able to distinguish the true vs. fake pairs despite implicit concept association. On the other hand, 50.7% of the randomly assigned sentences are also marked as matching sentences. The high false positive rate shows the ambiguous nature of GIF-sentence association in our dataset.
In contrast, for the TGIF dataset with clear explicit association, 95.2% of the positive pairs are correctly marked as relevant and only 2.6% of the irrelevant pairs are marked as being relevant. This human baseline demonstrates the challenging nature of GIF-sentence association in our dataset, due to their implicit rather than explicit association.
A.4 Application: Animated GIF Search
Animated GIFs are becoming increasingly popular bakhshi-chi16; more people use them to tell stories, summarize events, express emotion, and enhance (or even replace) text-based communication. To reflect this trend, several social networks and messaging apps have recently incorporated GIF-related features into their systems, e.g., Facebook users can create posts and leave comments using GIFs, Instagram and Snapchat users can put “GIF stickers” into their personal videos, and Slack users can send messages using GIFs. This rapid increase in popularity and real-world demand necessitates more advanced and specialized systems for animated GIF search.
Current solutions to animated GIF search rely entirely on concept tags associated with animated GIFs and matching them with user queries. The tags are typically provided by users or produced by editors at companies like GIPHY. In the former case, noise becomes an issue; in the latter, it is expensive and would not scale well.
One of the motivations behind collecting our MRW dataset is to build a text-based animated GIF search engine, targeted for real-world scenarios mentioned above. Existing video captioning datasets, such as TGIF li-cvpr16, are inappropriate for our purpose because of the explicit nature of visual-textual association, i.e., sentences simply describe what is being shown in videos. Rather, we need a dataset that captures various types of nuances used in social media, e.g., humor, irony, satire, sarcasm, incongruity, etc. Because our dataset provides video-text pairs with implicit visual-textual association, we believe that it has the potential to provide training data for building text-based animated GIF search engines targeted for social media.
To demonstrate the potential, we provide qualitative results on text-to-video retrieval using our dataset, shown in Figure 12. Each set of results show a query text and the top five retrieved videos, along with their ranks and cosine similarity scores. We would like the readers to take a close look at each set of results and decide which of the five retrieved videos depict the most likely visual response to the query sentence. The answers are provided below. For better viewing experience, we provide an HTML page with animated GIFs instead of static images. We strongly encourage the readers to check the HTML page to better appreciate the results. (Answers: 3, 5, 2, 4, 1, 5, 4)
Appendix B Baseline Implementation Details
In the experiment section, we provided baseline results for MS-COCO, TGIF, and MRW datasets. For MS-COCO, we provided previously reported results. For TGIF and MRW, on the other hand, we reported our own results because there has not been previous results on the datasets. Due to the space limit, we omitted implementation details of the baseline approaches; here we provide implementation details of the four baseline approaches: DeViSE frome-nips13, VSE++ faghri-bmvc17, Order Embedding vendrov-iclr16, and Corr-AE feng-mm14.
For fair comparison, all four baselines share the same video and sentence encoders as described in Section 3.1 of the main paper. The only difference is in the loss function we train the models with. Following the notation used in the main paper, we denote the output of the video and sentence encoders by and , respectively. We employ the following loss functions for the baselines:
DeViSE frome-nips13: We implement the conventional hinge loss in the triplet ranking setup; see Equation (9). It penalizes the cases when the distance between positive pairs (i.e., the ground truth) is further away than negative pairs (e.g., randomly sampled) with a margin parameter (we measure the cosine distance).
VSE++ faghri-bmvc17: We implement the hard negative mining version of the conventional hinge loss triplet ranking loss; see Equation (10). We have experimented with the original version and found that it fails to find a suitable solution to the objective, producing retrieval results that are almost identical to random guess. We suspect that the high noise present in both TGIF and MRW datasets makes the max function too strict as a constraint. We therefore replace the function with a “filter” function that includes only highly-violating cases while ignoring others.
Intuitively, we implement the filter function to be an outlier detection function based on z-scores, where any z-score greater than 3 or less than -3 is considered to be an outlier. Specifically, we compute the z-scores for all of possible combinations inside Equation (10) and discard instances if their absolute z-score is below 3.0. This way, we are considering multiple hard negatives instead of just one. We have empirically found this modification to be crucial to achieve reasonable performances on the TGIF and MRW datasets.
Order Embedding vendrov-iclr16: We used the original implementation provided by the authors of vendrov-iclr16.
Corr-AE feng-mm14: We implement the correspondence cross-modal autoencoder proposed by Feng et al. feng-mm14 (see Figure 4 in feng-mm14). Given the encoder output and , we build two autoencoders, one per modality, so that each autoencoder can reconstruct both and . The autoencoders have four fully-connected layers with hidden units, respectively. Each of the fully connected layers is followed by a ReLU activation and a layer normalization ba-arxiv16.
Appendix C Visualization of Multi-Head Self-Attention
Figure 10 shows examples of visual-textual attention maps on the MS-COCO dataset; the task is image-to-text retrieval. The first column shows query images with ground-truth sentences. Each of the other three columns shows visual (spatial) attention maps and their top-ranked text retrieval results, as well as their ranks and cosine similarity scores (green: correct, red: incorrect). We color-code words in the retrieved sentences according to their textual attention intensity values, normalized between .
A glimpse at the results in each row shows that the three attention maps attend to different regions of the query image. Looking closely, we notice that salient regions are typically attended by multiple attention maps. For example, all three attention maps in Figure 10 highlight: (a) the photographer, (b) the bench, (c) the fruit stand, (e) the pink flowers, (f) the stop sign, (h) the woman, (j) the fire hydrant. However, this is not always the case: In Figure 10 (i), none of the attention maps highlights the most salient object, the black dog, and each attention map highlights different regions in the image. Even though all three attention maps do not “attend to” the dog, their top-ranked text retrieval results are still highly relevant to the query image; all three retrieved sentences have the word dog in them. This is possible because our PIE-Net computes embedding vectors by combining global context with locally-guided features. In this example, the global context provides information about the black dog, while each of the three locally-guided features contains region-specific information, specifically, (first map): the book shelf, (second map): the floor, (third map): the brown cushion.
The most interesting observation is that there are subtle variations in the retrieved sentences depending on where the visual attention is focused on. For example, in Figure 10 (a), the first result focuses on the photographer as a whole, the second focuses on the tiny camera (the visual attention is more narrowly focused on the photographer), and the third focuses on the pizza on the table (notice the visual attention on the table). In Figure 10 (d), the first result focuses on the ship, the second focuses on the building, and the third on an (imaginary) bird that could have been flying over the buildings. In Figure 10 (g), the first result focuses on the boat and the muddy water (notice visual attention on the muddy water region at the lower left corner), while the second focuses on the table of people (notice visual attention on the table region). In Figure 10 (j), the first results focuses on the fire hydrant and the yellow wall that is right behind the hydrant, while the second focuses on the hydrant as well as the building with two windows (notice now the visual attention is more widely spread out than the first result). We encourage the readers to look closely at Figure 10 to appreciate the subtle variations in the retrieved sentences depending on their corresponding visual attention.
C.2 Video-to-Text Retrieval Results on TGIF
Figure 11 shows examples of visual-textual attention maps on the TGIF dataset; the task is video-to-text retrieval. In each set of results, we show: (top) a query video and its ground-truth sentence, (bottom three rows): three visual (temporal) attention maps and their top-ranked text retrieval results, as well as their ranks and cosine similarity scores (green: correct, red: incorrect). We color-code words in the retrieval results according to their textual attention intensity values, normalized between .
Similar to the results on MS-COCO, here we see that visual and textual attention maps tend to highlight salient video frames and words, respectively. Looking closely, we notice that the retrieved results tend to capture the concepts highlighted by their corresponding visual attention. For example, in Figure 11 (a), the top ranked result contain “lady dressed in black” and ”drinking a glass of wine”, and the visual attention highlights both the early part of the video, where a woman is drinking from a bottle of whisky, and the latter part, where her black dress is shown. For the second ranked result, the visual attention no longer highlights the latter part, and the retrieved text focuses solely on drinking action (no mention of her black dress). In Figure 11 (b), the top ranked result focuses on scoring a goal, while the second rank result also focus on the guy being hit in the face with the ball. Notice the difference of visual attention maps between the first and the second case.
C.3 Text-to-Video Retrieval Results on MRW
Figure 12 shows examples of text-to-video retrieval results on the MRW dataset. In each row, we show a query sentence and top five retrieved videos along with their ranks and cosine similarity scores. Unlike the previous two figures, here we do not directly show the ground-truth matches (but rather ask the readers to find them; we provide the answers above). The purpose of this is to emphasize the ambiguous and implicit nature of visual-textual association present in our dataset.
Most of the top five retrieved videos seem to be a good match to the query sentence. For example, Figure 12 (a) shows five videos that all contain a human face, each expressing subtly different emotions. Figure 12 (b) shows five videos that all contain an animal (squirrel, cat, etc), and most videos contain food. All five retrieved videos in Figure 12 show some form of awkward (dancing) moves.
We believe that the relatively poor retrieval performance reported in our main paper is partly explained by our qualitative results: visual-textual associations are highly ambiguous and there could be multiple correct matches. This calls for a different metric that measures the perceptual similarity between queries and retrieved results, rather than exact match. There has been some progress on perceptual metrics in the image synthesis literature (e.g., Inception Score salimans-nips16). We are not aware of a suitable perceptual metric for cross-modal retrieval, and this could be a promising direction for future research.