Similarity Reasoning and Filtration for Image-Text Matching

Haiwen Diao, Ying Zhang, Lin Ma, Huchuan Lu

Introduction

Image-text matching refers to measuring the visual-semantic similarity between image and text, which is becoming increasingly significant for various vision-and-language tasks, such as cross-modal retrieval (Wang et al. 2020), image captioning (Anderson et al. 2018), text-to-image synthesis (Xu et al. 2018), and multimodal neural machine translation (Toyama et al. 2017). Although great progress has been made in recent years, image-text matching remains a challenging problem due to complex matching patterns and large semantic discrepancies between image and text.

To accurately establish the association between the visual and textual observations, a large proportion of methods (Liu et al. 2017; Nam, Ha, and Kim 2017; Lee et al. 2018; Song and Soleymani 2019; Wang et al. 2019c; Li et al. 2019; Wang et al. 2020) utilize deep neural networks to firstly encode image and text into compact representations, and then learn to measure their similarity under the guidance of a matching criterion. For example, Wang et al. (Wang, Li, and Lazebnik 2016) and Faghri et al. (Faghri et al. 2017) map the whole image and the full sentence into a common vector space, and compute the cosine similarity between the global representations. To improve the discriminative ability of the unified embeddings, many strategies such as semantic concept learning (Huang et al. 2018; Shi et al. 2019) and region relationship reasoning (Li et al. 2019) are developed to enhance visual features by incorporating local region semantics. However, these approaches fail to capture the local interactions between image regions and sentence fragments, leading to limited interpretability and performance gains. To address this problem, Karpathy et al. (Karpathy and Li 2015) and Lee et al. (Lee et al. 2018) propose to discover all the possible alignments between image regions and sentence fragments, which produce impressive retrieval results and inspire a surge of works (Wang et al. 2019c; Hu et al. 2019; Zhang et al. 2020; Chen et al. 2020; Wehrmann, Kolling, and Barros 2020) to explore more accurate fine-grained correspondence. Although noticeable improvements have been made by designing various mechanisms to encode more powerful features or capture more accurate alignments, these approaches neglect the importance of similarity computation, which is the key to explore the complex matching patterns between image and text.

To be more specific, there are three defects in previous approaches. Firstly, these methods compute scalar-based cosine similarities between local features, which may not be powerful enough to characterize the association patterns between regions and words. Secondly, most of them aggregate all the latent alignments between regions and words simply with max pooling (Karpathy and Li 2015) or average pooling (Lee et al. 2018; Chen et al. 2020), which hinders the information communication between local and global alignments, and thirdly, fails to consider the distractions of less-meaningful alignments, such as the alignments built with "a" and "in", as shown in Figure 1.

To address these problems, in this paper we propose a novel Similarity Graph Reasoning and Attention Filtration (SGRAF) network for image-text matching. Specifically, we start with capturing the global alignments between the whole image and the full sentence, as well as the local alignments between image regions and sentence fragments. Instead of characterizing these alignments with scalar-based cosine similarity, we propose to learn the vector-based similarity representations to model the cross-modal associations more effectively. Then we introduce the Similarity Graph Reasoning (SGR) module, which relies on a Graph Convolution Neural Network (GCNN) to reason more accurate image-text similarity via capturing the relationship between local and global alignments. Furthermore, we develop the Similarity Attention Filtration (SAF) module to aggregate all the alignments attended by different significance scores, which reduces the interferences of non-meaningful alignments and achieves more accurate cross-modal matching results. Our main contributions are summarized as follows:

We propose to learn the vector-based similarity representations for image-text matching, which enables greater capacity in characterizing the global alignments between images and sentences, as well as the local alignments between regions and words.

We propose the Similarity Graph Reasoning (SGR) module to infer the image-text similarity with graph reasoning, which can identify more complex matching patterns and achieve more accurate predictions via capturing the relationship between local and global alignments.

We attempt to consider the interferences of non-meaningful words in similarity aggregation, and propose an effective Similarity Attention Filtration (SAF) module to suppress the irrelevant interactions for further improving the matching accuracy.

Related Work

Feature Encoding Many prior Approaches (Karpathy and Li 2015; Song and Soleymani 2019; Liu et al. 2017; Nam, Ha, and Kim 2017; Lee et al. 2018; Wang et al. 2019c; Li et al. 2019; Wang et al. 2020) focused on feature extraction and optimization for cross-modal retrieval. For textual features, Frome et al. (Frome et al. 2013) employed Skip-Gram (Mikolov et al. 2013) to extract word representations. Klein et al. (Klein et al. 2015) explored Fisher Vectors (FV) (Perronnin and Dance 2007) for text representation. Kiros et al. (Kiros, Salakhutdinov, and Zemel 2014) adopted a GRU as the text encoder. For visual features, Liu et al. (Liu et al. 2017) adapted Recurrent Residual networks to refine global embeddings. (Song and Soleymani 2019; Wei et al. 2020) employed multi-head self-attention to combine global context with locally-guided features. Besides, Some works (Nam, Ha, and Kim 2017; Ji et al. 2019) exploited block-based visual attention to gather semantics on feature maps, while (Lee et al. 2018; Wang et al. 2019c, b; Li et al. 2019; Wang et al. 2020; Chen and Luo 2020) followed (Anderson et al. 2018) to obtain region-based features of visual objects with the pre-trained model on Visual Genomes (Krishna et al. 2017). Especially, (Chen and Luo 2020) explored Bi-GRU to gain high-level object features, while (Li et al. 2019; Wang et al. 2020) proposed GCN-based networks to generate relationship-enhanced object features. We employ self-attention (Vaswani et al. 2017) on region or word features to get image or text representation. We concentrate on the similarity encoding mechanism that models global image-text and local region-word alignments comprehensively and fully encodes fine-grained relations between image and text.

Similarity Prediction Most existing works (Faghri et al. 2017; Wang, Li, and Lazebnik 2016; Zheng et al. 2017; Vendrov et al. 2016; Gu et al. 2018) for image-text matching learned the joint embedding and the similarity measures for cross-modal matching. For global alignments, some works (Faghri et al. 2017; Wang, Li, and Lazebnik 2016; Liu et al. 2017; Song and Soleymani 2019; Nam, Ha, and Kim 2017; Li et al. 2019) explored a joint space and calculated the inner product (e.g. cosine distance) for similarity computation. Others (Vendrov et al. 2016; Gu et al. 2018) introduced an ordered representations to measure antisymmetric visual-semantic hierarchy. For local alignments, most networks (Karpathy and Li 2015; Lee et al. 2018; Hu et al. 2019; Wang et al. 2019b; Chen et al. 2020) computed scalar-based alignments and adopted simple operation (e.g. sum and average) to fuse local alignments. For example, Lee et al. (Lee et al. 2018) studied the latent semantic alignments among region-words pairs and integrated local cosine alignments by average or LogSumExp. Differently, our network aggregates similarities by exploring global-local relationships among vector-based alignments and reducing the distraction from less-meaningful ones.

Graph Convolution Network

The researches based on Graph modeled the dependencies between concepts and facilitated graph reasoning such as GCNN (Duvenaud et al. 2015; Kipf and Welling 2017), and Gated Graph Neural Network (GGNN) (Li et al. 2016). These graph neural networks have been widely employed in various visual semantic tasks, such as image captioning (Yang et al. 2019), VQA (Teney, Liu, and van den Hengel 2017), and grounding referring expressions (Wang et al. 2019a). In recent years, there are several approaches to utilize graph structures to enhance single visual or textual features referring to image-text matching. Shi et al. (Shi et al. 2019) adopted Scene Concept Graph (SCG) by using image scene graphs and frequently co-occurred concept pairs as scene common-sense knowledge. Li et al. (Li et al. 2019) proposed Visual Semantic Reasoning to build up connections between image regions and generate visual representations with semantic relationships. Wang et al. (Wang et al. 2020) employed visual scene graph and textual scene graph, each of which separately refines visual and textual features including objects and relationships. They all focus on ”feature encoding” by learning single-modality contextualized representations, while our SGR targets at ”similarity reasoning” and explores more complex matching patterns with global and local cross-modal alignments.

Attention Mechanism

The attention mechanism has been applied to adaptively filter and aggregate information in natural language processing. When it comes to image-text matching, it has been intended to attend to certain parts of visual and textual data. (Lee et al. 2018; Wang et al. 2019b) developed Stacked Cross Attention to match latent alignments using both image regions and textual words as context. (Liu et al. 2019; Hu et al. 2019; Wang et al. 2019c) designed more complicated Cross Attentions to improve image-text matching. Chen et al. (Chen et al. 2020) proposed an Iterative Matching with Recurrent Attention Memory to explore fine-grained region-word correspondence progressively. We adopt textual-to-visual attention (Lee et al. 2018) with region-word pairs and calculate textual-attended alignments. In this paper, our SAF aims to discard less-semantic alignments instead of exploiting precise cross-modal attention.

Method

In this section, we focus on improving the visual-semantic similarity learning via capturing the relationship between local and global alignments, and suppressing the disturbance of less-meaningful alignments. As illustrated in Figure 2, we begin with introducing how to encode the visual and textual observations, and then compute the similarity representations of all local and global representation pairs. Afterwards, we elaborate on the proposed Similarity Graph Reasoning (SGR) module for relation-aware similarity reasoning and Similarity Attention Filtration (SAF) module for representative similarity aggregation. Finally, we present the detailed implementations of training objectives and inference strategies with both the SGR and SAF modules.

Textual Representations.

Similarity Representation Learning

Global Similarity Representation.

We compute the similarity representation between the global image feature vˉ\bar{\boldsymbol{v}} and sentence features tˉ\bar{\boldsymbol{t}} with Eq. (1),

Local Similarity Representation.

To exploit local similarity representations between local features of visual and textual observations, we apply textual-to-visual attention (Lee et al. 2018) to attend on each region with respect to each word. Attention weight for each region is computed by

Here the weight αij{\alpha}_{ij} is calculated by the softmax function with a temperature parameter λ\lambda. cijc_{ij} indicates the cosine similarity between region feature vi\boldsymbol{v}_{i} and word feature tj\boldsymbol{t}_{j}, c^ij=[cij]+/∑j=1L[cij]+2{\hat{c}_{ij}}={\left[c_{ij}\right]}_{+}/\sqrt{{\sum}_{j=1}^{L}{\left[c_{ij}\right]}_{+}^{2}} aims to normalize the cosine similarity matrix, and [x]+=max(x,0)\left[x\right]_{+}=max(x,0).

Then we generate the attended visual features ajv\boldsymbol{a}_{j}^{v} with respect to jj-th word by

and finally we compute the local similarity representation between ajv\boldsymbol{a}_{j}^{v} and tj\boldsymbol{t}_{j} as

Similarity Graph Reasoning

To achieve more comprehensive similarity reasoning, we build a similarity graph to propagate similarity messages among the possible alignments at both local and global levels. More specifically, we take all the word-attended similarity representations and the global similarity representation as graph nodes, i.e. N={s1l,....,sLl,sg}\mathcal{N}=\{\boldsymbol{s}_{1}^{l},....,\boldsymbol{s}_{L}^{l},\boldsymbol{s}^{g}\}, and follow (Kuang et al. 2019) to compute the edge from node sq∈N\boldsymbol{s}_{q}\in\mathcal{N} to sp∈N\boldsymbol{s}_{p}\in\mathcal{N} as

Graph Reasoning.

With the constructed graph nodes and edges, we perform similarity graph reasoning by updating the nodes and edges with

with sp0\boldsymbol{s}_{p}^{0} and sq0\boldsymbol{s}_{q}^{0} taken from N\mathcal{N} at step n=0n=0, and Wrn\boldsymbol{W}_{r}^{n}, Winn\boldsymbol{W}_{in}^{n}, Woutn\boldsymbol{W}_{out}^{n} are learnable parameters in each step. After current step of graph reasoning, the node spn\boldsymbol{s}_{p}^{n} is replaced with spn+1\boldsymbol{s}_{p}^{n+1}.

We iteratively reason the similarity for NN steps, and take the output of the global node at the last step as the reasoned similarity representation, and then feed it into a fully-connect layer to infer the final similarity score. The SGR module enables the information propagation between local and global alignments, which can capture more comprehensive interactions to facilitate the similarity prediction.

Similarity Attention Filtration

Although the exploitation of local alignments can boost the matching performance via discovering more fine-grained correspondence between image regions and sentence fragments, we notice that the less-meaningful alignments hinder the distinguishing ability when aggregating all the possible alignments in an undifferentiated way. Therefore we propose a Similarity Attention Filtration (SAF) module to enhance important alignments, as well as suppress ineffectual alignments, such as the alignments with "the", "be" and etc.

Given the local and global similarity representations, we calculate an aggregation weight βp{\beta}_{p} for each similarity representation sp∈N\boldsymbol{s}_{p}\in\mathcal{N} by

Then we aggregate the similarity representations with sf=∑sp∈Nβpsp\boldsymbol{s}_{f}=\sum_{\boldsymbol{s}_{p}\in\mathcal{N}}{\beta}_{p}\boldsymbol{s}_{p}, and feed sf\boldsymbol{s}_{f} into a fully-connect layer to predict the final similarity between the input image and sentence. The SAF module learns the significance scores to increase the contribution of more-informative similarity representations and meanwhile reduce the disturbance of less-meaningful alignments.

Training Objectives and Inference Strategies

We utilize the bidirectional ranking loss (Faghri et al. 2017) to train both the SGR and SAF modules. Given a matched image-text pair (v,t)(\boldsymbol{v},\boldsymbol{t}), and the corresponding hardest negative image v−\boldsymbol{v}^{-} and the hardest negative text t−\boldsymbol{t}^{-} within a minibatch, we compute the bidirectional ranking loss with

where γ\gamma is the margin parameter and Sr(⋅,⋅)\mathcal{S}_{r}(\cdot,\cdot) indicates similarity prediction function implemented with SGR. Similarly, we define the training objectives on SAF module as Lf\mathcal{L}_{f}.

In this paper, we explore different training and inference strategies with the proposed SGR and SAF modules: joint training and independent training. For joint training, we combine Lr\mathcal{L}_{r} and Lf\mathcal{L}_{f} to train SGR and SAF modules simultaneously, where the similarity representations are shared for the proposed two modules. For independent training, we train the SGR and SAF modules separately. At the inference stage, we average the similarities predicted by SGR and SAF modules for the retrieval evaluation.

Experiments

To verify the effectiveness of the our model, in this section we demonstrate extensive experiments on two benchmark datasets. We also introduce detailed implementations and training strategy of the proposed SGRAF model.

We evaluate our model on the MSCOCO (Lin et al. 2014) and Flickr30K (Young et al. 2014) datasets. The MSCOCO dataset contains 123,287 images, and each image is annotated with 5 annotated captions. The dataset is split into 113,287 images for training, 5000 images for validation and 5000 images for testing. We report results by averaging over 5 folds of 1K test images and testing on the full 5K images. The Flickr30K dataset contains 31,783 images with 5 corresponding captions each. Following the split in (Frome et al. 2013), we use 1,000 images for validation, 1,000 images for testing and the rest for training.

Protocols.

For image-text retrieval, we measure the performance by Recall at K (R@K) defined as the proportion of queries whose ground-truth is ranked within the top KK. We adopt R@1, R@5 and R@10 as our evaluation metrics.

Implementation Details.

For each image, we take the Faster-RCNN (Ren et al. 2015) detector with ResNet-101 provided by (Anderson et al. 2018) to extract the top K=36K=36 region proposals and obtain a 2048-dimensional feature for each region. For each sentence, we set the word embedding size as 300, and the number of hidden states as 1024. The dimension of similarity representation mm is 256, with smooth temperature λ=9\lambda=9, reasoning steps N=3N=3, and margin γ=0.2\gamma=0.2. Our model employs the Adam optimizer (Kingma and Ba 2015) to train the SGRAF network with the mini-batch size of 128. The learning rate is set to be 0.0002 for the first 10 epochs and 0.00002 for the next 10 epochs on MSCOCO. For Flickr30K, we start training the SGR (SAF) module with learning rate 0.0002 for 30 (20) epochs and decay it by 0.1 for the next 10 epochs. We select the snapshot with the best performance on the validation set for testing.

Quatitative Results and Analysis

In this section, we present the retrieval results on the MSCOCO and Flickr30K datasets, aiming to demonstrate the effectiveness and superiority of the proposed approach.

Table 1 and 2 report the experimental results on MSCOCO dataset with 1K and 5K test images, separately. We can see that our proposed SGRAF model outperforms the existing methods, with the best R@1=79.6%79.6\% for sentence retrieval and R@1=63.2%63.2\% for image retrieval with 1K test images. For 5K test images, the proposed approach maintains the superiority with an improvement of more than 3%3\% on the R@1 results. It should be noted that competitive retrieval performance can be also achieved with the SGR/SAF module alone, demonstrating the effectiveness and complementarity of our modules.

Comparisons on Flickr30K.

Table 1 compares the bidirectional retrieval results on Flickr30K dataset with the latest algorithms. We can observe that the SAF module alone produces comparable retrieval results and the SGR module achieves state-of-the-art performance with R@1 of 75.2%75.2\% and 56.2%56.2\% for sentence and image retrieval, separately. This verifies the effectiveness of exploiting the relationship between alignments to boost similarity reasoning. When we combine the SAF and SGR module, the performance is further improved to achieve the best R@1 of 77.8%77.8\% and 58.5%58.5\%.

Ablation Studies

In this section, we carry a series of ablation studies to explore the impact of different configurations for the SGR module, the similarity representation learning module and the process of training. We also compare different strategies of similarity prediction to demonstrate the superiority of SGR and SAF modules. All the comparative experiments are conducted on the Flickr30K dataset.

In Table 3 we investigate the effectiveness of each component in the SGR module. 1) Graph reasoning. We employ a framework without graph reasoning as the baseline(#1), which adopts a fully-connected layer and sigmoid function on the global alignment to obtain the final similarity. Comparing #1 and #6 based on R@1, the SGR module achieves 12.8%12.8\% improvement for sentence retrieval and 10.2%10.2\% for image retrieval. 2) Reasoning steps setting. Comparing #4, #5, #6 and #7, we set the step of the SGR module to 33 for maximum performance. 3) Global and local alignments. #2 and #3 only utilize local alignments for graph reasoning and adopt a mean-pooling operation on them after reasoning. Comparing #2, #4 and #3, #6, we discover that global similarity is beneficial for aggregating local similarities and exploring their relations which improves at least 1.6%1.6\% for sentence retrieval and 1.9%1.9\% for image retrieval on R@1.

Configurations for Similarity Computation.

Table 4 illustrates the impact of different strategies in similarity representation computation and the similarity score prediction. We test the results on local alignments and set the reasoning step of the SGR module to 3. we following(Lee et al. 2018) to explore two types of the cross-attention modes, i.e. I2T and T2I. Comparing #1, #2, #5 and #6, we find that averaging the local alignments calculated by a fully-connected layer and sigmoid function leads to better performance than averaging local cosine distance. Comparing #3 and #7, it is more reasonable for the SGR module to count on the local alignments attended by word features (T2I) than the ones by region features (I2T). Besides, the SGR module fails to achieve significant improvement on I2T which indicates that the region features are redundant, relatively independent and irregular in order. Therefore, it is difficult for the SGR module to exploit semantic connections compared with word features. In terms of #4 and #8, the SAF module achieves impressive progress both in I2T and T2I modes that demonstrates that the SAF module filters and aggregates plenty of discriminative local alignments steadily to improve the precision of image-text matching.

Configurations for Training Process.

In table 5, we report the results of different training strategies: joint learning and independent learning. Compared with the SGR/SAF module alone, joint learning can help the SAF module improve the performance of sentence retrieval, and also help the SGR module enhance the ability of image retrieval. In terms of independent learning, the SGRAF network gains an exact and impressive promotion. We assume that the SGR module frequently captures several crucial cues by propagating information between local and global alignments and throws out some relatively unimportant interactions. Moreover, the SAF module attempts to gather all the meaningful alignments and eliminates completely irrelevant interactions. Therefore, the global and local alignments for the SAF and SGR modules are seemingly not incompatible resulting in the unobvious improvement. It is worth noting that the SAF module tends to be more susceptible to the hard negative samples than the SGR module because of the high correlation. On the other hand, it is more challenging for the SGR module to resolve the transmission and integration of numerous semantic alignments. As a result, they can cooperate with each other and further achieve more accurate similarity prediction through independent training.

Qualitative Results and Analysis

As it is shown in Figure 3, we illustrate the distribution of attention weights learned by the SAF module. Given an image query, the SAF module captures the key cues ("dog runs", "green grass", "wooden fence") for positive image-text pairs, and also highlights the meaningful instances ("brown dog", "white paws", "trotting", "green grass") for negative pairs. Note that there exists a crucial discrepancy ("brown") which is submerged by AVE operation between negative text and image that depicts a black and white dog. Compared with the wrong matching of AVE, SAF module can stress on all the useful alignments including unmatched instance ("brown") and suppress irrelevant interactions ("of", "with", "is", and etc). On the other hand, the process of SGR module reinforces the role of the alignment ("brown"), which leads to lower similarity between hard negative and query image. Our implementation of this paper is publicly available on GitHub at: https://github.com/Paranioar/SGRAF.

Conclusion

In this work, we present a SGRAF network consisting of similarity graph reasoning (SGR) and similarity attention filtration (SAF) module. The SGR module performs multi-step reasoning based on global and local similarity nodes and captures their relations through information propagation, while the SAF module attends more to discriminative and meaningful alignments for similarity aggregation. We demonstrate that it is important to exploit the relationship between local and global alignments, and suppress the disturbances of less-meaningful alignments. Extensive experiments on benchmark datasets show that both SGR and SAF modules can effectively discover the associations between image and text and achieve further improvements when cooperating with each other.

Acknowledgments

The paper is supported in part by the National Key R&\&D Program of China under Grant No. 2018AAA0102001 and National Natural Science Foundation of China under Grant No. 61725202, U1903215, 61829102, 91538201, 61771088, 61751212 and the Fundamental Research Funds for the Central Universities under Grant No. DUT19GJ201 and Dalian Innovation Leader’s Support Plan under Grant No. 2018RD07.

References

References

Appendix Overview

This supplementary document for similarity reasoning and filtration is organized as follows: 1) more diagrams and descriptions of the SGRAF network: self-attention and SGR module; 2) more quantitative studies: the impact of graph dimension; 3) more qualitative studies: retrieval examples of bidirectional retrieval and visualization of our model.

Table 6 shows the detailed implementations of the proposed SGRAF network including generic representation extraction, similarity representation learning, similarity graph reasoning and attention filtration.

Generic Representation Extraction. Given an image, we first apply Faster R-CNN (Anderson et al. 2018) to extract the top KK=36 region proposals and obtain 2048-d feature for each region, then we add a FC layer to transform region features into 1024-d vectors V\boldsymbol{V}, and perform the self-attention mechanism (Vaswani et al. 2017) to output a 1024-d global visual vector v‾\overline{\boldsymbol{v}}. Given a sentence with LL words, we transform each word into a 300-d vector with word-embedding, and use Bi-GRU to encode words into 1024-d vectors T\boldsymbol{T}. Similarly, we exploit the self-attention mechanism (Vaswani et al. 2017) illustrated in Figure 5 to output a 1024-d global textual vector t‾\overline{\boldsymbol{t}}.

Similarity Representation Learning. We compute LL textual-attended 256-d similarity vectors sl\boldsymbol{s}^{l} with Eq.(5), and one global similarity vector sg\boldsymbol{s}^{g} with Eq.(2), which obtain LL+1 (local+global) 256-d similarity vectors N\mathcal{N}.

Similarity Graph Reasoning. As shown in Figure 4, we take the above-introduced LL+1 (256-d) similarity vectors N\mathcal{N} as graph nodes, and then compute the weight of each edge via Eq.(6) with learnable parameter matrices. Graph reasoning is conducted with Eq.(7-8), which means that, for each node sp\boldsymbol{s}_{p} at step nn, we learn the weight of its connected nodes (including itself) to aggregate their features from step nn-1, and then perform a non-linear transformation to update the feature of sp\boldsymbol{s}_{p} at step nn. In this way, the information from both local and global alignments is aggregated to produce more accurate similarity predictions. Then we feed the reasoned 256-d global vector sr\boldsymbol{s}_{r} into a FC+sigmoid layer to output a scalar similarity.

Similarity Attention Filtration. The SAF module takes the LL+1 (256-d) similarity vectors N\mathcal{N} as inputs to learn LL+1 attention weights β\beta with Eq.(9) and performs aggregation to output one 256-d similarity vector sf\boldsymbol{s}_{f}, which is then fed into another FC+sigmoid layer to output a scalar similarity.

Quantitative Studies

We evaluate the SGR module with different graph dimension mm as illustrated in Table 7. We test the results on global and local alignments and set the reasoning step to 3. The parameters during each step are not shared. We observe that the SGR module is insensitive to the dimension of similarity representation that implies the stabilization and robustness of the SGR module. Note that we set graph dimension mm to 256, which can yield the best results for image-text retrieval.

Qualitative Studies

In this section, we exhibit the retrieval examples of sentence retrieval in Figure 6. Retrieval examples of image retrieval are shown in Figure 7. Furthermore, we demonstrate additional visualization of the SGRAF model in Figure 8 where the local alignments are attended by textual words.

Retrieval Examples of Bidirectional Retrieval. For sentence retrieval, our proposed SGRAF model can efficiently retrieve the correct sentences. Note that the mismatch of F30K-Query3 is also reasonable, which includes highly relevant descriptions of concepts ("young boy", "handheld shovel") and scene ("dirt") with the image. For image retrieval, our network can distinguish hard samples well and retrieve the ground-truth image accurately, even if negative samples consist of the same semantic concepts, attributes, and relations with the text descriptions.

Visualization of the SGRAF Model. In Figure 8, the SAF module can selectively aggregate the discriminative alignments and meanwhile reduce the interferences of less-meaningful alignments, e.g. for the first image query, the SAF module can highlight the key alignments ("two man", "dancing", "street", "synchronized martial arts performance", etc.) and suppress irrelevant ones ("the", "of", "in", "a", "be", etc). Besides, the SGR module can capture fine-grained alignments to achieve comprehensive similarity reasoning, e.g. for the second image query, the SGR module stresses on the alignments ("young boy", "Texas") and produces larger gaps between matched and unmatched pairs.