Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders

Nicola Messina, Giuseppe Amato, Andrea Esuli, Fabrizio Falchi, Claudio Gennaro, Stéphane Marchand-Maillet

Introduction

Since 2012, deep learning has obtained impressive results in several vision and language tasks. Recently, various attempts have been made to merge the two worlds, and state-of-the-art results have been obtained in many of these tasks, including visual question answering (Hu et al., 2017; Anderson et al., 2018; Teney et al., 2017), image captioning (Zhou et al., 2020; Rennie et al., 2017; Huang et al., 2019; Cornia et al., 2019), and image-text matching (Chen et al., 2019; Lu et al., 2019; Faghri et al., 2018; Lee et al., 2018). In this work, we deal with the cross-modal retrieval task, with a focus on the visual and textual modalities. The task consists in finding the top-relevant images representing a natural language sentence given as a query (image-retrieval), or, vice versa, in finding a set of sentences that best describe an image given as a query (sentence-retrieval).

This cross-modal retrieval task is closely related to image-sentence matching, which consists in assigning a score to a pair composed of an image and a sentence. The score is high if the sentence adequately describes the image, and low if the input sentence is unrelated to the corresponding image. The score function learned by solving the matching problem can be then used for deciding which are the top-relevant images and sentences in the two image- and sentence- retrieval scenarios. The matching problem is often very difficult since a deep high-level understanding of images and sentences is needed for succeeding in this task.

Visuals and texts are used by humans to fully understand the real world. Although they are of equal importance, the information hidden in these two modalities has a very different nature. The text is already a well-structured description developed by humans in hundreds of years, while images are nothing but raw matrices of pixels hiding very high-level concepts and structures. Images and texts do not describe only static entities. In fact, they can easily portray relationships between the objects of interest, e.g.: ”The kid kicks the ball”. Therefore, it would be helpful to also understand spatial and even abstract relationships linking them together.

Vision and language matching has been extensively studied (Faghri et al., 2018; Carrara et al., 2018; Lu et al., 2019; Karpathy and Fei-Fei, 2015; Lee et al., 2018). Many works employ standard architectures for processing images and texts, such as CNNs-based models for image processing and recurrent networks for language. Usually, in this scenario, the image embeddings are extracted from standard image classification networks, such as ResNet or VGG, by employing the network activations before the classification head. Usually, descriptions extracted from CNN networks trained on classification tasks can only capture global summarized features of the image, ignoring important localized details. For this reason, recent works make extensive use of attention mechanisms, which are able to relate each visual object, extracted from the spatial locations of a feature map or an object detector to the most interesting parts of the sentence, and/or vice-versa.

Many of these works, such as ViLBERT(Lu et al., 2019), ImageBERT(Qi et al., 2020), VL-BERT(Su et al., 2020), IMRAM(Chen et al., 2020), try to learn a complex scoring function s=ϕ(I,C)s=\phi(I,C) that measures the affinity between an image and a caption, where II is an image, CC is the caption and ss is a normalized score in the range $.Theseareveryeffectivemodelsfortacklingthematchingtask,andtheyreachstate−of−the−artresults.However,theyremainveryinefficientforlarge−scaleimageorsentenceretrieval:theproblemwiththeseapproachesisthatitisnotpossibletoextractvisualandtextualdescriptionsseparately,asthepipelinesarestronglyentangledthroughcross−attentionormemorylayers.Thus,ifwewanttoretrieveimagesrelatedtoagivenquerytext,wehavetocomputeallthesimilaritiesusingthe. These are very effective models for tackling the matching task, and they reach state-of-the-art results. However, they remain very inefficient for large-scale image or sentence retrieval: the problem with these approaches is that it is not possible to extract visual and textual descriptions separately, as the pipelines are strongly entangled through cross-attention or memory layers. Thus, if we want to retrieve images related to a given query text, we have to compute all the similarities using the\phi$ function and then sort the resulting scores in descending order. This is unfeasible if we want to retrieve images or sentences from a large database in a few milliseconds.

In our previous work, we introduced the Transformer Encoder Reasoning Network (TERN) architecture (Messina et al., 2020), which is a transformer-based model able to independently process images and sentences to match them into the same common space. TERN is a useful architecture for producing compact yet informative features that could be used in cross-modal retrieval setups for efficient indexing using metric-space or text-based approaches. TERN processes visual and textual elements using transformer encoder layers, exploring and reasoning on the relationships among image regions and sentence words. However, its main objective is to match images and sentences as atomic, global entities, by learning a global representation of them inside special tokens (I-CLS and T-CLS) processed by the transformer encoder. This usually leads to performance loss and possibly poor generalization since fine-grained information useful for effective matching is lost during the projection to a fixed-sized common space.

For this reason, in this work, we propose TERAN (Transformer Encoder Reasoning and Alignment Network) in which we force a fine-grained word-region alignment. Fine-grained matching deals with the accurate understanding of the local correspondences between image regions and words, as opposed to coarse-grained matching, where only a summarized global descriptions of the two modalities is considered. In fact, differently from TERN, the objective function is directly defined on the set of regions and words in output from the architecture, and not on a potentially lossy global representation. Using this objective, TERAN tries to individually align the regions and the words contained in images and sentences respectively, instead of directly matching images and sentences as a whole. The information available to TERAN during training is still coarse-grained, as we do not inject any information about word-region correspondences. The fine-grained alignment is thus obtained in a semi-supervised setup, where no explicit word-region correspondences are given to the network.

Our TERAN proposal shares most of the previous TERN building blocks and interconnections: the visual and textual pipelines are forwarded separately and they are fused only during the loss computation, in the very last stage of the architecture, making scalable cross-modal information retrieval possible. At the same time, this novel architecture employs state-of-the-art self-attentive modules, based on the transformer encoder architecture (Vaswani et al., 2017), able to spot out hidden relationships in both modalities for a very effective fine-grained alignment.

Therefore, TERAN is able to produce independent visual and textual features usable in efficient retrieval scenarios implementing two simple visual and textual pipelines built of modern self-attentive mechanisms. In spite of its overall simplicity, TERAN is able to reach state-of-the-art results in the image and sentence retrieval task, even when compared with complex entangled visual-textual matching models. Experiments show that TERAN can generalize better with respect to the previous TERN approach.

In the evaluation of the proposed matching procedure, we used a typical information retrieval setup using the Recall@K metrics (with K={1,5,10}K=\{1,5,10\}.) However, in common search engines where the user is searching for related images and not necessarily exact matches, the Recall@K evaluation could be too rigid, especially when K=1K=1. For this reason, as in our previous work (Messina et al., 2020), in addition to the strict Recall@K metric, we propose to measure the retrieval abilities of the system with a normalized discounted cumulative gain metric (NDCG) with relevance computed exploiting caption similarities.

Summarizing, the contributions of this paper are the following:

we introduce the Transformer Encoder Reasoning and Alignment Network (TERAN), able to produce fine-grained region-word alignments for efficient cross-modal information retrieval.

we show that TERAN can reach state-of-the-art results on the cross-modal visual-textual retrieval task, both in terms of Recall@K and NDCG, while producing visually-pleasant region-words alignments without using supervision at the region-word level. Retrieval results are measured both on MS-COCO and Flickr30k datasets.

we quantitatively compare TERAN with our previous work (Messina et al., 2020), and we perform an extensive study on several variants of our novel model, including weight sharing in the last transformer layers, stop-words removal during training, different pooling protocols for the matching loss function, and the usage of different language models.

Related Work

In this section, we review some of the previous works related to image-text joint processing for cross-modal retrieval and alignment, and high-level relational reasoning, on which this work lays its foundations. Also, we briefly summarize the evaluation metrics available in the literature for the cross-modal retrieval task.

Image-text matching is often cast to the problem of inferring a similarity score among an image and a sentence. Usually, one of the common approaches for computing this cross-domain similarity is to project images and texts into a common representation space on which some kind of similarity measure can be defined (e.g.: cosine or dot-product similarities). Images and sentences are preprocessed by specialized architectures before being merged at some point in the pipeline.

Concerning image processing, the standard approach consists in using Convolutional Neural Networks (CNNs), usually pre-trained on image classification tasks. In particular, (Klein et al., 2015; Vendrov et al., 2016; Lin and Parikh, 2016; Huang et al., 2017; Eisenschtat and Wolf, 2017) use VGGs, while (Liu et al., 2017; Faghri et al., 2018; Gu et al., 2018; Huang et al., 2018a) use ResNets. Concerning sentence processing, many works (Karpathy and Fei-Fei, 2015; Faghri et al., 2018; Li et al., 2019; Lee et al., 2018; Huang et al., 2018a) employ GRU or LSTM recurrent networks to process natural language, often considering the final hidden state as the only feature representing the whole sentence. The problem with these kinds of methodologies is that they usually extract extremely summarized global descriptions of images and sentences. Therefore, a lot of useful fine-grained information needed to reconstruct inter-object relationships for precise image-text alignment is permanently lost.

For these reasons, many works try to employ region-level information, together with word-level descriptions provided by recurrent networks, to understand fine-grained alignments between words and localized patches in the image. Recent works (Huang et al., 2018c; Liu et al., 2019, 2020; Huang and Wang, 2019; Wang et al., 2019; Chen et al., 2020; Lee et al., 2018) exploit the availability of pre-computed region-level features extracted from the Faster-RCNN (Ren et al., 2015) object detector. An alternative consists in using the features maps in output from ResNets, without aggregating them, for computing fine-grained attentions over the sentences (Xu et al., 2020; Huang et al., 2018b; Wang et al., 2018; Wei and Zhou, 2020; Guo et al., 2020; Ji et al., 2020b).

Recently, the transformer architecture (Vaswani et al., 2017) achieved state-of-the-art results in many natural language processing tasks, such as next sentence prediction or sentence classification. The results achieved by the BERT model (Devlin et al., 2019) are a demonstration of the power of the attention mechanism to produce accurate context-aware word descriptions. For this reason, some works in image-text matching use BERT to extract contextualized word embeddings for representing sentences (Wu et al., 2019; Sarafianos et al., 2019; Qu et al., 2020; Wei et al., 2020). Drawing inspiration from the powerful contextualization capabilities of the transformer encoder architecture, some works use BERT-like processing on both visual and textual modalities, such as ViLBERT (Lu et al., 2019), ImageBERT (Qi et al., 2020), Pixel-BERT (Huang et al., 2020), VL-BERT (Su et al., 2020).

These latest works achieve state-of-the-art results in sentence and image retrieval, as well as excellent results on the downstream word-region alignment task (Chen et al., 2019). However, they cannot produce separate image and caption descriptions; this is an important requirement in real-world search engines, where usually, at query time, only the query element is forwarded through the network, while all the elements of the database have already been processed by means of an offline feature extraction process.

Some architectures have been designed so that they are natively able to extract disentangled visual and textual features. In particular, in (Faghri et al., 2018) the authors introduce the VSE++ architecture. They use VGG and ResNets visual features extractors, together with an LSTM for sentence processing, and they match images and captions exploiting hard-negatives during the loss computation. With their VSRN architecture (Li et al., 2019), the authors introduce a visual reasoning pipeline built of Graph Convolution Networks (GCNs) and a GRU to sequentially reason on the different image regions. Furthermore, they impose a sentence reconstruction loss to regularize the training process. The authors in (Huang et al., 2018b) use a similar objective, but employing a pre-trained multi-label CNN to find semantically relevant image patches and their vectorial descriptions. Differently, in (Sarafianos et al., 2019) an adversarial learning method is proposed, where a discriminator is used to learn modality-invariant representations. The authors in (Guo et al., 2020) use a contextual attention-based LSTM-RNN which can selectively attend to salient regions of an image at each time step, and they employ a recurrent canonical correlation analysis to find hidden semantic relationships between regions and words.

The works closer to our setup are SAEM (Wu et al., 2019) and CAMERA (Qu et al., 2020). In (Wu et al., 2019) the authors use triplet and angular loss to project the image and sentence features into the same common space. The visual and textual features are obtained through transformer encoder modules. Differently from our work, they do not enforce fine-grained alignments and they pool the final representations to obtain a single-vector representation. Instead, in (Qu et al., 2020) the authors use BERT as language model and an adaptive gating self-attention module to obtain context-enhanced visual features, projecting them into the same common space using cosine similarity. Unlike our work, they specifically focus on multi-view summarization, as multiple sentences can describe the same images in many different but complementary ways.

The loss used in our work is inspired by the matching loss introduced by the MRNN architecture (Karpathy and Fei-Fei, 2015), which seems able to produce very good region-word alignments by supervising only the global image-sentence level.

High-Level Reasoning

Another branch of research from which this work draws inspiration is focused on the study of relational reasoning models for high-level understanding. The work in (Santoro et al., 2017) proposes an architecture that separates perception from reasoning. They tackle the problem of Visual Question Answering by introducing a particular layer called Relation Network (RN), which is specialized in comparing pairs of objects. Object representations are learned using a four-layer CNN, and the question embedding is generated through an LSTM. The authors in (Messina et al., 2019, 2018) extend the RN for producing compact features for relation-aware image retrieval. However, they do not explore the multi-modal retrieval setup.

Other solutions try to stick more to a symbolic-like way of reasoning. In particular, (Hu et al., 2017; Johnson et al., 2017) introduce compositional approaches able to explicitly model the reasoning process by dynamically building a reasoning graph that states which operations must be carried out and in which order to obtain the right answer.

Recent works employ Graph Convolution Networks (GCNs) to reason about the interconnections between concepts. The authors in (Yao et al., 2018; Yang et al., 2019; Li and Jiang, 2019) use GCNs to reason on the image regions for image captioning, while (Yang et al., 2018; Li et al., 2018) use GCN with attention mechanisms to produce the scene graph from plain images.

Cross-Modal Retrieval Evaluation Metrics

All the works involved with image-caption matching evaluate their results by measuring how good the system is at retrieving relevant images given a query caption (image-retrieval) and vice-versa (caption-retrieval).

Usually the Recall@K metric is used (Faghri et al., 2018; Li et al., 2019; Qi et al., 2020; Lu et al., 2019; Lee et al., 2019), where typically K={1,5,10}K=\{1,5,10\}. On the other hand, (Carrara et al., 2018) introduced a novel metric able to capture non-exact results by weighting the ranked documents using a caption-based similarity measure.

We extend the metric introduced in (Carrara et al., 2018), giving rise to a powerful evaluation protocol that handles non-exact yet relevant matches. Relaxing the constraints of exact-match similarity search is an important step towards an effective evaluation of real search engines.

Review of Transformer Encoders

Our proposed architecture is based on the well established Transformer Encoder (TE) architecture, which heavily relies on the concept of self-attention. The self-attention mechanism tries to weight every vector of the sequence using a scalar value normalized in the range $computedasafunctionoftheinputvectorsthemselves.Inparticular,theattentioniscomputedbyusingaqueryvectorcomputed as a function of the input vectors themselves. In particular, the attention is computed by using a query vectorQandasetofkey−value(and a set of key-value (K,,V$) pairs derived from data using simple feed-forward networks, and processed as shown in Equation 1. More in detail, the attention-aware vector in output from the attention module is computed for every input element as a weighted sum of the values, where the weight assigned to each value is computed as a similarity score (scaled dot-product) between the query with the corresponding key:

Q,K,VQ,K,V are the query, the key, and the value respectively, while the factor dk\sqrt{d_{k}} is used to mitigate the vanishing gradient problem of the softmax function in case the inner product assumes too large values. In real implementations, a multi-head attention is used: the input vectors are chunked, and every chunk is processed independently using a different instantiation of the above-described mechanism. This helps in capturing the relationships between the different portions of every input vector.

Finally, the output from the TE is computed through a simple feed-forward layer applied to the Att(Q,K,V)Att(Q,K,V) vectors, with a ReLU activation function. This simple feed-forward layer casts in output a set of features having the same dimensionality of the input sequence. Two residual connections followed by layer normalization are also present around the self-attention and the feed-forward sub-modules. An overview of the transformer encoder architecture is shown in Figure 1.

Although the TE was initially developed to work on sequences, there are no architectural constraints that prevent its usage on sets of vectors instead of sequences. In fact, the TE module has not any built-in sequential prior which considers every vector in a precise position in the sequence. This makes the TE suitable for processing visual features coming from an object detector.

We argue that the transformer encoder self-attention mechanism can drive a simple but powerful reasoning mechanism able to spot hidden relationships between the vector entities, whatever nature they have (visual or textual). Also, the encoder is designed in a way that multiple instances of it could be stacked in sequence. Using multiple levels of attention helps in producing a deeper and more powerful reasoning pipeline.

Transformer Encoder Reasoning and Alignment Network (TERAN)

Our Transformer Encoder Reasoning and Alignment Network (TERAN) leverages our previous work (Messina et al., 2020) that introduced the TERN architecture. TERAN modifies the learning objective of our previous work by forcing a fine-grained alignment between the region and word features in output from the last transformer encoder (TE) layers so that meaningful fine-grained concepts are produced.

As TERN, our TERAN reasoning engine is built using a stack of TE layers, both for the visual and the textual data pipelines. The TE takes as input sequences or sets of entities, and it can reason upon these entities disregarding their intrinsic nature. In particular, we consider the salient regions in an image as visual entities, and the words present in the caption as textual entities.

More formally, the input to our reasoning pipeline is a set I={r0,r1,…,rn}I=\{r_{0},r_{1},\ldots,r_{n}\} of nn image regions (visual entities) representing an image II and a sequence C={w0,w1,…,wm}C=\{w_{0},w_{1},\ldots,w_{m}\} of mm words (textual entities) representing the corresponding caption CC. Thus, the reasoning module continuously operates on sets and sequences of nn and mm objects respectively for images and captions.

The TERN architecture in (Messina et al., 2020) produces summarized representations of both images and words by employing special I-CLS and T-CLS tokens that are forwarded towards the layers of the TEs. In the end, the processed I-CLS and T-CLS tokens gather important global knowledge from both modalities. Contrarily, TERAN does not produce aggregated fixed-sized representations for images and sentences. For this reason, it does not employ the global features constructed inside the I-CLS and T-CLS tokens. Instead, it tries to impose a global matching loss defined on the variable-length sets in output from the last TE layers that is able, as a side effect, to produce also good and interpretable region-word alignments.

The overall architecture is shown in Figure 2. We left in the scheme the I-CLS and T-CLS tokens connections for comparison with the TERN architecture presented in (Messina et al., 2020). These tokens are still used for a targeted experiment that exploits the combination of the TERN and TERAN losses (more details in Section 7). However, they are not used in the main TERAN experiments.

The visual features extracted from Faster-RCNN are conditioned with the information related to the geometry of the bounding-boxes. This is done through a simple fully-connected stack in the early visual pipeline before the reasoning steps. The two linear projection layers within the TE modules are used to project the visual and textual concepts in spaces having the same dimensionality. Then, the latest TE layers perform further processing before outputting the final features that are used to compute the final matching loss.

Differently from TERN, we initially do not share the weights of the last TE layers. We will discuss the effect of weight sharing in our ablation study, in Section 8.1.

At this point, the global similarity SklS_{kl} between the kk-th image and the ll-th sentence is computed by pooling this similarity matrix through an appropriate pooling function. Inspired by (Karpathy and Fei-Fei, 2015) and (Lee et al., 2018), we employ the max-sum pooling, which consists in computing the max over the rows of AA and then summing or, equivalently, max-over-regions sum-over-words (MrSwM_{r}S_{w}) pooling. We explore also the dual version, as in (Lee et al., 2018), by computing the max over the columns and then summing, or max-over-words sum-over-regions (MwSrM_{w}S_{r}) pooling:

Since both these similarity functions are not symmetric due to the diverse outcomes we obtain by inverting the order of the sum and max operations, we introduce also the symmetric form, obtained by summing the two:

Given the global image-sentence similarities SklS_{kl} computed through alignments pooling, we can proceed as in previous works (Faghri et al., 2018; Li et al., 2019) using a contrastive learning method: we use a hinge-based triplet ranking loss, focusing the attention on hard negatives, as introduced by (Faghri et al., 2018). Therefore, we used the following loss function:

where [x]+≡max(0,x)[x]_{+}\equiv max(0,x) and α\alpha is a margin that defines the minimum separation that should hold between the truly matching word-region embeddings and the negative pairs. The hard negatives k′{k}^{\prime} and l′{l}^{\prime} are computed as follows:

where (k,l)(k,l) is a positive pair. As in (Faghri et al., 2018), the hard negatives are sampled from the mini-batch and not globally, for performance reasons.

2. Region and Word Features Extraction

The I={r0,r1,…,rn}I=\{r_{0},r_{1},\ldots,r_{n}\} and C={w0,w1,…,wm}C=\{w_{0},w_{1},\ldots,w_{m}\} initial descriptions for images and captions come from state-of-the-art visual and textual pre-trained networks, Faster-RCNN with Bottom-Up attention and BERT respectively.

Faster-RCNN (Ren et al., 2015) is a state-of-the-art object detector. It has been used in many downstream tasks requiring salient object regions extracted from images. Therefore, Faster-RCNN is one of the main architectures implementing human-like visual perception. The work in (Anderson et al., 2018) introduces bottom-up visual features by training Faster-RCNN with a Resnet-101 backbone on the Visual Genome dataset (Krishna et al., 2017). Using these features, they can reach remarkable results on the two downstream tasks of image captioning and visual question answering.

Concerning text processing, we use BERT (Devlin et al., 2019) for extracting word embeddings. BERT already uses a multi-layer transformer encoder to process words in sentences and capture their functional relationships through the same powerful self-attention mechanism. BERT embeddings are trained on some general natural language processing tasks such as sentence prediction or sentence classification and demonstrated state-of-the-art results in many downstream natural language tasks. BERT embeddings, unlike word2vec (Mikolov et al., 2013), capture the context in which each word appears. Therefore, every word embedding carries information about the surrounding context, that could be different from caption to caption.

Since the transformer encoder architecture does not embed any sequential prior in its architecture, words are given a sequential order by mixing some positional information into the learned input embeddings. For this reason, the authors in (Vaswani et al., 2017) add sine and cosine functions of different frequencies to the input embeddings. This is a simple yet effective way to transform a set into a sequence.

Computational Efficiency of TERAN

A principled objective of our work is efficient feature extraction for cross-modal retrieval applications. In these scenarios, it is mandatory to have a separable network that can produce visual and textual features by independently forwarding the two disentangled visual and textual pipelines. Furthermore, the similarity function should be simple, so that it is efficient to compute.

TERAN, as well as TERN (Messina et al., 2020) and other works in literature (Wu et al., 2019; Qu et al., 2020; Li et al., 2019) adhere to this principle. In fact, if KK is the number of images and LL the number of sentences in the database, these methods have a feature space complexity, as well as a feature extraction time complexity, of O(K)+O(L)O(K)+O(L).

Other works that entangle the visual and textual pipelines, such as (Chen et al., 2020; Xu et al., 2020; Wang et al., 2019) require a feature space, and a number of network evaluations, scaling with O(KL)O(KL). These methods are impractical to deploy to real-world scalable search engines. Some of these works partially solve this issue by keeping the two representations separated up until a certain point in the network, so that these intermediate representations can be cached, as proposed in (MacAvaney et al., 2020). In all these cases, a new incoming query to the system needs O(K)O(K) or O(L)O(L) re-evaluations of the whole network (depending on whether we are considering image or sentence retrieval); in the best case, we need to re-evaluate the last attention layers, which could be similarly expensive.

Regarding the similarity computation, TERN uses simple dot products that enable quick and efficient document rankings in modern search engines. TERAN implements also a very simple similarity function, built of simple dot products and summations without including complex layers of memories or attentions. This possibly enables an implementation that uses metric space approaches to prune the search space and obtain very efficient image or sentence rankings for a given query. However, the implementation of the TERAN similarity function in real-world search engines is left for future research.

NDCG metric for Cross-Modal Retrieval

As of now, many works in the computer vision literature treating image-text matching measure the retrieval abilities of the proposed methods by employing the well known Recall@K metric. The Recall@K measures the percentage of queries able to retrieve the correct item among the first K results. This is a metric perfectly suitable for scenarios where the query is very specific and thus we expect to find the elements that match perfectly among the first search results. However, in common search engines, the users are not asked to input a very detailed query, and they are often not searching for an exact match. They expect to find in the first retrieved positions some relevant results, with relevance defined using some pre-defined and often subjective criterion.

For this reason, inspired by the work in (Carrara et al., 2018) and following the novel ideas introduced by our previous work on TERN (Messina et al., 2020), we employ a common metric often used in information retrieval applications, the Normalized Discounted Cumulative Gain (NDCG). The NDCG is able to evaluate the quality of the ranking produced by a certain query by looking at the first pp positions of the ranked elements list. The premise of NDCG is that highly relevant items appearing lower in a search result list should be penalized as the graded relevance value is reduced proportionally to the position of the result.

The NDCG until position pp is defined as follows:

reli\text{rel}_{i} is a positive number encoding the affinity that the ii-th element of the retrieved list has with the query element, and IDCGp\text{IDCG}_{p} is the DCGp\text{DCG}_{p} of the best possible ranking. Thanks to this normalization, NDCGp\text{NDCG}_{p} acquires values in the range $$.

We thus compute the relirel_{i} value in the following ways:

reli=τ(Cˉi,Cj)\text{rel}_{i}=\tau(\bar{C}_{i},C_{j}) in case of image retrieval, where CjC_{j} is the query caption

reli=τ(Cˉj,Ci)\text{rel}_{i}=\tau(\bar{C}_{j},C_{i}) in case of caption retrieval, where Cˉj\bar{C}_{j} is the set of captions associated to the query image IjI_{j}.

In our work, we use ROUGE-L(Lin, 2004) and SPICE(Anderson et al., 2016) as sentence similarity functions τ\tau for computing caption similarities. These two scoring functions capture different aspects of the sentences. In particular, ROUGE-L operates on the longest common sub-sequences, while SPICE exploits graphs associated with the syntactic parse trees, and has a certain degree of robustness against synonyms. In this way, SPICE is more sensitive to high-level features of the text and semantic dependencies between words and concepts rather than to pure syntactic constructions.

Experiments

We trained the TERAN architecture and we measured its performance on the MS-COCO (Lin et al., 2014) and Flickr30k datasets (Young et al., 2014), computing the effectiveness of our approach on the image retrieval and sentence retrieval tasks. We compared our results against state-of-the-art approaches on the same datasets, using the introduced NDCG and the already-in-use Recall@K metrics.

The MS-COCO dataset comes with a total of 123,287 images. Every image has associated a set of 5 human-written captions describing the image. We follow the splits introduced by (Karpathy and Fei-Fei, 2015) and followed by the subsequent works in this field (Faghri et al., 2018; Gu et al., 2018; Li et al., 2019). In particular, 113,287 images are reserved for training, 5,000 for validating, and 5,000 for testing. Differently, Flickr30k consists of 31,000 images and 158,915 English texts. Like MS-COCO, each image is annotated with 5 captions. Following the splits by (Karpathy and Fei-Fei, 2015), we use 29,000 images for training, 1,000 images for validation, and the remaining 1,000 images for testing. For MS-COCO, at test time the results for both 5k and 1k test-sets are reported. In the case of 1k images, the results are computed by performing 5-fold cross-validation on the 5k test split and averaging the outcomes.

We computed caption-caption relevance scores for the NDCG metric using ROUGE-L(Lin, 2004) and SPICE(Anderson et al., 2016), as explained in Section 6, and we set the NDCG parameter p=25p=25 as in (Carrara et al., 2018) in our experiments. We employed the NDCG metrics measured during the validation phase for choosing the best performing model to be used during the test phase.

For a better comparison with our previous TERN approach, we included three more targeted experiments. In the first two, called TERN MwSrM_{w}S_{r} Test and TERN MrSwM_{r}S_{w} Test we used the best-performing TERN model, trained as explained in (Messina et al., 2020), testing it using the MwSrM_{w}S_{r} and MrSwM_{r}S_{w} alignments criteria respectively. TERN is effectively able to output features for every image region or word; however, it is never constrained to produce meaningful descriptions out of these sets of features; hence, this trial is aimed at checking the quality of the alignment of the concepts in output from the previous TERN architecture. In the third experiment, called TERN w. Align, we tried to integrate the objectives of both TERN and TERAN during training, by combining their losses using the uncertainty weighting method proposed in (Kendall et al., 2018), and testing the model using the TERN inference protocol. Thus, in this experiment, we effectively reuse the I-CLS and T-CLS tokens as global descriptions for images and sentences, as described in (Messina et al., 2020). This experiment aimed to evaluate if the TERAN alignment objective can help TERN learn better fixed-sized global vectorial descriptions.

We employ the BERT model pre-trained on the masked language task on English sentences, using the PyTorch implementation by HuggingFace https://github.com/huggingface/transformers. These pre-trained BERT embeddings are 768-D. For the visual pipeline, we extracted the bottom-up features from the work by (Anderson et al., 2018), using the code and pre-extracted features provided by the authors https://github.com/peteanderson80/bottom-up-attention. Specifically, for MS-COCO we used the already-extracted bottom-up features, while we extracted from scratch the features for Flickr30k using the available pre-trained model.

In the experiments, we used the bottom-up features containing the top 36 most confident detections, although our pipeline already handles variable-length sets of regions for each image by appropriately masking the attention weights in the TE layers.

Concerning the reasoning steps, we used a stack of 4 TE layers for visual reasoning. We found the best results when fine-tuning the BERT pre-trained model, so we did not add further reasoning TE layers for the textual pipeline. The final common space, as in (Faghri et al., 2018), is 1024-dimensional. We linearly projected the visual and textual features to a 1024-d space and then we processed the resulting features using 2 final TEs before computing the alignment matrix.

All the TEs feed-forward layers are 2048-dimensional and the dropout is set to 0.1. We trained for 30 epochs using Adam optimizer with a batch size of 40 and a learning rate of 1e−51e{-5} for the first 20 epochs and 1e−61e{-6} for the remaining 10 epochs. The α\alpha parameter of the hinge-based triplet ranking loss is set to 0.2, as in (Faghri et al., 2018; Li et al., 2019).

2. Results

We compare our TERAN method against the following baselines: JGCAR (Wang et al., 2018), SAN (Ji et al., 2019), VSE++ (Faghri et al., 2018), SMAN (Ji et al., 2020b), M3A-Net (Ji et al., 2020a), AAMEL (Wei and Zhou, 2020), MRNN (Karpathy and Fei-Fei, 2015), SCAN (Lee et al., 2018), SAEM (Wu et al., 2019), CASC (Xu et al., 2020), MMCA (Wei et al., 2020), VSRN (Li et al., 2019), PFAN (Wang et al., 2019), Full-IMRAM (Chen et al., 2020), and CAMERA (Qu et al., 2020). We clustered these methods based on the visual feature extractor they use: VGG, ResNet, or Region CNN (e.g., Faster-RCNN). To have a better comparison with our method, we also annotated in the tables whenever they use BERT as the textual model, or if they use disentangled visual-textual pipelines for efficient feature computation, as explained in Section 5. Also note that many of the listed methods report the results using an ensemble of two models having different training initialization parameters, where the final similarity is obtained by averaging the scores in output from each model. Hence, we reported also our ensemble results, for a better comparison with these baselines. In the tables, we indicate ensemble methods postponing ”(ens.)” to the method name.

We used the original implementations from their respective GitHub repositories to compute the NDCG metrics for the baselines, where possible. In the case of missing pre-trained models, we were not able to produce consistent results with the original papers. In this case, we do not report the NDCG metrics (”-”).

On both the 1K and 5K test sets, our novel TERAN approach reaches state-of-the-art results on almost all the metrics. Concerning the results reported in Table 1 regarding 1K test set, the best performing TERAN model is the one implementing the max-over-regions sum-over-words (MrSwM_{r}S_{w}) pooling method, although the model using the symmetric loss reaches comparable results. We chose the same TERAN MrSwM_{r}S_{w} model to evaluate the ensemble, reaching an improvement of 5.7% and 3.5% on the Recall@1 metric on image and sentence retrieval respectively, with respect to the best baseline using ensemble methods, which is CAMERA (Qu et al., 2020). Notice, however, that even the basic TERAN model without ensemble is able to surpass CAMERA in many metrics. This confirms the power of the TERAN model despite its overall simplicity.

Table 2 reports the results for the 5K test set, which confirm the superiority of TERAN MrSwM_{r}S_{w} over all the baselines also on the full test set. In this scenario, we increase the Recall@1 performance by 11.3% and 7.6% on image and sentence retrieval with respect to the CAMERA approach. On the other hand, the max-over-words sum-over-regions (MwSrM_{w}S_{r}) method loses around 10% on the Recall@1 metrics with respect to the best performing TERAN non-ensemble model. In this case, the Recall@K metric does not improve over top results obtained by the current state-of-the-art approaches. Nevertheless, this model loses only about 1.5% during image-retrieval and about 3.5% during sentence-retrieval as far as the SPICE NDCG metric is concerned, reaching perfectly comparable results with our state-of-the-art method. In light of these results, we deduce that the MwSrM_{w}S_{r} model is not so effective in retrieving the perfect-matching elements; however, it is still very good at retrieving the relevant ones.

As far as image retrieval is concerned, in the TERN MwSrM_{w}S_{r}Test and TERN MwSrM_{w}S_{r}Test experiments we can see that the TERN architecture trained as in (Messina et al., 2020) performs fairly good when the similarity is computed as in the novel TERAN architecture, using the region and words outputs and not the I-CLS and T-CLS global descriptions. In particular, the use of max-over-words sum-over-regions similarity still works quite well compared to the similarity computed through I-CLS and T-CLS global visual and textual features as it is in TERN.

Notice instead that on the sentence retrieval task, the TERN MrSwM_{r}S_{w} Test experiment obtains a very low performance. This is the consequence of the fact that TERN is trained to produce global-scale image-sentence matchings, while it is never forced to produce meaningful fine-grained aligned concepts. This is further supported by the evidence that if we visualize the region-words alignments as explained in the following Section 8.5 we obtain random word groundings on the image, meaning that the concepts in output from TERN are not sufficiently informative.

In order to better compare TERAN with our previous TERN approach, in Figure 3 we report the validation curves for both NDCG and Recall@1 metrics, for both methods. We can notice how the NDCG metric overfits in our previous TERN model, especially when using the SPICE metric, while the Recall@ keeps increasing. On the other hand, TERAN demonstrates better generalization abilities on both metrics. This is a clear indication that TERAN is better able to retrieve relevant items in the first positions, as well as exact matching elements. Instead, TERN is more prone to overfitting to the SPICE metric, meaning that at a certain point in training, the network still searches for the top matching element, but with a tendency to push away possible relevant results compared to the novel TERAN approach.

However, looking at the results from the TERN w. Align experiment, we can notice that by augmenting the TERN objective with the TERAN alignment loss, we can slightly increase the TERN overall performance. This confirms that a more precise and meaningful region-word alignment has a visible effect also on the quality of the fixed-sized global embeddings produced by TERN.

In Table 3 we report the results on the Flickr30k dataset. Our single-model TERAN MrSwM_{r}S_{w} method outperforms the best baseline (CAMERA) on the image retrieval task while approaching the single-model CAMERA performance on the sentence retrieval task. Nevertheless, even on Flickr30k our TERAN MrSwM_{r}S_{w} method with model ensemble obtains state-of-the-art results with respect to all the baselines on all the metrics, gaining 4.6% and 1.5% on the Recall@1 metric on the image and sentence retrieval tasks respectively.

On the MS-COCO dataset, our system powered by a single GTX 1080Ti can compute a single image-to-sentence query in ∼0.12s\sim 0.12s on 5k sentences of the test split; in the sentence-to-image scenario, it can produce scores and rank the 1K images in ∼0.02s\sim 0.02s. These timings allow the TERAN scores to be effectively used, for example, in a re-ranking phase, where the first 1k images - 5k sentences have been previously retrieved using a faster descriptor (e.g., the one from TERN).

3. Qualitative Analysis for Image Retrieval

The visualization of image retrieval results is a good way to qualitatively appreciate the retrieval abilities of the proposed TERAN model. Figures 4 and 5 show examples of images retrieved given a textual caption as a query, with scores computed using the max-over-regions sum-over-words method. In particular, Figure 4 shows image retrieval results for a couple of flexible query captions. The red-marked images represent the exact-matching elements from the ground-truth. We can therefore conclude that the retrieved images in these examples are incorrect results for the Recall@1 metric (and for the first query even for Recall@5). However, in the very first positions, we find non-matching yet relevant images, due to the ambiguity of the query caption. These are common examples where NDCG succeeds over the Recall@K metric since we need a flexible evaluation for not-too-strict query captions.

Figure 5 reports instead image retrieval results for a couple of very specific query captions. For the first two queries, the network succeeds in positioning the only really relevant image in the first position (a dog sitting on a bench on the upper query, and Pennsylvania Avenue, uniquely identifiable by the street sign, on the lower query). In this case, the Recall@1 metric also succeeds, given that the query captions are very selective. The third example, instead, evidences a failure case where the model cannot deal with very subtle details. The (only) correct result is ranked 6th in this case; in the first ranking positions, the model can find images with a vase used as a centerpiece, but the table is not often visible, and when it is visible, it is not in the corner of the room.

Ablation Study

We tried to apply weight sharing for the last 2 layers of the TERAN architecture, those after the linear projection to the 1024-d space. Weight sharing is used to reduce the size of the network and enforce a structure able to perform common reasoning on the high-level concepts, possibly reducing the overfitting and increasing the stability of the whole network. We experimented with the effects of weight sharing on the MS-COCO dataset with 1K test set, for both the max-over-words sum-over-regions and the max-over-regions sum-over-words scenarios.

Results are shown in the 2-nd and 6-th rows of Table 4. It can be noticed that the values are perfectly comparable with the TERAN results reported in Table 1, suggesting that at this point in the network the abstraction is high enough that concepts coming from images and sentences can be processed in the exact same way. This result shows that vectors at this stage have been freed from any modality bias and they are fully comparable in the same representation space.

Also, in the max-over-words sum-over-regions scenario (6-th row), there is a small gain both in terms of Recall@K and NDCG. This confirms the slight regularization effect of the weight sharing approach.

2. Averaging Versus Summing

We tried to do average instead of sum during the last pooling phase of the alignment matrix. We consider only the case in which we average-over-sentences; in fact, since in our experiments the number of visual concepts is always fixed to the 36 more influent ones during the object detection stage, average-over-regions and sum-over-regions do not differ substantially.

Thus, we considered the case of max-over-regions average-over-words (MrAvgwM_{r}\text{Avg}_{w}):

If we compute the average instead of the sum in the max-over-regions sum-over-words scenario, the final similarity score between the image and the sentence is no more dependent on the number of concepts from the textual pipeline: the similarities are averaged and not accumulated.

In the 3-rd row of Table 4 we can notice that by averaging we lose an important amount of information with respect to the max-over-regions sum-over-words scenario (1-st row). This insight suggests that the complexity of the query is beneficial for achieving high-quality matching.

Another side effect of using average instead of the max is the premature clear overfitting on the NDCG metrics as far as image-retrieval is concerned. The effect is shown in Figure 6. The clear overfitting of the NDCG metrics resembles the training curve trajectories of TERN (Figure 3). This result demonstrates that although this model can correctly perform exact matching, it is pulling away relevant results from the head of the ranked list of images, during the validation phase.

3. Removing Stop-Words During Alignment

Some words may carry no substantial meaning by themselves, such as articles or prepositions. These words with a high and diffuse frequency of use are typically called stop-words and are usually removed in classical text analysis processes. In this context, removing stop-words may help the architecture to focus only on the important concepts. Doing so, the training process is simplified as the noise introduced by possibly irrelevant words is removed. Results are reported in the 4-th row of Table 4.

The overall performance, both in terms of Recall@ and NDCG is comparable, yet with a small decrease, with the one obtained without stop-words removal (1-st row of the table). This suggests that in this context stop-words are linguistic elements that bring some useful information to distinguish ambiguous scenes. Prepositions and adverbs often indicate the spatial arrangement of objects, thus ”chair near the table” is not the same as ”chair over the table”; distinguishing these fine-grained differences is beneficial for obtaining a precise image-text matching.

4. Using Different Language Models

Despite the power of BERT (Devlin et al., 2019) for obtaining contextualized representations for the words in a sentence, many works use recurrent bidirectional networks instead, such as Bi-GRU or Bi-LSTMs.

In the 5-th and 6-th row of Table 4 we report the results for the TERAN model with Bi-GRU and Bi-LSTM in substitution of the BERT model for language processing. We used 300-dimensional word embeddings, and a hidden size of 512 so that the final bi-directional sentence feature is a 1024-dimensional description; we used the same training protocol and hyper-parameters used for the main experiments. The results suggest that BERT is an essential ingredient for reaching top results on the Recall@K metrics, especially when K={1, 5}. In particular, Bi-LSTM and Bi-GRU lose around 14% on image retrieval and 12% on sentence retrieval on the Recall@1 metric compared to the TERAN MrSwM_{r}S_{w} single-model method. However, we can notice that TERAN with these recurrent language models still maintains a comparable performance with respect to the NDCG metric, especially on the image retrieval task.

5. Visualizing the Visual-Word Alignments

Inspired by the work in (Karpathy and Fei-Fei, 2015), we try to visualize the region-word alignments learned by TERAN on some images from the test set of MS-COCO dataset. We recall that no supervision was used at the region-word level during the training phase.

In Figure 7, we report some figures where every sentence word has been associated with the top-relevant image region. The affinity between visual concepts (region features) and textual concepts (word features) has been measured through cosine similarity, just as during the training phase.

We can see that the words have overall plausible groundings on the image they describe. Some words are really difficult to ground, such as articles, verbs, or adjectives. However, we can notice that phrases describing a visual entity and composed of nouns with the related articles and adjectives (e.g. ”a green tie”, or ”a wooden table”) are often grounded to the same region. This further confirms that the TERAN architecture can produce meaningful concepts, and it is also able to cluster them under the form of complete reasonable phrases.

We can notice some wrong word groundings in the images, such as the phrase ”eyes closed” that is associated with the region depicting the closed mouth. In this case, the error seems to lie on some localized misunderstanding of the scene (in this case the noun ”eyes” has probably been misunderstood since the mouth and the eyes are both closed). Overall, however, complex scenes are correctly broken down into their salient elements, and only the key regions are attended.

Conclusions

In this work, we introduced the Transformer Encoder Reasoning and Alignment Network (TERAN). TERAN is a relationship-aware architecture based on the Transformer Encoder (TE) architecture, exploiting self-attention mechanisms, able to reason about the spatial and abstract relationships between elements in the image and in the text separately.

Differently from TERN (Messina et al., 2020), TERAN forces a fine-grained alignment among the region and word features without any supervision at this level. We demonstrated that by enforcing this fine-grained word-region alignment at training time we can obtain state-of-the-art results on the popular MS-COCO and Flickr30K datasets. Besides, thanks to the overall simplicity of the proposed model, we can obtain effective visual and textual features for use in scalable retrieval setups.

We measured the performance of our TERAN architecture in the context of cross-modal retrieval using both the already-in-use Recall@K metric and the newly introduced NDCG with the ROUGE-L and SPICE textual relevance measures. In spite of its simplicity, TERAN can outperform current state-of-the-art models on these two retrieval metrics, competing with the currently very effective entangled visual-textual matching models, which on the contrary are not able to produce features for scalable retrieval. Furthermore, we showed that TERAN can successfully output visually-pleasant word-region alignments. We also observed that a further reduction of the network complexity can be obtained by sharing the weights of the last TE layers. This has important benefits also on the stability and in the generalization abilities of the whole architecture.

In the end, we think that this work proposes an interesting path towards efficient and effective cross-modal information retrieval.

References