Learning the Best Pooling Strategy for Visual Semantic Embedding

Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, Changhu Wang

Introduction

Recognizing and describing the visual world with natural language is an essential capability for artificial intelligence. It motivates the research of image-text matching, which challenges a learning agent to establish accurate and generalizable alignment between visual and textual data, so that one can identify images or videos by text queries or vice versa.

Visual semantic embedding (VSE) tackles this challenge by learning a semantic embedding space, where the distance between paired visual and textual instances in the embedding space is optimized to be small. The core idea of the VSE has three steps:

Extract a set (or sequence) of features from data, using feature extractors (e.g., ConvNets for visual data).

Contextualize and aggregate the extracted features to project them into the joint embedding space as holistic vectors, using feature aggregators.

Compute the matching score between embeddings with a similarity metric (e.g., cosine distance).

With the feature extractor determined, one might expect that a complex aggregator is required to achieve good results. However, we show (in § 3) that a surprisingly simple and efficient aggregator, a carefully selected pooling function (e.g., max pooling), can surpass prior state-of-the-art VSE methods with complex aggregators .

Such pooling functions are both simple and effective. However, searching for the optimal pooling requires extensive manual tuning and repetitive experiments (e.g., grid search) for each data modality and features, which is tedious and costly as it enumerates over a combinatorial number of configurations. This search procedure could be even more complicated when the sets of features have varying sizes.

Can we discover the best pooling strategy automatically? In this paper, we propose a novel parameterized pooling operator, Generalized Pooling Operator (GPO), to fully exploit the strengths of pooling-based feature aggregation. GPO generalizes over various pooling functions and learns to adjust itself to the best one for different data modalities and feature extractors. Specifically, GPO learns a generator that predicts the pooling coefficients to weight the elements of sorted feature vectors, and use their weighted sum as the pooling output. The coefficient generator is instantiated as a tiny sequence model to handle variable-sized features. GPO learns to adapt to the optimal pooling strategy, and improve VSE models at a negligible extra computational cost.

With the proposed GPO, we build our multi-modal matching system as VSE∞\infty, which extends a standard VSE framework by using GPO as the feature aggregators for both visual and text features. We train our system optimizing a margin-based triplet ranking objective similar to , with the online hard-negative mining.

Without bells and whistles, VSE∞\infty surpasses all previous state-of-the-art VSE-based methods on the image-text retrieval tasks, over COCO and Flickr30K . With a straightforward extension, variants of VSE∞\infty also achieve the best video-text retrieval results on two benchmark datasets, i.e., MSR-VTT and VaTeX . In additional experiments, we show that GPO consistently outperforms other alternative learnable poolings from the literature. To better understanding GPO, we further visualize the pooling strategy found by VSE∞\infty, and compare it with the one from a thorough grid search process.

Our contributions are summarized as the following:

We empirically find that carefully selecting simple pooling functions can outperform complex visual aggregators in prior VSE methods for image-text matching.

We propose a novel Generalized Pooling Operator (GPO) that generalizes various pooling functions. It learns to automatically discover the best pooling function for image, text, and video data with various feature extractors.

We build up VSE∞\infty with GPO, which achieves the new state-of-the-art performances among VSE methods on image-text and video-text retrieval.

We visualize the pooling strategies learned by GPO, and verify that GPO learns the best pooling strategies given the data by comparing VSE∞\infty with a thorough grid search over pooling functions of all modalities.

Visual Semantic Embedding for Multi-modal Matching

We begin by revisiting the formal formulation of Visual Semantic Embedding (VSE). A VSE model (illustrated in Figure 1) leverages a visual embedding function Φ(x)\mathbf{\Phi}({\bm{x}}) such as convolutional neural networks (e.g. CNNs ), and a text embedding function Ψ(t)\mathbf{\Psi}(\bm{t}) such as sequence models (e.g. LSTMs , Transformers ), to compute the set of visual features and text features, respectively:

The compatibility score is then defined as the cosine similarity between v\bm{v} and u\bm{u}, formally as:

During the inference, the s(x,t)\bm{s}_{({\bm{x}},\bm{t})} scores are used to rank a query text against all candidate images, and the top candidate are returned as the prediction. We note that the inference procedure is efficient as the visual and text embedding v\bm{v} and u\bm{u} can be pre-computed. The pair-wise scores are then computed by matrix multiplication.

Learning Multi-modal Matching To learn a VSE model, existing methods mostly optimize the hinge-based triplet ranking loss with online hard negative mining proposed by VSE++ . The concrete matching objective is defined by:

where α\alpha is a hyper-parameter. (x,t)({\bm{x}},\bm{t}) is a positive image-text pair in the dataset D\mathcal{D} and [x]+≡max(0,x)[x]^{+}\equiv\texttt{max}(0,x). We represent t^=argmaxt′≠ts(x,t′)\hat{\bm{t}}=\texttt{argmax}_{\bm{t}^{\prime}\neq\bm{t}}\bm{s}_{({\bm{x}},\bm{t}^{\prime})} and x^=argmaxx′≠xs(x′,t)\hat{\bm{x}}=\texttt{argmax}_{{\bm{x}}^{\prime}\neq{\bm{x}}}\bm{s}_{({\bm{x}}^{\prime},\bm{t})} as the hardest negative text and image examples measured by the learned VSE model within a mini-batch.

VSE∞\infty with Generalized Pooling Operator

In this section, we first present an empirical finding that highlights the effectiveness of well-selected pooling function in VSE model, which motivates our methodological pursuit (§ 3.1). We then propose our method, Generalized Pooling Operator (GPO), with a introduction of its formal definition (§ 3.2), followed by the details of GPO’s concrete model architecture (§ 3.3). Finally, we summarize our multi-modal system (VSE∞\infty) that leverages GPO (§ 3.4).

As aforementioned in § 1, complex aggregators ff have been investigated in the VSE literature , such as sequence-to-sequence encoder (Seq2Seq), graph convolution network (GCN), self-attention encoder (SelfAttn), etc. However, we surprisingly find that these aggregation models with millions of parameters underperform carefully selected pooling functions.

Table 1 highlights a comparison between different aggregators, across two widely used image feature extractors in the literature – Grid feature is the feature maps from ConvNets and Region feature is the ROI features from object detectors (details in § 5). The results are reported in recall@1 for text-based image retrieval (T→\rightarrowI) and vice versa. Given the candidates of Average Pooling (AvgPool), Max Pooling (MaxPool) and K-Max Pooling (K-MaxPool , details in § 3.2) with different K, it shows that the best among them consistently outperform complex aggregators. Here, the best results for Region and Grid feature are achieved by MaxPool and K-MaxPool (KK=20), respectively.

Analyses of the Empirical Findings. Most complex aggregators are designed to contextualize the input features spatially, leveraging the relationship between spatial grids or regions. However, these aggregators introduce a large set of parameters in addition to the vanilla VSE model, which causes a higher risk of over-fitting comparing to simple pooling functions. In this paper, instead of investigating why complex aggregators are suboptimal, we focus on maximizing the advantages of pooling-based aggregation.

While the optimal pooling strategy enjoys simplicity and effectiveness, searching it requires repetitive experiments over numerous configurations (e.g., different K for K-MaxPool), which is both tedious and costly. This process can be more complicated when the feature extractor changes, or when the features have variable lengths (e.g., text).

Motivated by these, we aim for a general and plug-and-play pooling operator that generalizes over different pooling patterns (e.g., Avg, Max and K-MaxPool with arbitrary K) for variable-sized inputs, and learns to automatically adapt itself to the best strategy according to the data (e.g., image, text, video etc.) and feature extractors. We denote our proposed module as the Generalized Pooling Operator (GPO).

2 Generalizing over Different Pooling Strategies

AvgPool The average pooling computes the mean value among the N\mathsf{N} elements, as vi=1N∑n=1Nϕni,∀i\bm{v}^{i}=\frac{1}{\mathsf{N}}\sum^{\mathsf{N}}_{n=1}\bm{\phi}^{i}_{n},\forall i.

MaxPool The max pooling computes the maximum value among the N\mathsf{N} elements, as vi=max1({ϕni}n=1N),∀i\bm{v}^{i}=\texttt{max}_{1}(\{\bm{\phi}^{i}_{n}\}^{\mathsf{N}}_{n=1}),\forall i.

K-MaxPool The K-max pooling computes the mean value of the top-K maximum values among the N\mathsf{N} elements, as vi=1K∑k=1Kmaxk({ϕni}n=1N),∀i\bm{v}^{i}=\frac{1}{\mathsf{K}}\sum^{\mathsf{K}}_{k=1}\texttt{max}_{k}(\{\bm{\phi}^{i}_{n}\}^{\mathsf{N}}_{n=1}),\forall i.

Main Idea As described above, GPO aims to generalize over various pooling strategies, so that the pooling operator can automatically find the most appropriate strategy for different features. Therefore, GPO learns to generate the pooling coefficients θ\bm{\theta}, and the pooling is defined as a weighted sum over sorted features :

Here, the coefficients θ\bm{\theta} are of the size N\mathsf{N}, with a scalar weight θk\bm{\theta}_{k} for the kk-th maximum value among the N\mathsf{N} elements. The constraint ∑k=1Nθk=1\sum_{k=1}^{\mathsf{N}}\bm{\theta}_{k}=1 is enforced via Softmax. The parameterized pooling operator can approximate AvgPool, MaxPool, K-MaxPool with arbitrary K\mathsf{K}, and more complex pooling functions. For instance, the learned pooling strategy could weight the top-K elements unevenly, or only set non-zero values for θ1\bm{\theta}_{1} and θN\bm{\theta}_{N}. We visualize some learned pooling coefficients in § 5.3.

Learning to Generate the Pooling Coefficients The most straightforward way to parameterize θ\bm{\theta} is to define it as a trainable vector, but this can only deal with the scenario where N\mathsf{N} is a constant integer (like did for fixed-size kernels). When the features are of variable sizes, which is common in video and text sequences, learning a fixed set of coefficients θ\bm{\theta} is no longer feasible. To address this issue, we propose to learn a parameterized function g(⋅,⋅)g(\cdot,\cdot) as the coefficient generator:

As a consequence, for each position kk, the coefficient generator g(⋅,⋅)g(\cdot,\cdot) outputs a coefficient θk\bm{\theta}_{k} to aggregate {ϕni}n=1N\{\bm{\phi}^{i}_{n}\}^{\mathsf{N}}_{n=1}.

3 Implementing Generalized Pooling Operator

Now we discuss the concrete implementation of the GPO function g(⋅,⋅)g(\cdot,\cdot). Figure 2 provides an illustration of the architecture. There are two major components in the GPO design: (1) A positional encoding function based on trigonometric function; (2) A sequence model that takes the positional encoding sequence to generate pooling coefficients, based on bidirectional Gated Recurrent Unit (BiGRU).

Encoding Position Every position index kk is uniquely represented by a dense vector, such that the vector can be further transformed to θk\bm{\theta}_{k} by parameterized functions. A common approach here is to learn an embedding matrix in which row kk is the embedding for kk. However, this presumes the input positions {1,…,k,…,N}\{1,\ldots,k,\ldots,\mathsf{N}\} orthogonal to each other. To make more efficient use of the prior information between position indices, we adopt the positional encoding strategy used in Transformers to vectorize positional indices:

where wj=1100002j/d3w_{j}=\frac{1}{10000^{2j/d_{3}}} and d3d_{3} is the number of dimensions for the positional encoding.

Here hk\bm{h}_{k} is the output of the BiGRU at the position kk.

Learning Generator with Diverse Set Sizes To make GPO’s coefficient generator g(⋅,⋅)g(\cdot,\cdot) better approximate different pooling patterns for variable-sized inputs, we perform a data augmentation strategy to allow it observing a larger variety of feature set sizes. During the training, we randomly drop 20%\% inputs vectors to perturb the size of the input feature set, which we call Size Augmentation. We show in Appendix that applying this strategy to both image and text effectively improve the performance of VSE models.

4 Building up VSE∞\infty using GPO

We build up our multi-modal matching model (dubbed VSE∞\infty) by pluging GPO into the standard VSE framework (§ 2). Specifically, we replace the visual and text aggregators in the standard VSE framework (i.e., AvgPool) with two GPOs. The two GPOs project the image feature vectors and text feature vectors independently into two holistic embeddings, to further compute the matching score. VSE∞\infty is closely related to previous VSE models. We adopt the learning framework of VSE++ (Eq. 1), which improves early VSE models with an additional online hard negative mining procedure. We refer to § 5 for more details.

Related Works

Existing image-text matching methods can be categorized differently based on how the cross-modal interaction is implemented. As aforementioned, Visual Semantic Embedding (VSE) learns a joint embedding space, such that the compatibility score can be computed as a inner-product between the two holistic image and text vectors. Therefore, VSE relies on learning strong image and text embedding functions to obtain high-quality joint embedding space. Frome et al. used this approach for zero-shot image recognition , via matching visual embeddings with semantic word embeddings. Kiros et al. extends the idea by using bi-directional LSTMs to encode sentence as the semantic embedding. Faghri et al. proposes VSE++, which learns with online hard-negative mining and further improves the quality of VSE models . VSE++ is one of the most fundamental VSE methods that use AvgPool as the feature aggregator. Beyond the above, more research along this line focused on improving the visual or text embedding function (especially the aggregator), or designing auxiliary training objectives .

Recently, methods using BERT models for vision-language data (V++L BERTs) learns to perform rich cross-modal interaction, via tailored mechanisms such as (single/multi-headed) cross-attention . These methods typically use a BERT as the text feature extractor and learn additional cross-modal Transformers for rich cross-modal interactions. At the same time, these methods perform large-scale visual-linguistic pre-training with a collection of datasets with paired images and text (e.g., the Conceptual Caption dataset ). Comparing to this family of methods, VSE models are inferior in empirical performances as its lack of strong cross-modal interaction. However, VSE models are orders of magnitude more efficient than V++L BERTs in terms of cross-modal retrieval as the latter requires the huge BERT model to forward over all pairs of images and texts. In § 5.1.1, we show that the best VSE∞\infty can attain a close image-text matching performance to the best V++L BERT method while being much faster in large-scale multi-modal retrieval.

Experiments

We conduct experiments to validate VSE∞\infty on image-text (§ 5.1.1) and video-text matching (§ 5.1.2). We compare GPO with alternative poolings in § 5.2, and analyze the learned GPO in § 5.3. We refer to the Appendix for complete experimental details and more ablation studies.

Multi-modal retrieval is typically evaluated using the metric of recall at K (R@K), with K={1,5,10}K=\{1,5,10\}. We follow to use rsum, which is defined as the sum of recall metrics at K={1,5,10}K=\{1,5,10\} of both I→\rightarrowT (I2T) and T→\rightarrowI (T2I) retrievals, as a summarizing metric to gauge retrieval model’s overall performances. In all experiments, we set the dimensions of the positional encoding and BiGRU to be 32. Therefore, GPO has 0.1MM parameter in total, which is less than 1%1\% of the entire model.

Setup For image-text retrieval, we perform experiments on MS-COCO and Flickr30K over various feature extractors. Each image of these two datasets is associated with five text descriptions. COCO contains 123,287 images, we use the data split of where there are 113,287 training images, 5000 test images, and 5000 validation images. Flickr30K contains 31,000 images, we also use the same data split as , where there are 29,000 training images, 1000 test images, and 1000 validation images. COCO results are reported in 5K and 1K, where the 1K results are averaged over the five 1K data folds. The image feature extractors are categorized into Region feature and Grid feature following the naming convention in , where grid feature represents the feature maps from a CNN, and region feature represents object-level features from a detector.

Implementation Details The dimension of the joint embedding space is 1024. We use pre-extracted object features as the region feature (butd feature). For grid feature, the CNN backbone is fine-tuned, and we increase the resolution of input images to 512×\times512 as suggested by . We experiment with two different CNNs: (1) ResNet-101 of Faster-RCNN pre-trained on ImageNet and Visual Genome (butd) and (2) ResNeXT-101(32×\times8d) pre-trained on Instagram (wsl) . Meanwhile, we use either BiGRU or BERT-base as the text feature extractor. We refer to the Appendix for full training details and more results.

Main Results Table 2 compares VSE∞\infty with VSE baselines over different feature extractors. VSE++ is the fundamental VSE method as described in § 4, we re-implement it (denoted as Our: VSE++) and apply it on latest feature extractors (e.g., butd image features, BERT, etc.). The major difference to its original implementation is the input image size for grid feature. LIWE , VSRN , and CVSE are state-of-the-art VSE methods proposed in recent two years (we compare with more baselines in the Appendix). We use numbers directly from original papers except for CVSE, for which we re-run the official code after removing unfair additional label inputs and fixing its 1K evaluation setting (details in Appendix). Region+Grid means training two separate models with region and grid feature and averaging their similarity outputs. Over all three combinations of feature extractors, VSE∞\infty outperforms the baselines without using complicated aggregator. Besides, VSE∞\infty with WSL+BERT as feature extractors achieves the best empirical results, improving over the second best feature extractors by a large margin. VSE∞\infty is better than the baselines in both performance and simplicity. We present the COCO 5K Test results, and results with additional feature extractors in the Appendix.

Comparing VSE∞\infty with V++L BERTs We further compare VSE∞\infty with state-of-the-art V++L BERTs in Table 3. We report results on COCO 5K as the 1K results reported by V++L BERTs is computed on the first 1K fold, instead of the average result over the five 1K folds. Without large-scale V+L pre-training, our VSE∞\infty (R+G) is no worse than three out of five V++L BERTs using the same feature extractors. By using the WSL CNN to compensate for the lack of pre-training, VSE∞\infty further outperforms UNITER and gets very close to OSCAR , which is the current best V++L BERT. This is a promising result since VSE models by design do not have any fine-grained cross-modal interaction as V+L BERTs (see § 4). Meanwhile, VSE methods are orders of magnitude faster for large-scale multi-modal retrieval as the holistic embeddings can be pre-computed or indexed , and matrix multiplication is all we need to compute the compatibility score. To demonstrate this, we perform an additional text-to-image retrieval experiment with increasing size of image candidates, and visualize the model’s inference time in Figure 3. When the number of image candidates is small, we observe that VSE is a hundred time faster than V++L BERT. As the number of image candidates grows, the gap of time cost increases almost quadratically. VSE∞\infty fully exploits existing feature extractors and pushes the performance of VSE-based methods to a new height, which have significant impact in real-world problems such as image search with text query.

Evaluating VSE∞\infty with Crisscrossed Captions We evaluate our best models (with BERT and Grid features on either butd or wsl backbones) on the Crisscrossed Captions(CxC) extension of COCO, which evaluates image-text matching systems more holistically with additional intra-modal and inter-modal semantic similarity annotations. Table 5 shows that our model can significantly outperform the baseline for both inter-modality and intra-modality (on butd features). Moreover, VSE∞\infty with wsl feature can further boost the performances.

1.2 Video-text Retrieval

Setup We evaluate our method on two video datasets: MSR-VTT and VATEX . MSR-VTT contains 10,000 videos while each video has 20 text descriptions, and we use the standard split with 6573 videos for training, 2990 for testing and 497 for validation. VATEX contains 25,991 videos for training, 6000 for testing and 3000 for validation, and the 10 English descriptions for each video are used in the experiments. We splits the original validation set into new validation and testing set, each with 1500 videos, as .

Implementation Details We use ResNet-152 pre-trained on ImageNet to extract frame features for MSR-VTT and use the official I3D feature for VATEX. All implementations are based on the official code of the video-text matching method HGR , and we re-train all models. BiGRU is the text backbone for all experiments and the VSE setting is similar to 5.1.1 except that visual features are frame-level video features. Complete details are in the Appendix.

Main Results Table 4b presents the effectiveness of VSE∞\infty on video-text matching. VSE++ for video-text matching is an extension of the image-text version. HGR is the current state-of-the-art method, which employs hierarchical matching strategies. By replacing the AvgPool on frames and text with GPO, VSE∞\infty clearly outperforms VSE++ in terms of RSUM. Additionally, we change the pooling function in the global-matching branch of HGR with GPO (denoted as HGR∞\infty), and get consistent improvements.

2 Comparing GPO to Alternative Poolings

We compare GPO with several representative learnable pooling methods across four combos of visual and text feature extractors. GPO’s Size Augmentation is used in all cases for fair comparison. The baselines include:

Generalized Mean Pooling (GeM) , an adaptive pooling function with a single trainable parameter, and is popular in image search literature;

Feature-sorting Pooling (FSPool) , a learnable pooling that handles variable-sized inputs by interpolating a fixed-size learnable vector, which was proposed to encode sets in permutation-invariant manner. FSPool generates different pooling coefficients for each feature dimension.

CLS token based aggregation, which is widely used for aggregating text features in the BERT models . For BiGRU, we simply take the feature of the first token for the CLS aggregation.

Figure 4 presents the comparison in R@1 of COCO I2T, we skip T2I results since the conclusions are the same. In the left part of visual pooling, FSPool is close to GPO on grid feature, but much worse on region feature. We note that the official GeM implementation is not numerical stable and it causes gradient explosion when training with region feature and BiGRU. In the right of Figure 4, we vary over different textual pooling strategies, with BiGRU or BERT being the text feature extractor. Again, GPO outperforms all alternative pooling methods. It is worth noting that BERT’s default CLS aggregator is far from being optimal in the context of multi-modal matching. Above all, GPO is the best pooling strategy on various combinations of features, and can serve as a plug-and-play per-modality feature aggregator.

3 Visualizing and Understanding GPO

To better understand the pooling patterns learned by GPO, we visualize the learned pooling coefficients of GPO in Figure 5. On butd region feature, GPO approximates MaxPool, which is consistent with the observation in § 3. On grid feature, the coefficients are less regular, but large position indices take up most large values. We additionally observe that GPO generates non-zero coefficients for the maximum and minimum values of BiGRU features, which goes beyond the pattern of K-MaxPool. The learned pooling strategy for BERT is close to MaxPool.

Comparing GPO against Grid Search. We recall that the main motivation of GPO is to fully exploit the advantages of simple pooling functions but eliminate the repetitive manual experiments for seeking the best pooling hyperparameter. To verify how GPO address this challenge, we conduct a manual grid search over K-MaxPool with different K values for image-text matching with butd region and BiGRU features. GPO’s Size Augmentation is used here for fair comparison as it improves performance (see Appendix for details). As shown by Figure 6, the best rsum given by the grid search is 520.4, which is slightly worse than the 520.8 of the corresponding GPO entry in Table 2, which means that GPO successfully refrains us from the costly repetitive search. Note that GPO generates a pooling strategy beyond K-MaxPool for BiGRU (Figure 5), although it does not make it significantly better than the best-selected K-MaxPool.

Figure 6 shows that the best combination of pooling functions for visual and text modalities are entangled with each other. For instance, the best textual pooling function varies when the visual pooling function is changed. Therefore, a n×nn\times n search is necessary to find the optimal combinations of K, where nn is the number of grids for each modality. This could become worse when the visual feature includes multiple feature extractors (e.g., region++grid), as the search complexity can further become O(n3)O(n^{3}).

In summary, GPO keeps the effectiveness and efficiency of best-selected pooling functions, and avoids the annoying grid search process. GPO can serve as a plug-and-play aggregation module to improve VSE models.

Conclusion

In this paper, we propose the Generalized Pooling Operator (GPO), which learns to automatically adapt itself to the best pooling strategy for different data and feature backbone. As a result, we build up our VSE∞\infty by extending the standard VSE model with GPO as the feature aggregators. VSE∞\infty outperforms previous VSE methods significantly on image-text retrieval benchmarks across popular feature extractors. We further demonstrate that our VSE model achieves comparable image-text matching performances to vision++language BERT models, without visual-linguistic pre-training. Comprehensive ablation experiments confirm that GPO discovers proper pooling strategies. With simple adaptations, variants of VSE∞\infty further demonstrate effectiveness by achieving the new state of the art on two video-text retrieval datasets.

In this Appendix, we provide implementation details and experiments omitted in the main text. The content is organized as follows:

Additional implementation details, including model settings and training setups for different experiments.

Ablation studies for better understanding GPO’s effects in VSE∞\infty.

Extensions of Table 2 of the main text, with full COCO 5K evaluation results, more combinations of feature backbones, and more related baselines.

A synthetic experiment for further verifying the design choices of GPO, as well as motivating the Size Augmentation.

Experiments for exploring more complex variants of GPO (i.e., data-dependent pooling and per-dimensional pooling), which provide potential explanations for the failure of complex feature aggregators.

Appendix A Additional Implementation Details

Model settings The dimensionality d3d_{3} of the joint embedding space is 1024 for all experiments. Note when d1≠d3{d_{1}}\neq{d_{3}} or d2≠d3{d_{2}}\neq{d_{3}}, MLPs are applied before GPO to transform the dimensionality of features. The CNN backbones used in our experiments are either ResNet or ResNeXt101 , thus d1d_{1} is 2048. For text backbones, we set d2=1024d_{2}=1024 for BiGRU, and pre-trained BERT has default d2d_{2} being 768. We convert the officially released butd CNN from Caffe to Pytorch for running experiments related to Grid feature. When using Grid feature, the dimensionality transformation only uses a single linear layer since the image branch already has a massive CNN. When using Region feature, the transformation uses a two-layer MLP with residual link.

Training details The VSE models for image-text matching are trained with AdamW optimizer with weight decay factor 10e-4. The batch size is always 128 and the margin α\alpha of the triplet ranking loss is 0.2. The initial learning rate is 5e-4 while different model components have different learning rate multiplier: (1) CNNs pre-trained on ImageNet: 0.1; (2) butd or WSL CNNs: 0.01; (3) BERT: 0.1. The models are trained for 25 epochs and the learning rate decays by a factor of 10 for the last 10 epochs. When fine-tuning CNN backbones (i.e., for Grid feature), we fix the running statistics of the Batch Normalization layers during the training.

Two warm-up strategies are used during training: (1) At the first epoch, the parameter of the ConvNet is not trained (only for Grid feature), and the triplet ranking loss uses all negative examples in the batch instead of using only the hardest negative example; (2) Starting from the second epoch, all parameters are trained end-to-end with Eq. 1 of the main text, and linear learning rate warmup is used.

All experiments are implemented with PyTorch v1.2.0 and run on Tesla V100 PCI-E GPU.

Fixes of CVSE for fair comparison We notice that the official implementation of CVSE shows two unfair experimental setups compared to other image-text matching methods in the literature:

The “concept labels”(i.e., words with semantic meaning) are provided as extra input for the model, which is one of the core contributions of CVSE. However, for each text input, its “concept labels” actually come from all five ground-truth captions (in COCO and Flickr30k, each image is associated with five captions). This is an unfair information leak compared to other methods, as it makes the model leverage information from five captions to match one image. The valid “concept labels” of each caption should only contain labels from itself.

Previous image-text matching methods report COCO 1K results by averaging over five 1K data folds of the test set, but CVSE is only evaluated on the first 1K fold.

We use the official released codehttps://github.com/BruceW91/CVSE to re-train and re-evaluate the model while fixing the above two setups, and report the results in the main paper.

A.2 Video-text Matching

Model settings & Training details We follow the official implementation of HGR https://github.com/cshizhe/hgr_v2t in the video-text matching experiments and plug in our GPO as the video and text feature aggregator. The same as the image-text matching experiments, the dimension of the joint embedding space is 1024 and the margin α\alpha of triplet ranking loss is 0.2. All video-text models are trained for 35 epochs with batch size 128. The initial learning rate is 1e-4, and the learning rate decays by a factor of 10 for the last 10 epochs.

We found that re-running the video-text VSE++ baseline in the official HGR code produces obviously higher results compared to those reported in the original paper, thus we used our re-running results to compare VSE++ and VSE∞\infty.

A.3 Setups of Figure 3

To measure the speed of text-based image retrieval with VSE and V++L BERT in Figure 3, we first pre-compute all features/embeddings that are not conditioned on the text query (i.e., holistic embeddings for VSE, and butd region features for V++L BERT). Then we compute the similarity scores between the text query and all image candidates. For VSE models, all similarity scores are calculated with a large matrix multiplication, while for V++L BERT, we need to forward the BERT model for nn times where nn is the number of image candidates. We tune the batch size so that the GPO memory is fully utilized.

Appendix B Additional Experiments and Results

(1) Do we need GPO for both visual and text features?

In Table 6, we use GPO to replace the standard pooling function for butd image region feature and BERT text feature (i.e., AvgPool and CLS). The comparisons confirm that GPO is effective for either modality alone, using GPO for both modalities can further improve the results by a large margin.

(2) Effectiveness of Size Augmentation Table 7 shows the effectiveness of Size Augmentation. The random dropping of either visual and text inputs boost the multi-modal retrieval performance. A synthetic experiment in Appendix provides more motivations for this augmentation strategy. Surprisingly, we find this strategy also improves the baseline VSE++ , potentially as a regularization method. State-of-the-art video-text matching model uses a similar strategy, feature dropout, which adds a dropout layer for input features. The difference is that it randomly set feature values to 0, while Size Augmentation randomly drops entire elements from the feature set.

(3) Different choices of the sequence model in GPO We also try using different sequence model to implement GPO. Table 8 shows that using a Transformer Encoder to replace the simple BiGRU does not yield improvements on two different combinations of features. The sequence model of GPO only takes the positional information as the input without using the exact feature vectors, thus we believe it’s not necessary for the sequence model to have large capacity. A simple BiGRU will suffice for both capacity and computational efficiency, and more complex mechanisms like the multi-head attentions of Transformers could even hurt the performance.

B.2 More results and comparisons for image-text matching

In Table 9 and Table 10, we provide results extending Table 2 of the main text. Grid feature with ImageNet-pretrained ResNet-152 was the standard image backbone for image-text matching before butd region features became popular. It is worth noting that our re-implementation of VSE++ improves the original VSE++ by a large margin, by increasing the image size from 224×224224\times 224 to 512×512512\times 512 as guided by the empirical study of . VSE∞\infty consistently outperforms the improved VSE++ and other baselines on ResNet-152 grid features.

We also include results of three non-VSE methods: SCAN , CAAN , and IMRAM . These methods rely on fine-grained cross-modality interactions to match image and text, and all of them use butd and BiGRU as the feature extractors. Under the same experimental setups, VSE∞\infty outperforms them without any complex cross-modal modeling.

B.3 Synthetic Experiment for Verifying GPO Design

We further verify the architecture of coefficient generator g(⋅,⋅)g(\cdot,\cdot) using synthetic data and pre-determined pooling coefficients (e.g., coefficients for the K-MaxPool with K=5K=5). As a concrete example, we generate a set of random feature vectors as the input to GPO, and then use the pre-determined "ground-truth" pooling strategy to generate the "ground-truth" output. Such synthetic input-output pairs are then used for learning a GPO module. As for evaluation, we took the coefficient generator from GPO and compare the predicted pooling coefficients to the "ground-truth" ones. We report the results in Root Mean Square Error (RMSE).

Evaluation Protocol We designed four types of synthetic pooling patterns: (1). AvgPool (denoted as A), (2). K-MaxPool (denoted as M-K), (3). Top-K%\% Pooling (denoted as T-K%\%), and (4). Linearly-decayed pooling weights (denoted as L), in which the pooling weight for the kk-th maximum linearly decreases to zero as kk goes from 11 to N\mathsf{N}. To better assess generalization, we set the training feature set sizes to range from 20 to 100, and we evaluate models on test data with the feature set size ranging from 10 to 120. We report results on both Seen and Unseen feature set sizes to investigate GPO’s generalization performance.

Different GPO Designs We compare different design choices of the architecture for g(⋅,⋅)g(\cdot,\cdot), including:

Cos/Sin+BiGRU This is the design introduced in the main paper, which uses the positional encoding and learn a BiGRU as the coefficients generator.

Interp. Learn a fixed size vector as the pooling coefficients, and use linear interpolation to get the pooling coefficients for various lengths N\mathsf{N}. FSPool uses this type of design to handle variable-length inputs.

Cos/Sin+MLP Instead of using a sequence model to handle variable-size features, simply using a MLP to map positional encodings into pooling coefficients. This design assumes the weight generation process for each index kk is unaware of the global length.

Index+BiGRU This model transforms the position index into embeddings with a learnable matrix, and learns a BiGRU as the coefficients generator. Without the positional encoding, this design applies no prior knowledge to the ordinal position indices.

Results Table 11b presents the results comparing different design of GPOs. Cos/Sin+BiGRU achieves the best overall performances. Interp. has clear difficulty in handling K-Max Pooling. Index+BiGRU produces slightly worse results, which shows the advantage of using Cos/Sin positional encoding. Moreover, generalizing to unseen feature sizes is indeed challenging for GPO, and generalizing to smaller feature sizes is harder than generalizing to larger feature sizes. To make GPO better generalize to inputs with different sizes, we propose the Size Augmentation as discussed in § 3 of the main paper.

B.4 Complex Variants of GPO

We have kept a simple architecture for GPO so that it only adds marginal extra computational cost to VSE models. However, it is worthwhile to verify whether more complex variants of GPO can indeed produce better results. We investigate two modifications:

Are per-dimension pooling coefficients helpful? Instead of generating shared pooling coefficients for all dimensions, we try to generate coefficients for each dimension separately. Per-dimensional pooling has stronger capacity and might improve the performance, like how FSPool is designed. However, results in Table 12 disproves the above statement in the context of VSE∞\infty. The per-dimension variant of GPO provides no improvements over two combinations of feature extractors, potentially due to over-fitting.

Would GPO be better if g(⋅,⋅)g(\cdot,\cdot) also takes feature as input? Another possible modification to GPO is to input both the feature itself and the position index into the coefficients generator g(⋅,⋅)g(\cdot,\cdot). With this modification, GPO can be considered as a special form of self attention. By intuition, this modified GPO can adaptively change the pooling coefficients according to the exact feature values. However, Table 13 shows that this modification does not bring improvements. Position index along suffices for generating good pooling coefficients.

In § 3.1 of the main paper, we observe that complex aggregators cannot outperform well-selected simple pooling function, and the above two experiments again show that complicated feature aggregation does not necessarily improve VSE models. A possible explanation for these experimental results is that: feature extractors have provided adequate information for multi-modal matching, so the feature aggregators do not have to further contextualize the feature vectors. Too complicated models for feature contextualization might increase the risk of over-fitting and hurt the performance at the end.

References