Cross Modal Retrieval with Querybank Normalisation

Simion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu, Samuel Albanie

Introduction

As the improving price-performance of hardware underpinning sensors, storage and networking continues to enable the expansion of humanity’s digital archives, the capacity to efficiently search data takes on greater commercial and scientific importance. An appealing way to search such data is via natural language queries, in which the user describes the target of their search exactly as they would to another human, rather than employing specialised database languages such as Structured Query Language (SQL).

Towards this goal, a rich body of research literature has studied the problem of cross modal retrieval, the task of searching a gallery of samples in one modality given a query in another. In particular, there has been significant progress in recent years for systems that can efficiently search images , audio and videos with natural language queries by employing cross modal embeddings.

The dominant cross modal embedding paradigm employs deep neural networks that project modality-specific samples into a high-dimensional, real-valued vector space in which they can be directly compared via an appropriate distance metric. A key challenge for such methods, intrinsic to such high-dimensional spaces, is the emergence of “hubs” —embedding vectors that appear amongst the nearest neighbour sets of disproportionately many other embedding vectors (Fig. 1, left). To illustrate this challenge, we show empirically in Sec. 3.2 and Fig. 2 that hubness is prevalent among a range of leading retrieval methods. Hubs have consequences: if left unaddressed, they lead to a significant degradation in the search ranking yielded by a retrieval system . The hubness problem has received considerable attention and a number of approaches have been proposed to address it , with notable contributions in the NLP literature focusing on bilingual word translation . One contribution of our work is to show how each of these methods can be interpreted within a single unifying conceptual framework termed Querybank Normalisation (QB-Norm, Fig. 1, right), that employs a querybank of samples during inference to reduce the influence of hubs in the gallery. We observe that existing methods have two challenges: (1) To date, these approaches have only been shown to work with concurrent access to multiple test queries—an assumption that is impractical for real-world retrieval systems; (2) They are sensitive to querybank selection, and indeed actively harm performance for certain querybanks (Tab. 2). To address the first challenge, we demonstrate through careful experiments (Tab. 1) that QB-Norm does not require concurrent access to test queries to be effective. To address the second challenge, we propose a new normalisation method, Dynamic Inverted Softmax (DIS), that operates as a module within the QB-Norm framework. We show that DIS provides effective normalisation, yet is more robust than prior approaches .

We make the following contributions: (1) We motivate our study by demonstrating that the longstanding problem of hubness remains a significant concern in modern cross modal embeddings for retrieval; (2) We propose Querybank Normalisation (QB-Norm), a simple non-parametric framework that brings significant gains in retrieval performance without requiring model fine-tuning; (3) We provide the first (to the best of our knowledge) demonstration that Querybank Normalisation methods retain their effectiveness for cross modal retrieval with no access to test queries beyond the current query; (4) We propose the Dynamic Inverted Softmax, a novel normalisation method for Querybank Normalisation that is more robust than prior literature; (5) We show that QB-Norm is highly effective across a broad range of tasks, models and benchmarks.

Related work

In this section, we summarise prior work from the literature that relates to our approach, focusing on cross-modal retrieval, external memory banks and hubness.

Cross-modal representations. Following initial studies in psychology , early frameworks for cross-modal retrieval included Gaussian Mixture Models modelling translation via EM , Topic Models , CCA , KCCA and rank optimisation . Motivated by the successes of deep metric learning and deep visual semantic embeddings , there has since been a Cambrian explosion of cross-modal embedding methods for text-image retrieval , text-video text-audio , image-audio and combinations of all the above . Recent research spanning these tasks has explored large-scale pre-training , domain adaptation and tight integration of multiple sensory modalities into one side of the embedding space .

Similarity search for retrieval: Tricks of the trade. A plethora of techniques have been developed to support and enhance similarity search for retrieval, including k-d trees , re-ranking , query expansion , vector compression schemes based on binary codes and quantization that help address the curse of dimensionality . Algorithms have been developed for approximate k-nearest neighbour graph construction on CPUs and GPUs, with the latter drawing on product quantization techniques to scale up to billion-scale searches.

Differently from the work on cross modal representations and improved similarity search described above, we focus specifically on tackling the problem of hubness in cross-modal embeddings, which we demonstrate (Sec. 3.2) to be a widespread issue among leading cross-modal embedding frameworks.

Memory bank augmented architectures. Memory banks in various forms have been studied as useful extensions to neural network architectures to facilitate general problem-solving , better image captioning and summarisation , enhance self-supervised training dynamics and to provide a mechanism to deal with rare instances . Our proposed Querybank Normalisation framework likewise stores embedding samples in an external memory bank, but targets a very different problem to these works, namely hubness mitigation.

The Hubness Problem. The hubness problem was formally characterised by Radovanovic et al. , who observed that in points sampled from a distribution with high intrinsic dimensionality, the distribution of “k-occurrences” (the number of times a point appears in the k nearest neighbours of other points) skews heavily to the right. Although there is disagreement about the cause of hubness , it has been conceptually linked to distance concentration in high-dimensions (high-dimensional points lie close to a hypersphere centred on the data mean, i.e., they all exhibit a similar distance to the mean ). It is thought that hubs then result from this phenomenon through the non-negligible variance in the distribution of distances to the mean in finite dimensions .

Hubness Mitigation. One paradigm has focused on rescaling the similarity space to account for asymmetries in nearest neighbour relations —a process that can be achieved through both local and global scaling schemes. Another work has focused on addressing the hub-like tendency of centroids in the data through Laplacian-based kernels and centring . Fedbauer et al. provide a comprehensive empirical comparison of these families of methods and note that while effective, these approaches scale quadratically, making their naive application unsuitable for large datasets. One exception is the CENT method , however we did not find this approach to be effective (experiments are provided in the supplementary). In the zero-shot learning literature, works have sought to address hubness by mapping (text) targets back into the (image) query space , and by minimising proxies for hubness and skewness in the k-occurrences distribution to improve 3D few-shot learning performance. More closely related to our work, propose general retrieval schemes for which queries are matched with targets for which they form the nearest neighbour. This work was built upon the NLP literature by , who propose a cross-domain local scaling scheme (which can be integrated into the loss ), and by , who introduce the Inverted Softmax (IS) to mitigate hubness when translating between dictionaries in different languages. We discuss the relationship of our approach to in more detail in Sec. 3 and compare these methods with our proposed Dynamic Inverted Softmax in Sec. 4.

Also related to our work, enforce a bipartite matching constraint between queries and test set items by applying an IS over the full set of test queries—a constraint that is unrealistic for practical retrieval systems that experience continuous operation from users. One contribution of this paper is to demonstrate that concurrent access to test queries is not required. A second contribution of our work, not considered in prior work, is to show that the techniques proposed above can actively damage retrieval performance for particular querybank selections, an issue that we address with our proposed Dynamic Inverted Softmax.

Method

We first define the task of retrieval with cross modal embeddings (Sec. 3.1), before outlining the motivation for our work by examining the hubness problem in the context of text-video retrieval (Sec. 3.2). Next, we introduce the Querybank Normalisation framework (Sec. 3.3), generalising several existing approaches to address this issue. Finally, we explore designs for framework components and introduce the proposed Dynamic Inverted Softmax for robust similarity normalisation (Sec. 3.4).

The choice of similarity measure used to define a “good match” is determined by the application domain. For instance, in the task of text-video retrieval with natural language queries, the objective is to rank a gallery of videos according to how well their content is described by a written free-form text query , whereas in image-audio retrieval the objective is typically to obtain audio samples from the gallery that share the same semantic category as the image query . In this work, we focus particularly on cross modal retrieval tasks with natural language queries, for two reasons: (1) these tasks have received limited attention in the hubness mitigation literature, (2) hubness has been shown to be particularly prevalent in embeddings with high intrinsic dimensionality . Since natural language queries can express more complex concepts than individual words (such as those considered in zero-shot learning image labelling tasks ), we expect might expect natural language queries to naturally induce cross modal embeddings with greater intrinsic dimensionality, and thus may have greater potential to benefit from hubness mitigation.

2 Motivation

It has long been observed that high-dimensional embedding spaces are prone to hubness , in which a small proportion of samples appear disproportionately frequently among the set of k-nearest neighbours of all embeddings. As noted by Berenzweig , this property can have damaging consequences for retrieval systems that employ nearest neighbour search to find the best gallery match for a given query. To illustrate this issue, we consider the problem of video retrieval with natural language queries. We plot the distribution of the number of times each gallery video was retrieved on the MSR-VTT retrieval benchmark for an array of text-video retrieval methods, including CE , TT-CE+ , MMT and CLIP2Video , the latter of which represents the current state of the art on this benchmark. In each case, we see striking evidence of hubness—a small number of videos are retrieved extremely often, while others are not retrieved at all. This phenomenon is not limited to a particular retrieval model, suggesting that the issue is not readily addressed by the use of multiple video modalities, attention mechanisms and large-scale pretraining implemented in various combinations by these approaches.

3 Querybank Normalisation

To address the hubness issues observed among cross modal embeddings for text-video retrieval in the previous section, we first turn to the existing literature on hubness mitigation. As noted in Sec. 2, hubness effects have been studied in several problem domains, including Zero-Shot Learning , NLP , biomedical statistics and music retrieval . Among this literature, we are particularly interested in methods that can be applied in a practical cross modal retrieval setting, namely, those methods whose complexity scales at most linearly with the size of the gallery (rather than quadratic complexity methods that seek to address hubness within a fixed embedding space ). To clarify relationships between existing approaches, we cast them into the Querybank Normalisation framework (Fig. 1), which comprises two components, querybank construction and similarity normalisation, described next:

Querybank construction. To mitigate hubness in the cross modal embedding space, we seek to alter the similarities between embeddings in a way that minimises the influence of hubs. To adjust similarities, we first construct a querybank of NN samples, B={b1,…,bN}\mathcal{B}=\{b_{1},\dots,b_{N}\} from the query modality, mqm_{q}, which will serve as a probe to measure the hubness of gallery samples.

In practice, the probe matrix employed for similarity normalisation can be precomputed and re-used across all queries (improving computational efficiency at the cost of higher memory). An overview of the resulting QB-Norm algorithm, and its application to ranking gallery samples for a collection of queries, Q\mathcal{Q}, is summarised in Alg. 1.

4 Design choices

The Querybank Normalisation framework admits a number of viable choices for both querybank construction and similarity normalisation. To illustrate this point, we first cast three techniques for hubness mitigation proposed in the NLP literature into the framework. We then introduce our proposed alternative, the Dynamic Inverted Softmax.

Inverted Softmax (IS) . Targeting bilingual word translation, this method constructs a querybank from the source vocabulary (corresponding to all possible queries of interest). For practical implementations, the authors recommend to uniformly randomly subsample a feasible number of queries. Similarity normalisation is implemented via:

where exp⁡[⋅]\exp[\cdot] denotes elementwise exponentiation and β\beta is a hyperparameter referred to as the “inverse temperature”.

Dynamic Inverted Softmax (DIS). In experiments with the methods described above (discussed in detail in Sec. 4) we observed an important practical issue: if the querybank does not effectively cover the space containing the gallery, performance is severely degraded such that it falls below the performance of unnormalised similarities. This characteristic renders them less desirable for a general-purpose solution: we would like something that not only enhances performance in favourable conditions, but also “does no harm” when curating a querybank to match the gallery closely is challenging. To address this issue, in addition to the querybank probe matrix described in Alg. 1, we also precompute a gallery activation set, A={j:j∈argmax⁡\mboxkls(bi,gl),i∈{1,…,N}}\mathcal{A}=\{j:j\in\stackrel{{\scriptstyle\mathclap{\scriptsize\mbox{k}}}}{{\operatorname*{argmax}}}_{l}s(b_{i},g_{l}),i\in\{1,\dots,N\}\}. Here, the notation argmax⁡\mboxklf(l)\stackrel{{\scriptstyle\mathclap{\scriptsize\mbox{k}}}}{{\operatorname*{argmax}}}_{l}f(l) denotes the kk-max-select operator that returns the kk values of ll that maximise f(l)f(l) (like jj, ll also runs over the gallery indices and kk is set as a hyperparameter). Intuitively, this set contains the indices of gallery vectors that our querybank probe has identified as potential hubs. We create a Dynamic Inverted Softmax by activating the inverted softmax only for nearest neighbour retrievals that fall within this set:

Since sq(j)s_{q}(j) is computed as an intermediate step in Eqn. 1, the only additional cost incurred by the Dynamic Inverted Softmax over the standard Inverted Softmax stems from the argmax operation in Eqn. 2. Fortunately, this computation can be performed extremely efficiently with almost no loss in precision, even at the scales of billions of gallery samples . We show through experiments in Sec. 4, the Dynamic Inverted Softmax is significantly more robust than GC, CSLS and IS: crucially, it does not harm performance when employed with suboptimal querybank selection.

Experiments

In this section, we first briefly describe the datasets and metrics used for our experiments (Sec. 4.1). We then conduct a series of experiments that: (i) demonstrate our claim that QB-Norm is effective without concurrent access to more than one test query; (ii) investigate the influence of querybank size; (iii) compare the Dynamic Inverted Softmax against prior methods; (iv) ablate other QB-Norm components (Sec. 4.2). Finally, we demonstrate the generality of Querybank Normalisation by applying it to a broad range of models, tasks and datasets (Sec. 4.3).

We conduct experiments on standard benchmarks for text-video retrieval: MSR-VTT , MSVD , DiDeMo , LSMDC , VaTeX and QueryYD . We also investigate QB-Norm on text-image retrieval (MSCoCo ), text-audio retrieval (AudioCaps ), and image-to-image retrieval (CUB-200-2011 , Stanford Online Products ). Detailed descriptions of each dataset are deferred to the supplementary. We report standard retrieval performance metrics: R@K (recall at rank K, higher is better) and MdR (median rank, lower is better). For each study, we report the mean and standard deviation over three randomly seeded runs.

2 Querybank Normalisation

We conduct initial studies on the MSR-VTT benchmark for text-video retrieval using TT-CE+ to address a series of questions relating to Querybank Normalisation.

Do we need access to more than one test query at a time to mitigate hubness? Prior work has investigated the use of IS for image and video retrieval with natural language queries, but only by assuming simultaneous access to the full test set of queries to construct the querybank . The motivation for this approach is to enforce a bipartite matching constraint that encodes the prior knowledge that each test query maps to exactly one gallery sample. Unfortunately, this approach is impractical to deploy for real world systems that experience sequential user queries. Therefore, we first ask whether we require concurrent access to all test set queries by constructing an alternative querybank from the training set. We evaluate performance with QB-Norm using DIS normalisation in which we construct querybanks from: (i) all test set queries; (ii) all validation set queries: (iii) a randomly subsampled subset of the training set matching the size of the test set (resampled once for each trained model to estimate variance). The results are reported in Tab. 1. Remarkably, we observe that training set querybanks perform comparably to test set querybanks. Given this finding, we conclude that test set querybanks are not necessary to mitigate hubness. We therefore restrict all querybank construction to use training set samples for all remaining experiments, ensuring valid comparisons on standard retrieval benchmarks.

What is the influence of querybank size on performance? To address this, we sample querybanks across a range of different scales, and report mean and standard deviations across metrics for three samplings of each scale using DIS normalisation. The results are shown in Fig. 3 (left), where we observe that performance increases with querybank size, but strong results can be obtained with a querybank of just a few thousand random training samples.

What is the influence of the similarity normalisation strategy on QB-Norm? To address this question, we first sample querybanks of 5,000 samples from the MSR-VTT training split and compare the normalisation strategies described in Sec. 3.4. Results are reported in the upper block (“In Domain”) of Tab. 2 where we observe that CSLS , IS and the proposed DIS strategy perform best, and that all querybank normalisation methods substantially outperform the baseline without normalisation. Next, to evaluate the robustness of the normalisation strategies to different querybank sampling distributions, we sample additional querybanks of 5,000 samples from the training splits of two different video retrieval datasets: MSVD (whose query domain closely matches MSR-VTT), and LSMDC (a collection of movies with audio descriptions, whose query domain is further away from MSR-VTT), and evaluate retrieval performance on MSR-VTT test. We report results in the middle blocks of Tab. 2 (“Close Domain” and “Far Domain”) where we observe that sampling the querybank from a closely overlapping domain (MSVD) works well for all methods (with DIS performing best), but that sampling from a different domain (LSMDC) degrades performance below the baseline without normalisation for all methods except GC and DIS.

To understand why the LSMDC querybank could be actively harmful for methods other than GC and DIS, we studied the samples closely and observed that LSMDC queries retrieve only a small subset of videos from the video gallery (and thus were ineffectual at their primary purpose of probing for hubs). To validate that this retrieval distribution was indeed the cause of the issue, we constructed an “adversarial” querybank from MSR-VTT by selecting the 5,000 training queries that achieved the smallest coverage (i.e. retrieved the lowest number of distinct videos) over the MSR-VTT test set. We report numbers in the Adversarial block of Tab. 2. We observe that despite sampling from the same dataset, all normalisation methods other than DIS are significantly harmed. In the lower block, Overall, we present the overall performance computed as geometric mean for all methods. Since DIS performs the best overall (presented in bold in Tab. 2), we use it as our normalisation strategy for QB-Norm for all remaining experiments.

Hyperparameter sensitivity. The IS and DIS normalisation strategies require the user to select an additional hyperparameter (the inverse temperature) that is absent from other methods. We evaluate the sensitivity of DIS to this hyperparameter in Fig. 3 (right), where we find that a value of 20 works best. In practice, we found that this value worked well consistently across datasets, and therefore we use it for all remaining experiments (with the exception of CLIP2Video where we used 1.99−11.99^{-1}, since the similarities are already scaled by the method). DIS normalisation introduces an additional hyperparameter (the kk maximum selection value described in Sec. 3.4). We observed that choosing k=1k=1 offers a good trade-off between good performance and robustness, so we simply use this value for all experiments.

Does QB-Norm mitigate hubness? The core motivation for QB-Norm is that existing cross modal retrieval methods are heavily affected by hubness (Fig. 2). To investigate whether this has been addressed by QB-Norm, we report the skewness of the k-occurrences distributionA detailed description of this calculation is given in the supplementary. (which indicates the hubness of an embedding space ) for four datasets in Tab. 3 using a querybank consisting from all the samples from the training set. We observe that in each case, skewness (and hence hubness) is significantly reduced.

3 Comparison with other methods

In this section, we conduct an extensive study to evaluate the effectiveness and generality of QB-Norm on several well established benchmarks.

The influence of applying QB-Norm to cross modal embeddings for text-video retrieval are reported in Tab. 4, 5, 17, 18, 8, 9. We provide further text-video retrieval results in the supplementary. In Tab. 10 we report results for the text-image retrieval task, while in Tab. 11, 12, we report results for the image-image retrieval task. Finally, in Tab. 13, we report results for text-audio retrieval. In Fig. 8 we also show a qualitative example. For the base models that provide weights for different seeds we report mean and standard deviation of QB-Norm applied on each seed. In each case, QB-Norm brings a significant improvement over all tested methods, benchmarks and tasks. We show in bold the best performing method.

Limitations and societal impact

Limitations All the normalisation techniques used with QB-Norm incur additional pre-computation costs. The proposed normalisation technique, DIS, adds a further small additional computational cost over other normalisation approaches. For a full discussion on complexity, please refer to the supplementary. We also show in Tab. 2 that adversarial querybank selection and significant domain gaps can reduce the benefits of Querybank Normalisation.

Societal impact Cross modal retrieval is a powerful tool with both positive applications and risks of harm. Cross modal search enables efficient content discovery for researchers, musicians, artists and consumers. However, this capability also lends itself to tools of political oppression: for example, it could enable efficient searching of social media content to discover signs of political dissent.

Conclusions

In this work, we introduced the Querybank Normalisation framework for hubness mitigation. We also proposed the Dynamic Inverted Softmax for robust similarity normalisation. We demonstrated its broad applicability across a range of tasks, models and benchmarks. Acknowledgements. This work was supported by Adobe, Google and Zhejiang Lab (NO. 2022NB0AB05), and a G-Research travel grant. The authors thank Bruno Korbar for useful suggestions, Andrew Zisserman and Jenny Hu for their support, and Adam Berenzweig for kindly sharing a copy of his earlier work. S.A. would like to acknowledge Z. Novak, N. Novak and S. Carlson in supporting his contribution.

References

Appendix A Dataset details

In this section, we describe the splits and datasets employed for all tasks considered in this work.

For the task of text-video retrieval we test our approach on seven current benchmarks.

MSR-VTT contains around 10k videos, each having 20 captions. For the task of text-video retrieval, we follow prior works and we report results on the official split (full) which contains 2,990 videos for testing and 497 for validation. Since a number of recent works also report results on the 1k-A split, we compare against these method on this split as well. The 1k-A split contains 1,000 videos for testing and around 9,000 for training. We use the same videos and captions as defined in which are used by other works for evaluation. We report the results using models trained for 100 epochs.

MSVD has 1,970 videos and around 80k captions. We report results on the standard split using in prior works which consists of 1,200 videos for training, 100 for validation and 670 for testing.

DiDeMo has 10,464 videos. They are collected from a large-scale creative commons collection and are varied in content (concerts, sports, pets etc.). For each video, there are 3-5 pairs of descriptions. For the task of text-video retrieval, we use the paragraph video retrieval protocol as defined in prior works . This means that we the split consisting of 8,392 for training, 1,065 validation and 1,004 test videos.

LSMDC contains 118,081 short video clips extracted from 202 movies. Each clip has a textual description which consist in a caption which is extracted either from the movie script or transcribed from descriptive video services (DVS) for the visually impaired. We use the official splits as defined in the Large Scale Movie Description Challenge (LSMDC). The testing split contains 1,000 videos.

VaTeX contains 3,4911 videos and has multilingual captions in Chinese and English. Each video has 10 captions for each language. As for the other datasets, we follow the same protocol as defined in prior works and use 1,500 videos for testing, while there are 1,500 videos for validation. Please note that in this work, we use only the English annotations.

QuerYD has 1,815 videos for training, 388 for validation and 390 for testing. The videos are extracted from YouTube and are varied in content. The dataset has 31,441 textual descriptions. 13,019 of these are precisely localized in the video with start time and end time annotations while the other 18,422 are coarsely localized. In this work, we do not use the localization annotations and report results on the official splits following prior work on text-video retrieval .

ActivityNet contains 20k videos and has around 100K descriptive sentences. The videos are extracted from YouTube. We use a paragraph video retrieval as defined in prior works . We report results on the val1 split. The training split consists of 10,009 videos, while there are 4,917 videos for testing.

A.2 Text-image retrieval

For text image retrieval, we report results on the MSCoCo dataset. It consists of 123k images with 5 captions for each sentence. We report results for the 5k test split.

A.3 Text-audio retrieval

For text audio retrieval, we report results on the AudioCaps dataset which comprises sounds with event descriptions. We use the same setup as prior work where 49,291 samples are used for training, 428 for validation and 816 for testing.

A.4 Image-to-image retrieval

CUB-200-2011 contains 11,788 images with 200 classes. The training split consist of the first 100 classes (5,863 images) while the testing split contains the remaining classes (5,924 images). We use the same setup as used in prior work .

Stanford Online Products contains 120,053 images with products from 22,634 classes. We use the provided train and test splits containing 59,551 and 60,502 images respectively, as used in prior works .

Appendix B The influence of the Top-k hyperparameter on DIS normalisation

In Tab. 14 we show the influence of kk in the Top-k selection employed when constructing the gallery activation set (introduced in Sec. 3.4 of the main paper). We observe that choosing k=1k=1 offers a good trade-off between good performance when constructing In Domain querybanks and robustness when constructing Far Domain querybanks. We therefore use k=1k=1 for all reported experiments.

Appendix C Can effective querybanks can be constructed from the training set with IS normalisation?

In the main submission, we showed that effective querybanks can be constructed from the training set when employing DIS normalisation. Here, we show that this property also applies to IS normalisation, supporting our hypothesis that Querybank Normalisation has the general property of not requiring concurrent access to multiple test queries for appropriate normalisation strategies. In Tab. 15 we report the results of selecting queries from training, validation or testing split to form the querybank when employing IS normalisation. Similarly to DIS, we observe that training set querybanks perform comparably to test set querybanks for IS normalisation.

Appendix D The influence of embedding dimensionality on the effectiveness of QB-Norm

Radovanovic et al. posit that hubness is a phenomenon that is: (i) inherent to high dimensional spaces; (ii) heavily influenced by the intrinsic dimensionality of the data. To investigate these perspectives, we study the improvement yielded by QB-Norm over embeddings of different dimensionality, reporting results in Fig. 5 (left). We observe that QB-Norm brings around the same gain when changing the embedding size. We can interpret this finding within the framework of as making the statement that changing the shared embedding dimensionality does not influence intrinsic dimensionality. To provide further analysis, we make a crude approximation to increasing/decreasing intrinsic dimensionality by increasing/decreasing the number of modalities employed in the video embedding. Intuitively, since audio provides a different “view” of a sample to visual data, we expect a joint embedding with access to more modalities to exhibit higher intrinsic dimensionality than one with only visual cues. We plot the effect of these changes in Fig. 5 (right). We observe a slight increase in performance gain when applying QB-Norm with an increased number of modalities, which accords with the Radovanovic hypothesis.

Appendix E Additional text-video retrieval results

In Tab. 16, 19 we report additional comparisons with state of the art on the MSR-VTT full split as well as ActivityNet . In both cases, we observe that QB-Norm yields improvements. We also explore the use of QB-Norm with CLIP4Clip —for this, we train models using the code made available by the authors. For CLIP4Clip experiments, we use a β\beta value of 0.450.45 with the exception of LSMDC where β\beta is 1.26−11.26^{-1}.

Appendix F The computational complexity of normalisation strategies

As discussed in the main paper in Sec.3.4, we use various normalization techniques in conjunction with QB-Norm. In this section, we describe the computational cost of each technique in the context of its influence on inference time. For clarity of exposition, we consider exact similarity searches, but note that in practice approximate nearest neighbour implementations are employed for large-scale deployments . All strategies incur an initial cost that corresponds to pre-computing the similarity between a test query and all the videos from the gallery, O(N)\mathcal{O}(N), where NN represents the number of videos in the gallery. We further assume that we have pre-computed and stored all similarities between each query in the querybank and videos from the gallery. This assumption incurs both computational and storage costs of O(NM)\mathcal{O}(NM), where MM represents the number of queries in the querybank.

Globally-Corrected (GC) retrieval involves determining the rank of the test query with respect to the querybank for each gallery item. Since we assume that we have pre-computed similarities between the querybank and the gallery, we also pre-compute an initial ranking over querybank elements for each gallery item. For each test query, we establish its rank amongst the querybank for every target item by performing a binary search over the sorted list of pre-computed similarities. This incurs an inference time cost of O(Nlog⁡M)\mathcal{O}(N\log M).

Cross-Domain Similarity Local Scaling (CSLS) consists of finding the most similar queries from the querybank for each gallery video and finding the KK gallery videos (here KK is a hyperparameter of CSLS) that are most similar to the test query. For the former, we can pre-compute, for each video in the gallery, the KK most similar queries from the querybank and store the average similarity into a vector of size NN. For the latter, we must compute (during inference) the average similarity of the KK most similar items among the gallery to our test query. Using quickselect, this can be done in O(N)\mathcal{O}(N) time on average (note that we do not require the top KK element similarities to be sorted, since they will be averaged).

Inverted Softmax (IS) involves normalizing the final similarity by the sum of the similarities given the querybank. However, the softmax denominator can be pre-computed by summing the querybank similarities for each gallery item and storing the results into a vector of size NN. During inference the similarities are divided by this pre-computed sum, which adds only constant-time overhead. Pre-computing the sum in this manner also reduces the storage cost associated with the querybank from O(NM)\mathcal{O}(NM) to O(N)\mathcal{O}(N) (since we can discard the memory allocated to store the similarities between each query in the querybank and each video in the gallery).

Dynamic Inverted Softmax (DIS). Since DIS involves applying IS dynamically, the computation of the normalization for each test query is done in constant time as described above for IS. The additional gallery activation set employed by DIS can be pre-computed and stored for an additional O(N)\mathcal{O}(N) storage cost. There is an additional cost during inference: the top-11 search to determine the video originally retrieved by the test query (which determines whether normalisation is performed). This can be done in linear time (O(N)\mathcal{O}(N)).

Appendix G Comparison to CENT

In Tab. 20 we show how CENT normalisation performs in comparison to an unnormalised baseline and Querybank Normalisation with DIS. Since we found CENT to consistently harm performance for cross-modal retrieval, we did not include it in all experiments in the main paper.

Appendix H Hubness and Skewness

We use the skewness metric as defined in to measure hubness:

where μNk\mu_{N_{k}} and σNk\sigma_{N_{k}} are the mean and standard deviation of NkN_{k}. NkN_{k} represents the k-occurrence distribution and is defined as follows Nk(x)=∑i=1npi,k(x)N_{k}(\mathbf{x})=\sum_{i=1}^{n}p_{i,k}(\mathbf{x}) where

Here x\mathbf{x} represents a video embedding and qi∈Qq_{i}\in Q a set of queries. To compute these statistics, In practice, we use the we use k=10k=10, following for the k-occurences distribution, employing the implementation of .

As shown in Tab. 3 in the main paper, skewness and hence hubness is reduced after applying QB-Norm. The same can be seen in Fig. 6 which depicts the distribution of number of times each video is retrieved before and after using QB-Norm. We observe that the maximum number of times a video is retrieved is reduced, indicating a hubness reduction.

Appendix I Additional ablations on other metrics

In the main paper, to maintain conciseness we report ablation plots for the influence of querybank size and inverse temperature using the geometric mean of R1, R5 and R10. For completeness, in this section we show results on each metric individually. As seen in Fig. 7, the individual metrics reflect the trend shown for the geometric means, aligning with the results shown in the main paper.

Appendix J Video and text embeddings (experts) description used for video retrieval

For this work, we used the pretrained weights provided by TT-CE+ and CE+ (https://github.com/albanie/collaborative-experts). These models use a set of pretrained experts. Below, we summarise how these experts were extracted.

Two action experts are used: Action(KN) and Action(IG). Action(KN) is a 1024-dimensional embedding produced by an I3D architecture trained on Kinetics . The embeddings are extracted from frame clips at 25fps and center cropped to 224 pixels. For Action(IG) the model is a 34-layer R(2+1)D , trained on IG-65m

Two forms of object experts: Obj(IN) and Obj(IG). For extracting Obj(IN) a SENet-154 model trained on ImageNet was used. For extracting Obj(IG) a ResNext-101 model trained on Instagram data with weakly labelled hashtags was used. Both of the embeddings are extracted at 25fps.

For producing an audio expert a VGGish model trained from audio classification on the YouTube-8m dataset was used.

For the scene expert a DenseNet-161 pretrained on Places365 was used. The scene embedding has a 2208 dimension.

For the speech expert, the Google Cloud API (to transcribe the speech content) is used.

For the text we use GPT2-xl finetuned as provided by the authors. The size of the final pre-trained embedding is 1600.

For CLIP2Video we used the model as it is provided online https://github.com/CryhanFang/CLIP2Video. The model receives as input the raw frames and raw queries. For CLIP4Clip , we use the online code https://github.com/ArrowLuo/CLIP4Clip and re-train the model for each dataset where we present results since weights are not available online. For the other tasks we followed the instructions given on the official repositories. For MMT-Oscar we used the pretrained weights and the features provided at https://github.com/UKPLab/MMT-Retrieval. For RDML we used the models provided at https://github.com/Confusezius/Deep-Metric-Learning-Baselines. For audio retrieval we used the pretrained weights and models provided at https://github.com/oncescuandreea/audio-retrieval.

Appendix K v2t performance metrics

In Tab. 21 we report metrics indicating the performance of QB-Norm on the reverse task of video-text retrieval (in which videos are used as queries to retrieve descriptions). We apply QB-Norm with DIS normalisation using all videos from the training split to construct the querybank. We observe that QB-Norm yields a striking boost in performance.

Appendix L Qualitative results

In Fig. 8 we provide some qualitative examples, illustrating cases for which the QB-Norm model correctly retrieves videos that are not retrieved without QB-Norm. Examining failure cases, we found qualitative examples for which the retrieval ranking produced with QB-Norm was more “reasonable” (as shown in the bottom set of Fig. 8). However, in line with prior work suggesting that hubness is a property of the distribution (rather than driven by individual samples), we did not observe consistent, obvious qualitative trends among the samples that were corrected, or remaining failure cases. As an example, we observed gains for queries with both shorter and longer, highly descriptive captions.