TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval

Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, Yang Liu

Introduction

The focus of this work is text-video retrieval—the task of identifying which video among a pool of candidates best matches a natural language query describing its content. Video search has a broad range of applications across domains such as wildlife monitoring, security, industrial process monitoring and entertainment. Moreover, as humanity continues to produce video at ever-increasing scale, the ability to perform such searches effectively and efficiently takes on critical commercial significance to video hosting platforms such as YouTube.

A central theme of recently proposed retrieval methods has been the investigation of how to best use multiple video modalities to improve performance. In particular, architectures based on mixtures-of-experts and multi-modal transformers have shown the benefit of making use of diverse sets of pre-trained models for related tasks (such as image classification, action recognition and ambient sound classification) as a basis for video encoding during training and testing.

In this work, we explore whether commensurate gains could be achieved by leveraging multiple text embeddings learned on large-scale written corpora. Different from video embeddings using multiple modalities and pretraining tasks, it is less obvious that there is sufficient diversity among collections of text embeddings to achieve a meaningful boost in performance. In fact, our inspiration stems from a careful investigation of the performance of different text embeddings across a range of retrieval benchmarks (Fig. 2). Strikingly, we observe not only that there is considerable variance in performance across text embeddings, but also that their ranking is not consistent, strongly supporting the idea of using multiple text embeddings.

Motivated by this finding, we propose a simple algorithm, TeachText, to effectively exploit the knowledge captured by collections of text embeddings. Our approach requires a \saystudent model to learn from a single or multiple \sayteacher retrieval models with access to different text embeddings by distilling their text-video similarity matrices into an enhanced supervisory signal. As shown in Fig. 1, TeachText is capable of delivering a significant performance gain. Moreover, this gain is complementary to that of adding more video modalities to the video encoder but importantly, unlike the addition of video modalities, does not incur additional computational cost during inference.

Our main contributions can be summarised as follows: (1) We propose the TeachText algorithm, which leverages the additional information given by the use of multiple text encoders; (2) We show that directly learning the retrieval similarity matrix between the joint query video embeddings, which to the best of our knowledge is novel, is an effective generalized distillation technique for this task (and we compare our approach to alternatives among prior work such as uni-modal relationship distillation ); (3) We show an application of our approach in eliminating noise from modern training datasets for the text-video retrieval task; (4) We demonstrate the effectiveness of our approach empirically, achieving state of the art performance on six text-video retrieval benchmarks.

Related Work

Video retrieval methods. The task of indexing video content to enable retrieval has a rich history in computer vision—sophisticated systems have been developed to find specific objects , actions , predefined semantic categories , irregularities and near-duplicates . In this work, we focus on the task of retrieving content that matches a given natural language description. For this particular task, there has been considerable interest in developing cross-modal methods that employ a joint-embedding space for text queries and video content . These joint video-text embeddings, which aim to map videos and text descriptions into a common space such that matching video and text pairs are close together, form an attractive computational model for tackling this problem, since they allow for efficient indexing (although hierarchical embeddings have also been investigated ). Recently, two key themes have emerged towards improving the quality of these embeddings. First, large-scale weakly supervised pretraining methods have sought to expand their training data by exploiting the speech contained in the videos themselves as a supervisory signal. Second, the integration of multiple modalities (which has long been considered important for semantic indexing ) has been shown to yield significant gains in performance . We focus on candidates from this latter theme as a basis for investigating our approach.

Text embeddings. The representation of language through learned embeddings has been widely studied and applied in a variety of natural language processing applications. Several works have demonstrated that even with large-scale pretraining, there still are benefits to finetuning the models on the target task and that larger models (often employing multiple attention heads) yield higher performance . Recently, provided a detailed comparisons on the importance of language features for vision applications and proposes a word embedding that is specifically designed for vision tasks. In this work, we first study how various pretrained language embeddings affect the performance for text-video retrieval and then propose a method to take advantage of the benefits of combining multiple text embeddings.

Knowledge Distillation/Privileged Information. The purpose of knowledge distillation is to transfer knowledge from one model (teacher) to another model (student). This idea was originally introduced in the context of decision tree simplification and model compression , and later extended by who formalised this knowledge transfer as the temperature-parameterised process of knowledge distillation. The concept was further generalised in the unifying framework of generalized distillation for learning with privileged information (via similarity control and knowledge transfer ), together with knowledge distillation . Our approach distills knowledge of the similarities between video and text samples into the student and therefore represents a form of generalized distillation. While most knowledge distillation methods train the student with the teacher’s outputs as targets, more recent methods propose different approaches . Of most relevance to our approach, transfer mutual relations of data examples and propose distance-wise and angle-wise distillation losses that penalize structural differences in relations instead of training the student to mimic the output of the teacher—we compare to their approach in Sec. 5.

Motivation and intuition

Recently, points out that even though language representation learning systems (such as ) are pre-trained on vast amounts of data, they are still sensitive to slight changes in the data distribution and task specification. In this way, most systems can be viewed as narrow experts rather than competent generalists.

Consequently, in Fig. 2 we investigate how the usage of different off-the-shelf pre-trained text embeddings affects the retrieval performance. We observe that there is significant variance both within and across datasets, suggesting that each embedding captures different types of information. Our intuition is that this information comes from the diversity of architectures, pretraining datasets and pretraining objectives, which differs across the text embeddings.

Next, we give details about the used text embeddings and summarise the key differences between them in relationship with our findings. Word2vec (w2v) is a lightweight text embedding that is widely used for vision tasks . Multi-task GrOVLE (mt_grovle) , is an extension of w2v that is specially designed for vision-language tasks (in our experiments, however, we find that it slightly under-performs w2v). The finetuned transformer language model (openai-gpt) embedding is trained on a book corpus containing long stretches of contiguous text. We observe that it performs well on datasets that have longer text queries such as ActivityNet. RoBERTa and ALBERT are based on the BERT architecture and are trained on the same data which consists of unpublished books and Wikipedia articles. RoBERTa focuses on hyperparameter optimization and shows that greater model capacity leads to better performance while ALBERT proposes some parameter-reduction techniques to reduce memory consumption and increase training speed. In our experiments, we observe a high variation in performance when comparing the two. In contrast to the other embeddings, gpt2 is trained on a crawled dataset that was designed to be as diverse as possible. We observe that gpt2 performs most robustly in our experiments, especially on smaller datasets such as MSR-VTT and MSVD. However, it nevertheless exhibits a domain gap to each corpus (highlighted by the fact that performance increases when fine-tuning gpt2-xl, termed gpt2-xl-F throughout the paper, on queries from the text-video retrieval datasets).

Additionally, in Fig. 3 we show how many correctly retrieved queries are shared between three text embeddings on MSR-VTT: gpt2-xl, gpt2-xl-F and w2v. Only around 19% (R1), respectively 42% (R5) queries are correctly retrieved by all the three considered text embeddings. This means that a significantly number of queries are sensitive to the used text embedding, consolidating our intuition.

Method

Motivated by the findings from Sec. 3, our work aims to study the influence of using multiple text embeddings for text-video retrieval.

where BB represents the batch size used during training, sij=F(xi)TQ(tj)s_{ij}=F(x_{i})^{T}Q(t_{j}) is the similarity score between the encoded video F(xi)F(x_{i}) and query Q(tj)Q(t_{j}) while mm is the margin.

The key idea behind our approach is to learn a retrieval model, MM, that, in addition to the loss described above, also has access to information provided by a collection of pre-trained \sayteacher retrieval models which are trained on the same task but ingest different text embeddings.

2 TeachText algorithm

3 Learning the similarity matrix

As noted in Sec. 4.1, the essence of the retrieval task is to create a model that is able to establish cross-modal correspondences between videos and texts/queries, assigning a high similarity value to a pairing in which a query accurately describes a video, and a low similarity otherwise. This renders the similarity matrix a rich source of information about the knowledge held by the model. In order to be able to transfer knowledge from the teachers to the student, we encourage the student to produce a similarity matrix that matches an aggregate of those produced by the teachers. In this way, we convey information about texts and video correspondences without strictly forcing the student to produce exactly the same embeddings as the teachers. To this end, we define the similarity matrix distillation loss as:

where BB represents the batch size, Φ=Φ(S1,…,SN)\Phi=\Phi(S_{1},\dots,S_{N}) represents the aggregate of the teacher similarity matrices and SsS_{s} represents the similarity matrix of the student. Finally, inspired from other distillation works such as , l represents the Huber loss and is defined as

We explored several forms of aggregation function and found that a simple element-wise mean, Φ(S1,…,SN)=1N∑k=1NSk\Phi(S_{1},\dots,S_{N})=\frac{1}{N}\sum_{k=1}^{N}S_{k}, worked well in practice.

The idea of learning directly the cross-modal similarity matrix is, to the best of our knowledge novel. It draws inspiration from the work of relational knowledge distillation which considered the idea of learning from relationships and introduced two algorithms to implement this concept in a uni-modal setting through pairwise and triplet distance sampling. We compare our matrix learning approach with theirs in Sec. 5.

4 Student model

A key advantage of our approach is that it is agnostic to the architectural form of the student and teachers, and thus the student (and teachers) can employ any method from the current literature. We test our TeachText algorithm using three different recent works MoEE , CE , MMT as the student and teacher base architectures. All these works employ multi-modal video encoders for the text-video retrieval task. For more details, please consult the original paper of each method.

Establishing a stronger baseline. In addition to these models, we also investigate our approach on a model which shares the CE architecture of but includes a series of small technical improvements to provide a stronger baseline against which we also test the TeachText algorithm. Starting from this base architecture, we refine the input embedding selection, finding that the face and OCR video modalities employed by do not consistently produce improvement so we remove them as inputs to the video encoder. We update the model to use the more powerful gpt2-xl text embedding of and following , we finetune this text embedding on captions from the target dataset to bring additional improvement. Combining all of these changes (ablations provided in Sec. 5.3 and Fig. 5a) results in the CE+ model which we include as an additional baseline. Thus, in summary we use four ( and CE+) different base architectures for the student model.

5 Teacher models

The teacher models use the same architecture as the student model. Concretely, for each of the four base architectures described in Sec. 4.4, we create a pool of multiple teachers, each using a different pre-trained text embedding as input. The candidate text embeddings we consider are: mt_grovle , openai-gpt , gpt2-large , gpt2-xl , w2v . So, we obtain a set of up to five models that form the teachers TkT_{k}, k=1..5k=1..5 used by TeachText.

6 Training and implementation details

In order to train our final student, we combine the retrieval loss and the proposed distillation loss L=Lr+Ld\mathcal{L}=\mathcal{L}_{r}+\mathcal{L}_{d}. Our model is trained in Pytorch using the Adam optimizer. TeachText does not add any additional trainable parameters or modalities to the final model. Moreover, when training the student using TeachText, only the additional loss term Ld\mathcal{L}_{d} is added, all other hyper-parameters remaining the same.

Experimental setup

To provide an extensive comparison we test our approach on seven video datasets that have been explored in recent works as benchmarks for the task of text-video retrieval: LSMDC , DiDeMo , MSVD , MSRVTT , ActivityNet , VaTeX and QuerYD . We follow the same experimental setup as prior works .

2 Metrics

To assess performance, we follow prior work (e.g ) and report standard retrieval metrics, including R@K (recall at rank K, where higher is better) and MdR (median rank where lower is better). For certain analyses, to maintain conciseness we report the geometric mean of R@1, R@5 and R@10 rather than individual metrics (this statistic aims to be representative of overall retrieval performance). The numbers are reported for the task of retrieving a video given text queries t2v which is more common in real world applications. The numbers for the reverse task v2t and the number of parameters for each model are reported in the Suppl. Mat. For each experiment, we report the mean and standard deviation of three randomly seeded runs.

3 Ablations

In this section we present an extensive study of our proposed approach. Following the setup used in prior works we conduct ablations on the MSR-VTT dataset , except where otherwise stated.

Baseline improvements. We propose CE+ as an additional baseline which consists of a series of technical improvements to the model of . As seen in Fig. 5a each modification described in Sec. 4.4 brings additional gain over the base architecture. We observe in particular that finetuning the text embedding on the target dataset has a high influence, further highlighting the critical role played by text embeddings and justifying their study. In addition to other changes we found that certain video embedding expert features were highly sensitive to compression choices used in video pre-processing, which we correct accordingly (more details in Suppl. Mat.). Please note that for a fair comparison, in Sec. 5.4 we report the numbers of re-training the methods using these embeddings extracted with the updated pre-processing which yields a higher performance than the ones reported in the original papers. Using multiple text embeddings during inference. TeachText makes no use of additional information at test time. However, it is natural to ask whether the additional text embeddings can be trivially included as part of the model architecture. In Fig. 5(b) we compare our approach with some relatively simple text embedding aggregation techniques, which require access to multiple text embeddings during both training and inference. We observe that TeachText outperforms these aggregation techniques such as direct concatenation or mean of the text embeddings, suggesting that the proposed method is effective in capturing the additional information given by multiple text embeddings. Moreover, the text encoder of existing systems typically employs many parameters, so adding multiple text embeddings to the architecture adds a significant number of parameters (100M+). For example, the concatenation of two text embeddings (provided that they have the same size) almost doubles the total number of parameters for CE+. In contrast, when employing TeachText, no parameters are added. Teacher variation.

The teacher models share the same architecture with the student, but use a different text embedding. We next conduct an ablation on the influence of the number of used teachers. We observe in Fig. 6a that performance increases with the addition of more teachers. Since the combined performance of the teachers after adding more than 3 remains about the same, we do not obtain a further improvement. Thus, for our final experiments presented in Sec. 5.4 we use a combination of three teachers, having the following text embeddings: w2v , gpt2-xl and gpt2-xl-F (gpt2-xl finetuned on the captions from the target dataset). A study of how each individual text embedding affects the final performance can be found in the Suppl. Mat. section Teacher study, where we observe that even when using a teacher with lower performance (w2v), the student has a significant boost in performance. Distillation ablation.

We compare the proposed learning of the similarity matrix with other distillation alternatives. As seen in Fig.6b, our proposed approach is effective in capturing the relationships between video and text. We first provide comparisons between TeachText and several possible instantiations of relational distillation . Indeed, given the highly general nature of , TeachText can be interpreted within this framework as a particular relational configuration that employs cross-modal distillation through batches of similarity matrices. Since the original work of considered single-modality applications, we explore two variations of as baselines for the text-video retrieval task. The first one (Relational), preserves the same intra-text and intra-video relationships independently. We use the same cost function as in and enforce it on both video and text embeddings. The second approach (Pdist), uses the cross modal pairwise distances as a relation measure between text and video as opposed to the similarity matrix. While these methods indeed bring a gain, we observe that TeachText is more effective. We also provide a baseline inspired by the work of which highlights the importance of looking only at the top K predictions given by the teacher. To do so, we enforce the same similarities using TeachText only for the top K ranks given by the teacher rather than for the whole mini-batch. We show the performance for K=1 and K=10 (Rank 1 and Rank 10 presented in Fig.6b). Restricting to only top K predictions when distilling the similarity matrix results in a slight drop in performance. Method generality. To demonstrate the generality of TeachText, we test it against three state of the art methods in addition to the proposed CE+ baseline. In Tab. 1 we observe a consistent gain in performance, independent of the base architecture. Moreover, a gain is achieved across all the datasets that we tested, having over 5% absolute gain on DiDeMo and ActivityNet datasets for MoEE, CE and CE+ models. Note that for MMT we report results on the datasets included in the public implementation provided by the authorshttps://github.com/gabeur/mmt. Method application – Denoising. One immediate application of our method is data denoising. Existing real-world text-video datasets for the retrieval task suffer from label noise which can harm training. More concretely, in crowd-sourced datasets such as MSR-VTT there are some captions that are highly ambiguous/generic (e.g ”A tutorial is presented”, ”Clip showing different colours”, ”A man is writing”) and can describe multiple videos from the dataset. We therefore propose to use TeachText teachers to filter out such cases. For this scenario, we simply remove low-ranked predictions given by teachers and re-train the student using only the new samples. Specifically, we remove all sentences for which the correct video is not ranked in top 40 from the training set. This method is best-suited for datasets where multiple captions per video are available, ensuring that we can remove noisy captions without removing the video itself from training. Following this, we apply the denoising on MSR-VTT and MSVD datasets with the CE+ model. As seen in Fig. 7a, this can be an effective way of further improving the results. Please note, denoising is not used in any other ablations. TeachVideo – Extension to video modalities. While the focus of this work is the use of multiple text embeddings, it is natural to consider whether this approach can be extended to the video encoder modalities. Thus, we introduce the TeachVideo algorithm which follows the same setup as the original TeachText, but now the teacher has access to multiple video modalities instead of multiple text modalities. In this study, all students and all teachers use the same text embedding, so we can assess the gains due to TeachVideo. By employing TeachVideo we retain the computational advantage of requiring fewer video modalities during inference. As it can be seen from our experiments presented in Fig. 7b, the method is effective and brings a boost over the original student. We believe this extension may be useful in scenarios in which limited computational resources are available during inference.

Qualitative examples and other ablation studies are presented in Suppl. Mat.

4 Comparison to prior work

As it can be seen in Tab.2,3,4,5,6,7,8,9 our approach is effective and achieves state of the art results on six datasets. All methods are trained for the retrieval task using only the samples from the target datasets. In order to be as fair as possible, we included the results of our TeachText (abbreviated TT) applied also to the best existing method for each dataset. So, the architecture and the used features are identical during inference (e.g. TT-CE has the same architecture and uses the same video and text embeddings as CE). We highlight in bold the best performing method.

Conclusion

In this paper, we present a novel algorithm TeachText for the text-video retrieval task. We use a teacher-student paradigm where a student learns to leverage the additional information given by one or multiple teachers, sharing the architecture, but each using a different pre-trained text embedding at input. In this way, we achieve state of the art results on six benchmarks. Finally, we present an application of our approach for denoising video retrieval datasets. Acknowledgements. This work was supported by EPSRC Programme Grants Seebibyte EP/M013774/1 and VisualAI EP/T028572/1, and a gift from Adobe. M.L. was supported by UEFISCDI, under project EEA-RO-2018-0496. The authors would like to thank Gyungin Shin and Iulia Duta for assistance. S.A. would like to acknowledge the support of Z. Novak and S. Carlson in enabling his contribution.

References

Appendix A Video embeddings (experts) description

In this work, we used the set of pretrained experts considered by the authors of . For completeness, we summarise here the manner in which these experts were extracted.

Two form of action experts are used: Action(KN) and Action(IG). The former is an I3D architecture trained on Kinetics , which produces 1024-dimensional embeddings from frame clips extracted at 25fps and center cropped to 224 pixels. The Action(IG) model is a 34-layer R(2+1)D model that has ben trained on IG-65m : it operates on frames extracted at 30 fps in clips of 8 at 112 ×\times 112 pixel resolution.

Two forms of object experts are used, named Obj(IN) and Obj(IG). They are produced from frame-level embeddings extracted at 25fps. The Obj(IN) model consists of an SENet-154 backbone which has been trained on ImageNet for image classification. Obj(IG) is formed from a ResNext-101 extractor which was trained on Instagram data that was weakly labelled with hashtags . For both models, frames are resized to 224×\times224 pixels.

The face expert uses a ResNet50 that has been trained for task of face classification on the VGGFace2 dataset , producing a 512-dimensional embedding for each detected face following detection.

The audio expert is produced using the VGGish model, trained for audio classification on the YouTube-8m dataset and described by .

The scene expert is a 2208-dim embedding that is extracted frames (at 25 fps) for a center crop of 224×\times224 pixels. The model, which is pretrained on Places365 , uses a DenseNet-161 architecture.

The speech expert is produced using the Google Cloud API (to transcribe the speech content).

The OCR expert is a word2vec encoding of text detected in frames using .

Apart from dropping the OCR and face experts as described in the main paper, one small modification we propose to the expert selection made by is to replace the Action(IG) from an I3D model to an R2P1D model (matching the architecture Action(IG)) which has also been pretrained on IG-65m and then finetuned on the Kinetics dataset .

Appendix B Text embeddings description

We use several text embeddings. In addition to the Sec. 3 from the main paper, further technical details about each of them are given below:

mt_grovle is a \sayvision-sensitive language embedding which is adapted from w2v using WordNet and an original vision-language graph built from Visual Genome . The size of the final pre-trained embedding is 300.

OpenAI-GPT is a pre-trained text embedding which uses transformers and language modeling on a large corpus (the Toronto Book Corpus) (the final model has 110M params). The size of the final pre-trained embedding is 768.

RoBERTa is a BERT-based embedding . The model is trained longer with bigger batch size on more data, having 125M params. The size of the final pre-trained embedding is 768.

ALBERT is a lightweight modification to BERT which overcomes some memory limitations, having 11M params. The size of the final pre-trained embedding is 768.

GPT2-large is a transformer-based model trained on even more data (40 Gb of text) without any supervision, having 774M params. The size of the final pre-trained embedding is 1280.

GPT2-xl is similar to GPT2-large, but has more parameters (1558M params). The size of the final pre-trained embedding is 1600.

W2V is one of the most popular text embeddings used in vision tasks. It uses a neural network model to learn word representations. The size of the final pre-trained embedding is 300.

Appendix C Further Details for Fig. 1

In Fig.1 from the main paper we highlight that the gain for a model that uses multiple text embeddings (last bar) is comparable with the gain of a model that uses multiple video modalities (middle bar), having as comparison a model that uses only one video modality (first bar). The first bar represents the CE model trained with one video embedding, namely Obj(IG) (the performance of the model is 19.8±0.1 in geometric mean of R1-R5-R10). The second bar represents a CE model using 7 video modalities both for inference and training (the performance of the model is 24.4±0.1 in geometric mean of R1-R5-R10). In the third and final bar of the chart we present the performance of using three different text embeddings with TeachText at training, while using only one text embedding at inference time (the performance of the model is 30.4±0.0 in geometric mean of R1-R5-R10). All the numbers are presented after the modification of the pre-processing pipeline (please see Sec. E for further details). All the experts used by CE are described in Sec.A.

Appendix D Optimization setup

CE+ models are trained in Pytorch using the Adam optimizer . We use a learning rate of 0.001 and weight decay of 1E-5. When using a base architecture different to the proposed CE+, we use the same hyper-parameters as in the public codebase for the underlying method (CEhttps://shorturl.at/ksxIS and MMThttps://github.com/gabeur/mmt). For MoEE, we use the re-implementation provided by the authors of the CE method .

Appendix E Modification to pre-processing pipeline

During our preliminary analysis, we found out that some pretrained expert models produce embeddings that are fairly sensitive to jpeg compression artifacts. To address this, we re-extracted features from video frames densely extracted with minimal jpeg compression (corresponding to the use of ffmpeg and the -qscale:v 2 flag). In order to be fair in our comparisons, we apply this corrections everywhere. Due to this factor, we re-train MoEE and CE and report higher numbers.

Appendix F Dataset details

To provide an extensive comparison we test our approach on seven video datasets that have been explored in recent works as benchmarks for the task of text-video retrieval. Next, we give details about all the datasets used. MSRVTT contains 10k videos, each having 20 captions. In order to test the retrieval performance, we report results on the official split which contains 2990 videos for the test split and 497 for validation, following the setup used in . We perform most of our ablations on this split. To enable comparison with as many other methods as possible, we also report results on the 1k-A split as used in . For this split, we report the performance after training 100 epochs. The split contains 1000 video candidates for testing and 9000 for training. We use the same candidates as defined in which are used by all the other works , using each of the 20 captions associated to each video independently during evaluation and averaging performance across them. MSVD contains 80k English descriptions for a total of 1970 videos. We use the standard split of 1200 (training), 100 (validation) and 670 (testing) as used in other works . The videos from MSVD do not have audio streams. DiDeMo contains 10464 videos sourced from a large-scale creative commons collection and features moments of unedited, diverse content (concerts, sports, pets etc.). The dataset comprises 3-5 pairs of descriptions per video. We adopt the paragraph-video retrieval protocols used by and use splits corresponding to 8392 train, 1065 validation and 1004 test videos. LSMDC contains 118081 short video clips extracted from 202 movies. Each clip is described by a caption that is either extracted from the movie script or from transcribed DVS (descriptive video services) for the visually impaired. There are 7408 clips in the validation set and the testing is performed on 1000 videos from movies that are disjoint from the training and val sets as described in the Large Scale Movie Description Challenge (LSMDC)https://shorturl.at/cdrI6. ActivityNet contains 20k videos extracted from YouTube and has around 100K descriptive sentences. We follow the same paragraph-video retrieval setup as used in prior works and report results on the val1 split. So, we use 10009 videos for training and 4917 videos for testing. VaTeX contains 34911 videos with multilingual captions (Chinese and English). There are 10 captions per video for each language. We follow the same protocol as in and split the validation set equally (1500 validation and 1500 testing videos). In this work, we only use the English annotations. QuerYD contains 1815 videos in the training split and 388 and 390 for validation and testing. The videos are sourced from YouTube and cover a diverse range of visual content. The dataset contains 31441 descriptions, from which 13019 are precisely localized in the video content (having start time and end time annotations) and the other 18422 are coarsely localized. For this work, we do not use the localization annotations and report results for the official splits.

Appendix G Ablations

In this section, we present additional ablations.

In Fig. 8a we vary the batch size for the MSR-VTT dataset in order to see how the performance is affected. As can be seen, we obtain the best value using the same batch size as for the method without applying TeachText algorithm (64 in this case).

G.2 Similarity matrix aggregation study

In Fig. 8b we present several similarity matrix aggregation possibilities: min, max and average. We observe that using the mean of the similarity matrices is more effective. Because of this, we use the mean as the final aggregation technique in our TeachText algorithm.

G.3 Denoising

In Fig. 9a we vary the threshold used to filter out sentences from the training set. As can be seen, this denoising method is effective and it can provide a significant gain in performance. In this experiment we have found out that the best threshold for MSRVTT is rank 40. Additionally, we present denoising results in Fig. 9b for the MSVD dataset using the 100 threshold. This method turns out to be effective in reducing noise for retrieval datasets. Denoising is not use in any other ablation studies. The final results when comparing with other state of the art methods are presented using denoising on MSRVTT and MSVD datasets.

G.4 Distillation setup

As stated in the main paper, the distillation setup admits a number of variants. In addition to the methods presented in Fig. 6b from the main paper, in Fig. 10a we present several additional comparisons. More exactly, we test our approach against a more classical distillation setup where we directly regress the embeddings given by the teacher (Embd regress). This setup does not follow the idea of relational distillation. Additionally, we also apply the angle distillation as introduced by (Relational angle) where we use exactly the loss as in the public codehttps://github.com/lenscloth/RKD. Please note that the drop in performance as opposed to the student without distillation can be explained by some technical challenges that we encountered in order to make the angle loss compatible with the ranking loss used for this task. Last but not least, we also show that a small improvement can be obtained by using a self learning technique, where the teacher has the exact same architecture and inputs as the student.

G.5 Loss study

In the main paper, we follow recent literature and use the Huber loss for distillation. However, we wanted to see how various losses affect the performance. We test with L1 and L2 losses. As can be seen in Fig. 10b, the Huber loss performs better than L1 loss and a bit better than L2.

G.6 Mixture of architectures

Our TeachText assumes that the only difference between the teacher and the student is the used pre-trained text embedding fed to the model. However, our method is not limited to this constraint. In this section, we show how having multiple teachers, that now have a different underlying architecture affect the performance of our method. Please note that in all other ablations, the architecture is shared between student and teacher. This is the only exception. Our preliminary results shown in Fig. 11 suggest that there isn’t much improvement that may be achieved by using a mixture of architectures as teachers. This is somehow expected, since these methods usually share the same video modalities so there isn’t much additional information that may be captured by the combination of multiple architectures. However, we expect to get a further boost if we diversify the set of used modalities.

G.7 Architecture extension

In addition to the main paper, we also introduce a new CE-L base architecture. This is similar to the CE and CE+, but uses w2v as the text embedding. In this way, the number of parameters are greatly reduced, making this the most lightweight architecture in terms of number of parameters that we can create. In Tab. 10, you can see that our method TeachText is effective even when using this lightweight architecture. This architecture also has the lowest numbers of parameters when compared to other state of the art methods as can be seen in Tab.11,12,13,14,15,16,17,18.

G.8 Model complexity

Changes in the pretrained text embedding strongly affect the number of parameters. Because of this factor, using more text embeddings at test time may strongly affect the total number of learnable parameters available to the model (in addition to adding the requirement to extract additional text embeddings during inference). While the simple ’Mean’ aggregation from Fig.5b in the main paper, does not change the number of parameters, the ’Concat’ aggregation adds a significant quantity (approx 240M learnable parameters, yielding total model sizes of 503.98M vs 262.73M for CE+). The proposed TeachText approach leaves the number of parameters untouched.

Since changing the text embedding to CE+ results in an increase in number of learnable parameters, we also study a CE-L architecture as a lightweight alternative in this Suppl. Mat. (described in Sec. 10), which demonstrates that the gain from the proposed TeachText approach is not limited to models with many learnable parameters. Please check Tab.11,12,13,14,15,16,17,18 for the exact number of params for every used architecture.

G.9 Amount of training data vs performance.

We next study how training data quantity influences the proposed method. In Fig. 12a we observe that by using the TeachText with more and more data, the performance gap increases, suggesting that its benefit may prove to be useful even in larger scale dataset scenarios.

G.10 Teacher study

In Fig. 12b, we study how each embedding affects the final performance. We observe that even though the model ingesting w2v embeddings has a significant lower performance than the student model without using TeachText, there is a significant gain when learning from the teacher which uses w2v. This again indicates that there is additional information captured by using a different text embedding which can be exploited by TeachText.

G.11 Influence of distillation over the correctly retrieved samples

In Fig. 13 we present the shares of correctly retrieved samples in terms of R1 on the test set of the MSR-VTT dataset for the student with and without TeachText and for the teacher. In Fig. 13a we present results when we learn from the three teachers and in Fig. 13b we considered the case when we learn only from one teacher (namely w2v). There is a significant share of correctly retrieved sample between the student using TeachText and the teacher.

Appendix H Comparison to prior work

In Tab.11,12,13,14,15,16,17,18 we make an extensive comparison of our method with other methods from the literature. In addition to the numbers reported in the main paper, we also report results for the v2t task. Moreover, we present the number of parameters of each method where available. As can be seen, our TeachText algorithm brings a clear improvement and the total number of parameters remains the same as for the base architecture. In addition to the main paper, we also introduce a new CE-L base architecture. This is similar to the CE and CE+, but uses w2v as text embedding. In this way, the number of parameters is greatly reduced. As can be seen from the tables, this lightweight architecture combined with our TeachText algorithm has very good results showcasing the effectiveness of TeachText across different parameter regimes. Moreover, some qualitative results can be seen in Fig.14.