DiffusionRet: Generative Text-Video Retrieval with Diffusion Model

Peng Jin, Hao Li, Zesen Cheng, Kehan Li, Xiangyang Ji, Chang Liu, Li Yuan, Jie Chen

Introduction

In recent years, text-video retrieval has made significant progress, allowing humans to associate textual concepts with video entities and vice versa . Existing methods for video-text retrieval typically model the cross-modal interaction as discriminant models . Under the discriminant paradigm based on contrastive learning , the primary focus of mainstream methods is to improve the dense feature extractor to learn better representation. This has led to the emergence of a large number of discriminative solutions , and recent advances in large-scale vision-language pre-training models have pushed their state-of-the-art performance even further.

However, from a probabilistic perspective, discriminant models only learn the conditional probability distribution, i.e., p(candidates∣query)p(\textit{candidates}|\textit{query}). This leads to a limitation of discriminant models that they fail to model the underlying data distribution , and their latent space contains fewer intrinsic data characteristics p(query)p(\textit{query}), making it difficult to achieve good generalization on unseen data . In contrast to discriminant models, generative models capture the joint probability distribution of the query and candidates, i.e., p(candidates,query)p(\textit{candidates},\textit{query}), which allows them to project data into the correct latent space based on the semantic information of the data. As a typical consequence, generative models are often more generalizable and transferrable than discriminant models. Recently, generative models have made significant progress in various fields, such as generating high-quality synthetic images , natural language , speech , and music . One of the most popular generative paradigms is the diffusion model , which is a type of likelihood-based model that gradually remove noise via a learned denoising model to generate the signal. In particular, the coarse-to-fine nature of the diffusion model enables it to progressively uncover the correlation between text and video, making it a promising solution for cross-modal retrieval. Therefore, we argue that it is time to rethink the current discriminant retrieval regime from a generative perspective, utilizing the diffusion model.

To this end, we propose a novel diffusion-based text-video retrieval framework, called DiffusionRet, which addresses the limitations of current discriminative solutions from a generative perspective. As shown in Fig. 1, we model the retrieval task as a process of gradually generating joint distribution from noise. Given a query and a gallery of candidates, we adopt the diffusion model to generate the joint probability distribution p(candidates,query)p(\textit{candidates},\textit{query}). To improve the performance of the generative model, we optimize the proposed method from both generation and discrimination perspectives. During training, the generator is optimized by common generation loss, i.e., Kullback-Leibler divergence . Simultaneously, the feature extractor is trained with the contrastive loss, i.e., InfoNCE loss , to enable discriminative representation learning. In this way, DiffusionRet has both the high performance of discriminant methods and the generalization ability of generative methods.

The proposed DiffusionRet has two compelling advantages: First, the generative paradigm of our method makes it inherently generalizable and transferrable, enabling DiffusionRet to adapt to out-of-domain samples without requiring additional design. Second, the iterative refinement property of the diffusion model allows DiffusionRet to progressively enhance the retrieval results from coarse to fine. Experimental results on five benchmark datasets for text-video retrieval, including MSRVTT , LSMDC , MSVD , ActivityNet Captions , and DiDeMo , demonstrate the advantages of DiffusionRet. To further evaluate the generalization of our method to unseen data, we propose a new out-domain retrieval task. In the out-domain retrieval task , labeled visual data with paired text descriptions are available in one domain (the “source”), but no data are available in the domain of interest (the “target”). Our method not only represents a novel effort to promote generative methods for in-domain retrieval, but also provides evidence of the merits of generative approaches in challenging out-domain retrieval settings. The main contributions are as follows:

To the best of our knowledge, we are the first to tackle the text-video retrieval from a generative viewpoint. Moreover, we are the first to adapt the diffusion model for cross-modal retrieval.

Our method achieves new state-of-the-art performance on text-video retrieval benchmarks of MSRVTT, LSMDC, MSVD, ActivityNet Captions and DiDeMo.

More impressively, our method performs well on out-domain retrieval without any modification, which may have a significant impact on the community.

Related Work

Text-Video Retrieval. Text-video retrieval is one of the most popular cross-modal tasks . Most existing text-video retrieval works employ a mapping technique that places both text and video inputs in the same latent space for direct similarity calculation. For instance, CLIP4Clip transfers knowledge from the text-image pre-training model, i.e., CLIP , to enhance video representation. EMCL-Net improves representation capabilities by bridging the gap between text and video. HBI models video-text as game players with multivariate cooperative game theory to handle the uncertainty during fine-grained semantic interaction. These methods are discriminant models and focus primarily on maximizing the conditional likelihood p(candidates∣query)p(\textit{candidates}|\textit{query}), without considering the underlying data distribution p(query)p(\textit{query}). As a result, identifying out-of-distribution data becomes a challenging task. In contrast, our DiffusionRet adopts a generative viewpoint to tackle the task and provides evidence of the merits of generative approaches in a challenging, out-domain retrieval. To the best of our knowledge, we are the first to tackle text-video retrieval from a generative viewpoint.

Diffusion models. Diffusion models are a type of neural generative model that uses the stochastic diffusion process, which is based on thermodynamics. The process involves gradually adding noise to a sample from the data distribution, and then training a neural network to reverse this process by gradually removing the noise. Recent developments in diffusion models have focused on generative tasks, e.g., image generation , natural language generation , and audio generation . Some other works have attempted to adapt the diffusion model for discriminant tasks, e.g., image segmentation , visual grounding , and detection . However, there are no previous works that adapt generative diffusion models for cross-modal retrieval. This work addresses this gap by modeling the correlation between text and video as their joint probability and utilizing the diffusion model to gradually generate the joint probability distribution from noise. To the best of our knowledge, we are the first to adapt the diffusion model for cross-modal retrieval.

Method

In this paper, we tackle the tasks of text-to-video and video-to-text retrieval. In the task of text-to-video retrieval, we are given a text query tt and a gallery of videos V\bm{V}. The goal is to rank all videos v∈Vv\in\bm{V} so that the video corresponding to the text query tt is ranked as high as possible. Similarly, the goal of video-to-text retrieval is to rank all text candidates t∈Tt\in\bm{T} based on the video query vv.

A common method for the retrieval problem is similarity learning . Specifically, the caption as well as the video are represented in a common multi-dimensional embedding space, where the similarity can be calculated as the dot product of their corresponding representations. In the task of text-to-video retrieval, such methods use the logits to calculate the posterior probability:

where τ\tau is the temperature hyper-parameter. θt\bm{\theta_{t}} and θv\bm{\theta_{v}} are parameters of the text feature extractor and video feature extractor, respectively. Finally, existing methods rank all video candidates based on the posterior probability p(v∣t;θt,θv)p(v|t;\bm{\theta_{t}},\bm{\theta_{v}}). Similarly, in the video-to-text retrieval, they rank all text candidates based on p(t∣v;θt,θv)p(t|v;\bm{\theta_{t}},\bm{\theta_{v}}). The parameters {θt,θv}\{\bm{\theta_{t}},\bm{\theta_{v}}\} of text feature extractor and video feature extractor are optimized by minimizing the contrastive learning loss :

where D\mathcal{D} is a corpus of text-video pairs (t,v)(t,v). Such learning strategy is equivalent to maximizing conditional likelihood, i.e., ∏(t,v)∈Dp(v∣t)+∏(t,v)∈Dp(t∣v)\prod_{(t,v)\in\mathcal{D}}p(v|t)+\prod_{(t,v)\in\mathcal{D}}p(t|v), which is called discriminative training .

Since existing methods directly model the conditional probability distribution p(v∣t)+p(t∣v)p(v|t)+p(t|v), without considering the input distribution p(t)p(t) and p(v)p(v), they fail to achieve good generalization on unseen data.

2 DiffusionRet: Generation Modeling

Our method reformulates the retrieval task from a generative modeling perspective. As Feynman’s mantra “what I cannot create, I do not understand” shows, we argue that it is time to rethink the current discriminative retrieval regime.

Inspired by the great success of diffusion models , we adopt the diffusion model as the generator. Concretely, given a query and NN candidates, our goal is to synthesize the distribution x1:N={xi}i=1Nx^{1:N}=\{x^{i}\}_{i=1}^{N} from Gaussian noise N(0,I)\mathcal{N}(0,\bm{\textrm{I}}). In contrast to the prior works, which typically optimize the posterior probabilities p(v∣t;θt,θv)+p(t∣v;θt,θv)p(v|t;\bm{\theta_{t}},\bm{\theta_{v}})+p(t|v;\bm{\theta_{t}},\bm{\theta_{v}}), our method builds the joint probabilities:

where ϕ\bm{\phi} is the parameter of the generator. It is worth noting that the learning objective of the generator is equivalent to approximating the data distribution, i.e., ∏(t,v)∈Dp(v,t)\prod_{(t,v)\in\mathcal{D}}p(v,t), which is called generative training . The overview of our DiffusionRet is shown in Fig. 2.

To extract the joint encoding of text and video, we propose the text-frame attention encoder, which takes text representation as query and frame representation as key and value.

where τ′{\tau}^{{}^{\prime}} is the trade-off hyper-parameter. The smaller τ′{\tau}^{{}^{\prime}} allows visual features to take more textual information into account when aggregated.

2.2 Query-Candidate Attention Denoising Network

Different from other generation tasks which only focus on the authenticity and diversity of generation, the key to the retrieval task is to mine the correspondence between the query and candidates. To this end, we propose the query-candidate attention denoising network to capture the correspondence between query and candidates in the generation process. The overview of the denoising network is shown in Fig. 3.

where WQtW_{Q_{t}}, WKvW_{K_{v}} and WVvW_{V_{v}} are projection matrices. “Proj” projects the noise level kk into DD dimensional embedding. To give more weight to the video candidates with higher joint probabilities of the previous noise level, we add the distribution xkx_{k} to the attention weight. The attention mechanism can be defined as:

Similarly, in the video-to-text retrieval, we feed the projection of video representation as query and the projection of text representation as key and value into an attention module. The output distribution is calculated in the same way.

2.3 Optimization from both Generation Perspective and Discrimination Perspective

Compared to the generation methods, discriminant methods usually give good predictive performance . To leverage the benefit of both generative and discriminative methods, we optimize the proposed generation model from both generation and discrimination perspectives.

Probabilistic Diffusion (Generation Perspective). In the generation perspective, we model the distribution x=p(v,t∣ϕ)x=p(v,t|\bm{\phi}) as the reversed diffusion process of the KK-length Markov Chain. Specifically, our method learns the joint distribution of text and video by gradually denoising a variable sampled from Gaussian distribution. In a forward diffusion process q(xk∣xk−1)q(x_{k}|x_{k-1}), noised sampled from Gaussian distribution is added to a ground truth data distribution x0x_{0} at every noise level kk:

where βk{\beta}_{k} decides the noise schedule which gradually increases. We can sample xkx_{k} by the following formula:

where αˉk=∏i=1kαi\bar{\alpha}_{k}=\prod_{i=1}^{k}\alpha_{i} and αk=1−βk{\alpha}_{k}=1-{\beta}_{k}. ϵ\epsilon is a noise sampled from N(0,1)\mathcal{N}(0,1). Instead of predicting the noise component ϵ\epsilon, we follow and predict the data distribution itself, i.e., x^0=fϕ(v,t,xk)\hat{x}_{0}=f_{\bm{\phi}}(v,t,x_{k}). The training objective of the diffusion model can be defined as:

This loss maximizes the likelihood of p(v,t)p(v,t) by bringing fϕ(v,t,xk)f_{\bm{\phi}}(v,t,x_{k}) and x0x_{0} closer together.

Contrastive Learning (Discrimination Perspective). In the discrimination perspective, we optimize the features which are input into the generator so that these features contain discriminant semantic information. Inspired by , we align representations of text and video at the token level. Specifically, we take all tokens output by the text encoder as word-level features {wi}i=1Nt\{w^{i}\}^{N_{t}}_{i=1}, where NtN_{t} is the length of the text. Frame-level features {fj}j=1Nv\{f^{j}\}^{N_{v}}_{j=1} are all tokens output by the video encoder, where NvN_{v} is the length of the video. Then, we calculate the alignment matrix A=[aij]Nt×NvA=[a_{ij}]^{N_{t}\times N_{v}}, where aij=(wi)⊤fj∥wi∥∥fj∥a_{ij}=\frac{(w^{i})^{\top}f^{j}}{\|w^{i}\|\|f^{j}\|} is the cosine similarity between the ithi_{th} word and the jthj_{th} frame. The total similarity score consists of two parts: text-to-video similarity and video-to-text similarity. For the text-to-video similarity, we first calculate the maximum alignment score of the ithi_{th} word as maxj aij\underset{j}{\textrm{max}}\ a_{ij}. We then take the weighted average maximum alignment score over all words. For the video-to-text similarity, we take the weighted average maximum alignment score over all frames. The total similarity score can be defined as:

where \{g_{t}^{i}\}_{i=1}^{N_{t}}\!=\!\textrm{Softmax}\big{(}\textrm{MLP}_{t}(\{w^{i}\}_{i=1}^{N_{t}})\big{)} and \{g_{v}^{j}\}_{j=1}^{N_{v}}\!\!=\!\!\textrm{Softmax}\big{(}\textrm{MLP}_{v}(\{f^{j}\}_{j=1}^{N_{v}})\big{)} are the weights of the text words and video frames, respectively. Then, the contrastive loss can be formulated as:

where st,vs_{t,v} is the similarity score between text tt and video vv. τ^\hat{\tau} is the temperature hyper-parameter. This loss brings semantically similar texts and videos closer together in the representation space, thus helping the diffusion model to generate the joint distribution of text and video.

Experiments

Datasets. MSRVTT contains 10,000 YouTube videos, each with 20 text descriptions. We follow the 1k-A split with 9,000 videos for training and 1,000 for testing. LSMDC contains 118,081 video clips from 202 movies. We follow the split of with 1,000 videos for testing. MSVD contains 1,970 videos. We follow the official split of 1,200 and 670 as the train and test set, respectively. ActivityNet Captions contains 20,000 YouTube videos. We report results on the “val1” split of 10,009 and 4,917 as the train and test set. DiDeMo contains 10,464 videos annotated 40,543 text descriptions. We follow the training and evaluation protocol in .

Metrics. We choose Recall at rank K (R@K), the Sum of Recall at rank {1,5,10}\{1,5,10\} (Rsum), Median Rank (MdR), and mean rank (MnR) to evaluate the retrieval performance.

Implementation Details. Following previous works , we utilize the CLIP (ViT-B/32) as the pre-trained model. The dimension of the feature is 512. The temporal transformer is composed of 4-layer blocks, each including 8 heads and 512 hidden channels. The temporal position embedding and parameters are initialized from the text encoder of the CLIP. We use the Adam optimizer and set the batch size to 128. The initial learning rate is 1e-7 for the text encoder and video encoder and 1e-3 for other modules. We set the temperature τ^\hat{\tau} to 0.01 and τ′{\tau}^{{}^{\prime}} to 1. For short video datasets, i.e., MSRVTT, LSMDC, and MSVD, the word length is 32 and the frame length is 12. For long video datasets, i.e., ActivityNet Captions and DiDeMo, the word length is 64 and the frame length is 64. The training is divided into two stages. In the first stage, we train the feature extractor from the discrimination perspective. In the second stage, we optimize the generator from the generation perspective. For the MSRVTT and LSMDC datasets, the experiments are carried out on 2 NVIDIA Tesla V100 GPUs. For the MSVD, ActivityNet Captions, and DiDeMo datasets, the experiments are carried out on 8 NVIDIA Tesla V100 GPUs. In both of the tasks of text-to-video and video-to-text retrieval, we assume that only the candidate sets are known in advance. In the inference phase, we consider both the distance of video and text representations in the representation space and the joint probability of video and text.

2 Comparison with State-of-the-art

We compare the proposed DiffusionRet with other methods on five benchmarks. In Tab. 1, we show the retrieval results on the MSRVTT dataset. Our model outperforms the recently proposed state-of-the-art methods on both text-to-video retrieval and video-to-text retrieval tasks. Tab. 2 shows text-to-video retrieval results on the LSMDC, MSVD, ActivityNet Captions, and DiDeMo datasets. DiffusionRet achieves consistent improvements across different datasets, which demonstrates the effectiveness of our method.

3 Ablation Study

Generation loss type. In Tab. 3a, we compare two common generation losses, i.e., mean-squared loss (MSE) and Kullback-Leibler (KL) divergence. Results demonstrate that the KL divergence achieves optimal retrieval performance. We explain that it is because KL divergence can better measure the distance between probabilities than MSE, so it is more suitable for probability generation.

Sampling strategy. Denoising diffusion probabilistic models (DDPM) learn the underlying data distribution by a Markov chain as shown in Eq. 7. To accelerate the sampling process of diffusion models, denoising diffusion implicit models (DDIM) formulate a Markov chain that reverses a non-Markovian perturbation process. As shown in Tab. 3b, we find that the two sampling strategies have similar performance. To speed up the sampling process, we use DDIM by default in practice.

Schedule of β\beta. The schedule of β\beta controls how the step size increases. We compare the linear schedule and cosine schedule in Tab. 3c. As shown in Tab. 3c, we find that the cosine schedule performs well, so we adopt the cosine schedule by default, which is the same as the default setting in the motion generation task .

Training strategy. In Tab. 3d, we compare different training strategies. Similar to the findings of the pioneer work , we find that pure discriminant training has better performance on limited data than pure generative training. The hybrid training method we proposed achieves the best results, which indicates that our hybrid training strategy can combine the advantages of the two training methods.

The number of steps. In Tab. 3e, we explore the influence of the number of diffusion steps. Results demonstrate that the number of steps of 50 achieves optimal performance in the retrieval task, outperforming the standard value of 1000 of the image generation task . We consider that this is due to the simpler probability distribution in retrieval compared to the pixel distribution in natural images. Therefore, compared with the image generation task, the cross-modal retrieval task only needs fewer diffusion steps.

Scale of β\beta. The scale of β\beta indicates the signal-to-noise ratio of the diffusion process . We evaluate the scale range [0.1,2.0][0.1,2.0] as shown in Tab. 3f. We find that the model achieves the best performance at the scaling of 1.0, so we set the scale of β\beta to 1.0 in practice, which is the same as the default setting in the image generation task .

4 The Efficiency of DiffusionRet

We provide comparisons of the inference time and memory consumption of our method with other methods under different conditions (the number of diffusion steps and the size of the candidate gallery) in Tab. 4. As shown in Tab. 4, our method is as efficient as the existing state-of-the-art methods during the inference stage, which we explain in the following three aspects. (1) Lightweight denoising network. Our denoising network is lightweight (2.50 M parameter), while other methods use complex matching networks, e.g., TS2-Net which uses the token shift transformer and token selection transformer for fine-grained alignment. (2) Efficient feature extractor. About 80% of the inference time is spent on feature extraction such as extracting query text features. Compared with other methods with bulky feature extractors, e.g., X-Pool which uses an additional transformer-based pooling module to aggregate features, our method only uses vanilla transformer to extract features and is therefore more efficient. (3) Scalability. We can increase the number of diffusion steps to boost performance at a negligible time cost, indicating the scalability of our method. Compared with other methods, our method is more flexible and better suited to different retrieval scenarios where accuracy and speed are required differently.

5 Out-domain Retrieval

Current text-video retrieval methods are mainly evaluated on the same dataset. To further evaluate the generalization to unseen data, we conduct out-domain retrieval : we first pre-train a model on one dataset (the “source”) and then measure its performance on another dataset (the “target”) that is unseen in the training. As shown in Tab. 5, we compare the proposed DiffusionRet with other baselines (i.e., CLIP4Clip and EMCL-Net ) in out-domain retrieval settings. We find that the discriminant approaches fail to migrate the performance of in-domain retrieval to out-of-domain retrieval well. For example, the performance of EMCL-Net is significantly higher than that of CLIP4Clip in in-domain retrieval. However, in out-domain retrieval, the performance of EMCL-Net is slightly lower than that of CLIP4Clip. In contrast, DiffusionRet performs well for both in-domain and out-of-domain retrieval.

To further illustrate the benefits of our generative modeling, we provide the visualization of the similarity distribution in in-domain retrieval and out-domain retrieval. As shown in Fig. 4, compared to the baseline models, our DiffusionRet maximally keeps the distributions of positive and negative pairs separate from each other in both in-domain and out-domain settings. These results confirm that our method is more generalizable and transferrable for the unseen data than typical discriminant models.

6 Qualitative Analysis

The coarse-to-fine nature of the diffusion model enables it to progressively uncover the correlation between text and video, rendering it an effective approach for cross-modal retrieval. To better understand the diffusion process, we show the additional visualization of the diffusion process in Fig. 5. These results demonstrate that our method can progressively uncover the correlation between text and video.

7 Why Diffusion Models

Diffusion models have demonstrated remarkable generative power in various fields. Besides the powerful generative power of diffusion models, we explain other advantages of applying the diffusion model rather than other generative approaches to cross-modal retrieval, mainly in two aspects. First, the coarse-to-fine nature of the diffusion model enables it to progressively uncover the correlation between text and video, rendering it a more effective approach for retrieval tasks than other generation training methods, such as generative adversarial network and variational autoencoder . Second, the many-to-many nature of the diffusion model makes it more suitable for generating joint probabilities than the auto-regressive networks . In Fig. 5, we show the visualization of the diffusion process. These results further demonstrate the above two advantages of applying the diffusion model rather than other generative approaches to cross-modal retrieval.

Conclusion

In this paper, we propose DiffusionRet, the first generative diffusion-based framework for text-video retrieval. By explicitly modeling the joint probability distribution of text and video, DiffusionRet shows promise to solve the intrinsic limitations of the current discriminant regime. It successfully optimizes the DiffusionRet from both the generation perspective and discrimination perspective. This makes DiffusionRet principled and well-applicable in both in-domain retrieval and out-domain retrieval settings. To the best of our knowledge, we are the first to tackle text-video retrieval from the generative viewpoint. Besides, we show that our generation modeling method is superior to existing discriminant methods in terms of performance and generalization ability. This work is the first effort to promote generative methods for text-video retrieval, which may be meaningful and helpful to the community.

Acknowledgements. This work was supported in part by the National Key R&D Program of China (No. 2022ZD0118201), Natural Science Foundation of China (No. 61972217, 32071459, 62176249, 62006133, 62271465, 62202014). Li Yuan is also sponsored by CCF Tencent Open Research Fund.

References

Appendix A Appendix

This appendix provides the descriptions of datasets (Sec. A.1.1), implementation details (Sec. A.1.2), video-to-text retrieval performance on the LSMDC, MSVD, and ActivityNet Captions datasets (Sec. A.2.1), additional experiments for the out-domain retrieval (Sec. A.2.2), the reason for applying diffusion models (Sec. A.2.3), discussion of the limitations (Sec. A.2.4), the additional visualization of the diffusion process (Sec. A.3.1), the visualization of the text-frame attention map (Sec. A.3.2), and the visualization of the text-to-video retrieval examples (Sec. A.3.3).

We compare the proposed DiffusionRet with other methods on five benchmark text-video retrieval datasets, including MSRVTT , LSMDC , MSVD , ActivityNet Captions , and DiDeMo .

MSRVTT. MSRVTT contains 10,000 YouTube videos, each with 20 text descriptions. We follow the training protocol in and evaluate on text-to-video and video-to-text search tasks on the 1K-A testing split with 1K video or text candidates defined by .

LSMDC. LSMDC contains 118,081 video clips from 202 movies. The duration of videos in the LSMDC dataset is short. We follow the split of with 1,000 videos for testing.

MSVD. MSVD contains 1,970 videos. Each video has approximately 40 associated text description. Videos in the MSVD dataset are short in duration, lasting about 10 to 25 seconds. We follow the official split of 1,200 and 670 as the train and test set, respectively.

ActivityNet Captions. ActivityNet Captions consists densely annotated temporal segments of 20K YouTube videos. Following , we concatenate descriptions of segments in a video to construct “video-paragraph” for retrieval. We report results on the “val1” split of 10,009 and 4,917 as the train and test set.

DiDeMo. DiDeMo contains 10,464 videos annotated 40,543 text descriptions. We concatenate descriptions of segments in a video to construct “video-paragraph” for retrieval. We follow the training and evaluation protocol in .

A.1.2 Implementation Details.

Following previous works , we utilize the CLIP (ViT-B/32) as the pre-trained model. The dimension of the feature is 512. The temporal transformer is composed of 4-layer blocks, each including 8 heads and 512 hidden channels. The temporal position embedding and parameters are initialized from the text encoder of the CLIP. We use the Adam optimizer and set the batch size to 128. The initial learning rate is 1e-7 for the text encoder and video encoder and 1e-3 for other modules. We set the temperature τ^\hat{\tau} to 0.01 and τ′{\tau}^{{}^{\prime}} to 1. For short video datasets, i.e., MSRVTT, LSMDC, and MSVD, the word length is 32 and the frame length is 12. For long video datasets, i.e., ActivityNet Captions and DiDeMo, the word length is 64 and the frame length is 64.

The training is divided into two stages. In the first stage, we train the feature extractor from the discrimination perspective. In the second stage, we optimize the generator from the generation perspective. For the MSRVTT and LSMDC datasets, the experiments are carried out on 2 NVIDIA Tesla V100 GPUs. For the MSVD, ActivityNet Captions, and DiDeMo datasets, the experiments are carried out on 8 NVIDIA Tesla V100 GPUs. In both of the tasks of text-to-video and video-to-text retrieval, we assume that only the candidate sets are known in advance. In the inference phase, we consider both the distance of video and text representations in the representation space and the joint probability of video and text. Code is available at https://github.com/jpthu17/DiffusionRet.

A.2 Additional Results and Discussions

We compare the proposed DiffusionRet with other meth- ods on five benchmark. In addition to the text-to-video retrieval results in the main paper, we provide video-to-text retrieval results on the LSMDC, MSVD, ActivityNet Captions, and DiDeMo datasets in Tab. A. Extensive experiments on five datasets, including MSRVTT, LSMDC, MSVD, ActivityNet Captions, and DiDeMo, demonstrate that our method is capable of dealing with both short and long videos. DiffusionRet achieves consistent improvements across different datasets, which demonstrates the effectiveness of our method.

A.2.2 Out-domain Retrieval

Most text-video retrieval methods are evaluated using the same dataset, which may not reflect their ability to generalize to unseen data. To this end, we perform out-domain retrieval by pre-training a model on one dataset (referred to as the “source”) and evaluating its performance on another dataset (referred to as the “target”) that is not included in the training. In addition to the out-domain retrieval experiments in the main paper, we provide additional experiments in the out-domain retrieval setting (MSRVTT->ActivityNet Captions) in Tab. B. We find that discriminant approaches do not transfer well from in-domain to out-of-domain retrieval. For instance, EMCL-Net outperforms CLIP4Clip in in-domain retrieval, but its performance is slightly lower than CLIP4Clip in out-domain retrieval. In contrast, DiffusionRet achieves good performance in both in-domain and out-of-domain retrieval.

A.2.3 Why Diffusion Models

Diffusion models have demonstrated remarkable generative power in various fields. Besides the powerful generative power of diffusion models, we explain other advantages of applying the diffusion model rather than other generative approaches to cross-modal retrieval, mainly in two aspects. First, the coarse-to-fine nature of the diffusion model enables it to progressively uncover the correlation between text and video, rendering it a more effective approach for retrieval tasks than other generation training methods, such as generative adversarial network and variational autoencoder . Second, the many-to-many nature of the diffusion model makes it more suitable for generating joint probabilities than the auto-regressive networks . We recommend further investigation of the potential of the generative method for discriminant tasks in future research. In our future work, we will explore our algorithm in segmentation and visual question answering .

A.2.4 Limitations of our Work

Generative models have focused on generative tasks, e.g., image generation , natural language generation , and audio generation . Some other works have attempted to adapt the generative models for discriminant tasks, e.g., image segmentation , visual grounding , and detection . However, these precursor methods require additional discriminative training. To train on limited data, we optimize the proposed generation model from both generation and discrimination perspectives. Although such a hybrid training method can improve model performance with limited data, we believe that pure generative training is a more promising solution when the data is sufficient. We suggest exploring a pure generative training approach to the retrieval problem in the future.

A.3 Additional Visualizations

The coarse-to-fine nature of the diffusion model enables it to progressively uncover the correlation between text and video, rendering it an effective approach for cross-modal retrieval. To better understand the diffusion process, we show the additional visualization of the diffusion process in Fig. A. These results demonstrate that our method can progressively uncover the correlation between text and video.

A.3.2 Text-Frame Attention Map

To extract the joint encoding of text and video, we propose the text-frame attention encoder, which takes text representation as query and frame representation as key and value. To better understand the process of joint encoding of text and video, we show the visualization of the text-frame attention map in Fig. B. As shown in Fig. B, the text-frame attention encoder adaptively extracts the frames that are similar to the text so that fine-grained video features can be extracted. These results demonstrate that our method can capture the correlation between text and frames.

A.3.3 Text-to-Video Retrieval

We show two retrieval examples from the MSRVTT testing set for text-to-video retrieval in Fig. C. As shown in Fig. C, our method successfully retrieves the ground-truth video. These results demonstrate that our method can mine the correlation between text and video effectively.