Unified Coarse-to-Fine Alignment for Video-Text Retrieval

Ziyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius, Mohit Bansal

Introduction

The fields of computer vision and natural language processing have both seen significant progress in recent years. Thus, the cross-modal alignment , which involve developing techniques to connect these two domains, has seen considerable attention and progress. As a direct application of cross-modal alignment, the video-text retrieval task aligns video(text) candidates with text(video) queries to identify the most relevant videos, and the standard practice is to align the video and text features extracted by vision and language encoders. Recently, the emergence of large-scale image-text pretrained models prompted several methods to utilize CLIP image and text encoder to achieve strong performance on many video-text retrieval benchmarks. As a direct extension of CLIP, Luo et al. proposed temporal fusion modules to aggregate the features of different video frames and then perform the cross-modal alignment on video and text features. Later, to capture more correspondences between video and text, several works propose to conduct the alignment between frame and text features. Ma et al. take a step forward and leverage an alignment between frame and word features for more detailed information.

Although the aforementioned methods have achieved impressive results, they only rely on high-level visual information (frame and video) to perform the cross-modal alignment. This coarse-grained alignment only captures the high-level visual clues (scene, action, etc) that connect to the text query. As shown in the first row of Figure 1, the coarse-grained alignment only captures the scene of the stage with the huge audience” and the action of “singing (possibly talking)”, thus leading to the incorrect retrieval result. On the other hand, Zou et al. build a fine-grained alignment between patch tokens from the video and word tokens from the text query. As illustrated in the second row of Figure 1, the fine-grained alignment does capture the detailed information like “microphone”, but it might overlook high-level clues like scene information (“stage with the huge audience”). These results reveal that video-text retrieval requires an understanding of the both high-level and low-level correspondence between text and video. Thus, in this work, we aim to jointly consider coarse-grained and fine-grained cross-modal alignment and how to unify them to get the correct answer (as shown in the last row of Figure 1).

To this end, we propose UCoFiA, a Unified Coarse-to-fine Alignment model for video-text retrieval. Our approach aims to capture the multi-grained similarity between the text and video by performing alignment at different granularity. We begin with a coarse-grained alignment between the entire video and the query sentence (video-sentence). Next, we perform frame-sentence alignment by matching individual video frames and the query sentence. Finally, we conduct a fine-grained alignment between the video patches and query words (patch-word).

However, while this multi-grained information provides richer, more diverse detailed information, it also brings significant irrelevant information to the cross-modal alignment. For instance, several frames in the video might not contain information related to the query, and some patches in a frame might only correspond to the background information unrelated to any subjects in the query. The irrelevant information could impede the model from learning precise cross-modal correspondence. To address these issues, we first propose an Interactive Similarity Aggregation module (ISA) that considers the importance of different visual features while aggregating the cross-modal similarity to obtain a similarity score for each granularity. For frame-sentence alignment, our ISA module jointly considers the cross-modal similarity and the interaction of frame features while aggregating the frame-sentence similarity. Compared to the previous methods that ignore the temporal clues between video frames, our ISA module can better capture the important information within continuous video frames. Note that the ISA module is a general similarity aggregation approach regardless of the feature granularity, and we further extend it to a bidirectional ISA module for patch-word alignment.

Next, once we obtain the similarity score for each level of alignment, we can sum them to one score as the final retrieval similarity. However, we find that similarity scores across different videos are highly imbalanced, and we empirically show that correcting this imbalance before summation improves the performance. Concretely, sometimes the sum of retrieval similarities between one specific video and all texts (we called this term marginal similarity) might be much higher than that of the other videos, meaning that this video is over-represented and will lower the probability of the other video being selected. To address this, we utilize the Sinkhorn-Knopp algorithm to normalize the similarity scores and make sure the marginal similarities for different videos are almost identical so that each video has a fair chance to be selected after normalization. We then unify the scores of different levels by performing the algorithm separately on the similarities of different levels and summing them together.

We validate the effectiveness of our UCoFiA model on diverse video-text retrieval benchmarks. Specifically, UCoFiA achieve a text-to-video retrieval R@1 of 49.4%49.4\% on MSR-VTT and 45.7%45.7\% on ActivityNet, thus, outperforming the current state-of-the-art CLIP-based methods by 2.4%2.4\% and 1.4%1.4\%, respectively.

Related Work

Video-text Retrieval. Video-text retrieval is a fundamental topic in the vision-language domain and has attracted significant research attention. To retrieve the correct video candidate given the text query, it is crucial to align the features of the related video and text sample together. To this end, early works in video-text retrieval focus on designing fusion mechanisms for the alignment between pre-extracted and frozen video and text features. Later, ClipBERT proposes a sparse sampling strategy on video data to accomplish end-to-end training and apply image-text pretraining for video-text retrieval. Afterward, Bain et al. utilize a curriculum learning schedule to accomplish joint image and video end-to-end training on cross-modal data. With the great success of large-scale image-text pretraining model CLIP , several works utilize the powerful CLIP encoder for video-text retrieval tasks and achieve state-of-the-art results with an efficient training paradigm. Thus, in this work, we also use CLIP as our image-text backbone to enable a fair comparison with existing methods.

Moreover, most cross-modal alignment approaches can be divided into two categories: coarse-grained alignment which leverage the frame-level (or video-level) visual features and fine-grained alignment which utilize the information of patch within each video frames. Recently, several coarse-grained CLIP-based methods use frame aggregation strategies to convert frame features to video features and perform coarse-grained alignment between the video and query features. Follow-up works investigate various similarity calculation schemes for better cross-modal alignment. TS2-Net designs a cross-modal alignment model between the frame feature and sentence feature of the text query. X-CLIP utilizes a cross-grained alignment between coarse-grained video features and text features, including video-sentence, video-word, frame-sentence, and frame-word contrast. However, the coarse-grained alignment fails to capture detailed correspondence to due the limited information within high-level features. To this end, TokenFlow proposes a fine-grained alignment function for token-wise similarity calculation. The fine-grained alignment does capture more subtle correspondence between text and video, but it could overlook the high-level information like scene and action. In this work, we aim to combine the advantage of both coarse-grained and fine-grained alignment to better capture the correspondence between text query and video candidates.

Normalization for Video-text Retrieval. To retrieve the most relevant video candidate given a text query, common video-text retrieval models compute a similarity matrix between video and text input and retrieve the candidate with the highest similarity. Several previous works focus on the normalization of this similarity matrix to improve the performance. CAMoE introduces a Dual Softmax Loss (DSL) as a reviser to correct the similarity matrix and achieve the dual optimal match. Later, NCL reveals that cross-modal contrastive learning suffers from incorrect normalization of the sum retrieval probabilities of each text or video instance and proposes Normalized Contrastive Learning that computes the instance-wise biases that properly normalize the sum retrieval probabilities. Empirically, we find that the logits from the multi-level similarity matrix are imbalanced and make some videos and texts over- or under-representative. To mitigate the issue, we propose to balance the multi-level alignments by separately normalizing each similarity matrix and aggregating the normalized matrix for better retrieval prediction.

Methodology

In this selection, we present our proposed UCoFiA model. As shown in Figure 2, UCoFiA consists of four components: (1) text and video encoders, (2) coarse-to-fine alignment module, (3) Interactive Similarity Aggregation module, (4) and multi-granularity unification module with the Sinkhorn-Knopp algorithm. First, the video and text encoders extract multi-grained visual and textual features. Afterward, multi-grained features are fed into a coarse-to-fine alignment module that calculates the different levels of cross-modal similarity. Then, the Interactive Similarity Aggregation module fuses the similarity vector (or matrix, depending on the input type) and obtains the similarity score for each granularity. Finally, the multi-granularity unification module aggregates the similarity scores from all granularity and obtains the final unified similarity score for retrieval prediction. Below, we discuss each of these components in more detail.

2 Coarse-to-fine Alignment

Our proposed coarse-to-fine alignment module calculates the cross-modal similarity from multi-grained visual and textual inputs to address the weakness of only considering either coarse alignment or fine-grained alignment (as shown in Figure 1). First, we adopt a video-sentence alignment to obtain the similarity score of video and sentence features. Then, we leverage a frame-sentence alignment to capture the similarity between each frame and text query and obtain a frame-sentence vector. Lastly, we apply the most fine-grained patch-word alignment to model the similarity between each patch and word representation and obtain a patch-word matrix. Below, we describe each of these alignments in more detail.

Overall, the coarse-to-fine alignment module allows our model to capture cross-modal similarity from the different granularity of features. In Table 3, we demonstrate the effectiveness of each level alignment quantitatively.

3 Interactive Similarity Aggregation

Next, we describe how we aggregate the cross-modal similarity vector and matrix for different alignments. Due to the high redundancy of video features, irrelevant information within the similarity vector could impede the model from learning precise cross-modal correspondence. Existing methods propose a softmax-based weighted combination of the similarity vector to reduce the impact of irrelevant information. However, the softmax weights fail to capture the information between input features. For instance, while aggregating the frame-sentence alignment, softmax weights ignore the temporal information across frames. As a result, the weighted combination only focuses on the cross-modal relevance between text query and video frames and ignores the interaction between different frames. To this end, we propose a simple, yet effective interactive similarity aggregation module (ISA).

where s\textscfs{s_{\textsc{fs}}} denotes the frame-sentence similarity score between the text query and video candidate. To sum up, the ISA module is capable of aggregating the similarity vector to obtain a similarity score by jointly considering cross-modal relevance and feature interaction. Regardless of the feature dimension, the ISA module is flexible enough to deal with similarity vectors with different feature granularity. Thus, we extend the ISA module to aggregate the patch-word similarity matrix.

A direct idea is to flatten the patch-word matrix to a large vector and apply the ISA module to obtain the similarity score. However, due to the modality gap, it is difficult to model the feature interaction across video and text for the ISA module. The alternative idea is to split each row or column of the similarity matrix into a similarity vector and leverage the ISA module on each vector. Afterward, we aggregate the similarity score from each vector to another similarity vector and apply another ISA to obtain the patch-word score. In that way, we can separately model the feature interaction between patches and words to provide better similarity aggregation. We consider two ways of aggregation: patch-then-word (see the top of Figure 3(b)) and word-then-patch (see the bottom of Figure 3(b)). Empirically we find that jointly considering these two directions provides better aggregation for the patch-word matrix. To this end, we combine the strength of both directions and name the module as a bi-directional ISA module (Bi-ISA, Figure 3(b)). To describe the module formally, we first adopt a patch-level ISA module Ap\mathcal{A}_{p} on C\textscpw{\textbf{C}_{\textsc{pw}}} to obtain a word-level similarity vector. We then adopt a word-level ISA module Aw\mathcal{A}_{w} to aggregate the word-level similarity vector to the patch-then-word score. Similarly, we can obtain the word-then-patch score by leveraging a word-level ISA module and a patch-level ISA in a reverse way. The whole process of Bi-ISA can be formulated as:

where s\textscpw{s_{\textsc{pw}}} denotes the patch-word similarity score. Conceptually, our ISA module jointly considers the cross-modal relevance and the interaction between different features while aggregating the similarity vector (matrix). In Table 5 and Table 6, we validate the effectiveness of our ISA and Bi-ISA module compared to different aggregation mechanisms. In the next section, we further aggregate the different levels of similarity score to one final score for retrieval.

4 Unifying Coarse and Fine-grained Alignments

However, we find that scores across different videos are highly imbalanced in the similarity matrices of each level, and we empirically find that correcting the issue before summing the similarities leads to a better result in multi-level alignment methods . The imbalance issue is similar to the findings of Park et al. . Specifically, sometimes the summation of retrieval similarities between one specific video and all texts (we called this term marginal similarity in the following) might be much higher than that of the other videos, meaning that this video is over-represented and will lower the probability of the other video being selected. To address this, we re-scale the similarity matrix to normalize the marginal similarity of every video to be a similar value. One approach is to apply dual softmax operation on the similarity matrix, but this is not realistic in the testing phase since it requires obtaining all the testing videos and queries at hand.

We then obtain a normalized similarity matrix for testing videos. We apply the algorithm separately on the similarity matrix of different alignments before summing them together, and we empirically find out this is better than doing summation first and then normalization. Finally, the final retrieval score R can be written as:

Note that we only apply Equations 3 and 4 in the inference phase. Similarly, for video-to-text retrieval, we normalize the similarity matrix by adding the test text bias. In Table 7, we validate the effectiveness of applying the above normalization. We also provide a visualization of the reduction of over-/under-representation by applying this technique in the supplementary material.

5 Training and Inference

During inference, to perform video-text retrieval, we will compute the similarities between all videos and the query, normalize the similarities by the method introduced in Section 3.4, and retrieve the video with the highest similarity. We conduct the procedure similarly but in the other direction for video-to-text retrieval.

Experimental Setup

We evaluate UCoFiA on five popular video-text retrieval datasets: MSR-VTT , MSVD , ActivityNet and DiDeMo .

MSR-VTT contains 10,00010,000 videos, each annotated with 2020 text captions. The video length is ranged from 1010 to 3232 seconds. Following , we train UCoFiA on 9,0009,000 videos and report the results on 1,0001,000 selected video-text pairs (the 1kA test set).

Activity-Net consists of 20,00020,000 YouTube videos with 100,000100,000 captions. The average video length is 180180 seconds. We follow to concatenate the multiple text descriptions of a video into one paragraph and perform paragraph-to-video retrieval on ActivityNet. We train our model on 10,00010,000 videos and use the ’val1’ split for evaluation which contains 5,0005,000 videos.

DiDeMo is comprised of 10,00010,000 videos and 40,00040,000 captions. The average video length is 30 seconds. Similar to ActivityNet, we evaluate paragraph-to-video retrieval on DiDeMo. There are 8,3958,395 videos in the training set, 1,0651,065 videos in the validation set, and 1,0041,004 videos in the test set. We report the results on the test set for evaluation.

MSVD contains 1,9701,970 videos, each with a length that ranges from 11 to 6262 seconds. Each video contains approximately 4040 captions. Following , we split the train, validation, and test set with 1,2001,200, 100100, and 670670 videos. We follow the common multiple caption evaluation setting in which each video in the test set is associated with multiple text captions.

VATEX consists of 34,99134,991 video clips with multiple captions per video. We follow HGR’s split protocol. There are 25,99125,991 videos in the training set, 1,5001,500 videos in the validation set and 1,5001,500 videos for evaluation.

2 Evaluation Metrics

Following , we use standard video-text retrieval metrics, including R@1, R@5, and Mean Rank (MnR) to validate the effectiveness of our UCoFiA model. We report results with more metrics (including R@10, Median Recall) in supplements.

3 Implementation Details

Experimental Results

In this section, we compare UCoFiA with several recent methods on the five video-text retrieval datasets and conduct comprehensive ablation studies to verify our design choices. We also provide a qualitative analysis to show the effectiveness of our model designs. Moreover, we display quantitative results with full metrics (R@1,5,10, MdR, MnR) for each dataset, more experiments about applying UCoFiA to more advanced backbone model , and quantitative analysis on computational cost and training strategy in the supplements.

In Table 1, we compare UCoFiA with existing methods that are either with (in the middle section of the table) or without (in the upper section of the table) CLIP on text-to-video retrieval. We also compare UCoFiA with existing methods on video-to-text retrieval in Table 2. We observe that UCoFiA achieves better performance than existing methods on most of the metrics on both text-to-video and video-to-text retrieval settings.

On MSR-VTT, compared to the recent multi-level alignment method X-CLIP , UCoFiA gives a significant 3.3%3.3\% improvement on text-to-video R@1 metric and comparable video-to-text retrieval results despite X-CLIP leveraging more levels of coarse alignments (video-sentence, video-word, frame-sentence, and frame-word) than us. This result verifies our motivation that building a coarse-to-fine alignment is useful for video-text retrieval. We also achieve 2.0%2.0\% improvement on text-to-video R@1 on VATEX dataset.

Similarly, on ActivityNet and DiDeMo datasets, we notice that UCoFiA is capable of handling longer text queries and achieves observe 5.2%5.2\% and 4.7%4.7\% improvement on paragraph-to-video retrieval compared to CLIP4Clip which only utilizes video-sentence alignment. Furthermore, we observe 1.4%1.4\% and 1.3%1.3\% gain on ActivityNet and DiDeMo compared to X-CLIP under paragraph-to-video retrieval and 2.4%2.4\% and 2.9%2.9\% gain on video-to-paragraph retrieval. These results show the importance of fine-grained correspondence even on retrieval for long videos.

Meanwhile, our method achieves comparable results on the MSVD dataset which evaluates a multiple-caption setting, which further verifies the generalizability of our model. However, we observe our normalization before the summation strategy still introduces some performance gain on MSVD even though the multiple-caption setting breaks our assumption that one video has one corresponding query, showing the robustness of our approach.

2 Ablation study

In this section, we study the different design choices of our UCoFiA model and verify their effects on the video-text retrieval performance on MSR-VTT under text-to-video retrieval setting. Specifically, we investigate (1) the effect of different alignment schemes, (2) the comparison of our fine-grained alignment design to others, (3) different similarity aggregation methods and (4) the effect of the unification module.

The Effect of Different Alignment Schemes. First, we validate the effectiveness of our different levels of alignment. As shown in Table 3, adding frame-sentence and patch-word alignments improves the base model (that only leverages video-sentence alignment) with a significant margin. Specifically, we observe that adding patch-word alignment improves the R@1 while keeping a similar R@5, which indicates that the fine-grained alignment helps the model choose the most relevant video (top-11) from several similar video candidates (top-55) by capturing the subtle differences between these video candidates. This result justifies our motivation for using fine-grained alignment as complementary to coarse alignment.

Comparison of Our Fine-grained Alignment Design to Others. In Table 4, we study the different designs of fine-grained alignment in our model. We compare our patch-word similarity with patch-sentence alignment and the ensemble of these two alignments. Based on the results, we observe that our patch-word alignment is the best design, and meanwhile, adding patch-sentence alignment degrades the performance, possibly caused by the mismatch between patch and sentence representations, where one conveys local information and the other contains global information. We also provide qualitative analysis on different alignment designs in supplements.

Different Similarity Aggregation Methods. To validate the effectiveness of our ISA module for frame-sentence alignment and Bi-ISA module for patch-word alignment in Section 3.3, we compare our modules with several other aggregation methods. In Table 5, we show the effectiveness of our interactive similarity attention (ISA) on the frame-sentence score. Specifically, we compare the vanilla mean pooling strategy and softmax-based weighted combination . Note that we only adopt video-sentence and frame-sentence alignment (remove patch-word alignment and the Bi-ISA module) for the experiments in Table 5 to study the effect of ISA independently. Results show that using the ISA module achieves better performance compared to other aggregation methods. We also compare our Bi-ISA with other aggregation methods in Table 6, where we have adopted ISA for the frame-sentence score. We observe that the design of bi-directional aggregation achieves a significant gain in performance. In general, our experiments show that adding one linear layer to learn the temporal information across frames before aggregation is crucial.

The Effect of Unification Module. To verify the importance of the unification module for different levels of similarity, we compare our UCoFiA model with the variant that removes the Sinkhorn-Knopp normalization (abbreviated as SK norm) in Table 7. We notice that adding the SK norm provides better performance on video-text retrieval with a 1.2%1.2\% gain on both R@1 and R@10.

3 Qualitative Analysis

In this section, we report the qualitative analysis of the effectiveness of the ISA module. We show the comparison of the model with and without ISA module in Figure 4. We can see from the upper part of Figure 4 that the ISA module improves the softmax weight by highlighting the most relevant frame to the query. As a result, the calculated similarity score of ISA is significantly increased since the irrelevant information is eliminated. In the lower part, by recognizing the most related frame, our model with ISA produces a lower similarity to the unmatched video. For either the ground truth or wrongly retrieved video, we find the softmax weights used in TS2-Net (without frame-wise interaction) tend to be more uniformly distributed. On the contrary, our ISA module is capable of finding and assigning the highest score to the most relevant frame (the player is making shots) through the interaction between frames, which shows video candidates and results in the correct retrieval. This analysis verifies the effectiveness of our ISA module.

Conclusion

In this paper, we present UCoFiA, which jointly considers cross-modal correspondence from different granularity and accomplishes the unification of multi-grained alignment. It achieves state-of-the-art results on multiple video-text retrieval benchmarks. UCoFiA is a simple but effective model that achieves state-of-the-art results on five diverse video-text retrieval benchmarks. In the future, we plan to extend our method to other video-language tasks such as video question answering and video reasoning.

Acknowledgment

We thank the reviewers and Shoubin Yu for their helpful discussions. This work was supported by ARO Award W911NF2110220, ONR Grant N00014-23-1-2356, DARPA KAIROS Grant FA8750-19-2-1004, NSF-AI Engage Institute DRL211263, Sony Faculty Innovation award, and Laboratory for Analytic Sciences via NC State University.

References

Appendix

In this Appendix, we present the following items:

Appendix A Additional Quantitative Results

In this section, we report additional quantitative results for our UCoFiA model. First, we report the results with full video-text retrieval metrics (including video-to-text retrieval) on MSR-VTT, ActivityNet, and DiDeMo. Results indicate our UCoFiA model achieves better results on both text-to-video and video-to-text retrieval compared to the current state-of-the-art CLIP-based approaches. Meanwhile, we show UCoFiA is capable of adapting to other advanced backbone models. Then, we compare the performance and computational cost of UCoFiA with previous work and validate our methods accomplish significant improvement with limited additional computation. Lastly, we ablate the different training settings for encoders and the model design of our bi-directional ISA module (Bi-ISA).

In this section, we report the video-text retrieval results on MSR-VTT , ActivityNet and DiDeMo with full video-text retrieval metrics, including the results on video-to-text retrieval setting.

MSR-VTT. As shown in Table 8, UCoFiA achieves state-of-the-art results on most metrics. Specifically, compared to the most recent multi-level alignment method X-CLIP , UCoFiA achieves a 3.3%3.3\% gain on text-to-video R@1 metric and obtains comparable results on video-to-text retrieval metrics. Compared to another recent state-of-the-art CLIP-based method TS2-Net , our model gets 2.4%2.4\% and 1.8%1.8\% improvement on R@1 metric for text-to-video and video-to-text retrieval. These results verify the effectiveness of the UCoFiA model. Moreover, replacing the visual backbone (ViT-3232) with a larger model (ViT-1616) would improve the model performance, especially on video-to-text retrieval.

ActivityNet. As shown in Table 9, UCoFiA outperforms the current state-of-the-art CLIP-based methods on a wide range of metrics on ActivityNet benchmark . Concretely, our model achieves 1.4%1.4\% and 2.4%2.4\% gain on the R@1 metric on text-to-video and video-to-text retrieval compared to the state-of-the-art approaches. This indicates our UCoFiA model is capable of tackling long video retrieval, thus validating the generalization ability of our method.

DiDeMo. As shown in Table 10, compared to the current state-of-the-art models, UCoFiA achieves better results on most evaluation metrics. Specifically, our model outperforms the recent state-of-the-art CLIP-based approach X-CLIP with a significant margin of 1.3%1.3\% on text-to-video R@1 and 2.9%2.9\% on video-to-text R@1.

A.2 Adapt UCoFiA to Other Backbone Model

In this section, we apply UCoFiA to the recent CLIP-ViP’s backbone model, which is a video-text model pretrained on 100100M video-text pairs. As shown in Table 11, UCoFiA improves the backbone CLIP-ViP model on all metrics on the MSR-VTT text-to-video retrieval task. This indicates that our method is able to generalize to a more advanced backbone model and verifies the robustness of our method.

A.3 The Computational Cost of UCoFiA

In this section, we compare our model with the recent X-CLIP model on the balance of model performance and computational cost in Table 12. Results show that UCoFiA is 3.3% better than X-CLIP on text-to-video retrieval on MSR-VTT dataset while only requiring 1.2% additional parameters and 1.4 GB memory per GPU (train on 44 GPUs). Therefore, UCoFiA achieves significant improvement with limited additional computational cost compared to previous works.

A.4 Different Training Settings for Encoders

A.5 Additional Ablation Study for Bi-ISA

In the main paper, we mention that empirically we find that jointly considering two directions of patch-word matrix aggregation (patch-then-word and word-then-patch) provides better aggregation for the patch-word matrix. In Table 13, we compare our bi-directional solution with single-directional methods on the MSR-VTT dataset. For better comparison, we do not apply the Sinkhorn Knopp algorithm to normalize the retrieval similarities. Results show that leveraging both aggregation directions achieves better results, validating the effectiveness of our Bi-ISA design.

Appendix B Additional Qualitative Results

In this section, we provide additional qualitative results of UCoFiA. First, we visualize the imbalanced retrieval results and show how our unification module mitigates this issue. Then, we visualize the video samples retrieved by methods focusing on different alignment levels to validate the effectiveness of our coarse-to-fine alignment design.

As discussed in the main paper, we find that scores across different videos are highly imbalanced in the similarity matrices of each level. As a result, the video candidate could be over-/under-represented by the retrieval model due to the imbalanced summation of retrieval similarities. As shown in Figure 5, the left part denotes the video candidates haven’t been retrieved in the inference stage which corresponds to under-representative. The right part denotes the video candidates have been retrieved more than twice (including twice) in the inference stage which corresponds to over-representative. The middle part denotes the video candidates have been retrieved once, which is the ideal situation. The blue column in Figure 5 represents the model without the Sinkhorn Knopp algorithm. The results show that only 43%43\% video candidates are retrieved once in the inference stage while 34%34\% video candidates are under-represented and 23%23\% video candidates are over-represented. After applying the Sinkhorn Knopp algorithm in the unification module (the orange column in Figure 5), the under-representative issue is mitigated and more than 5050 under-represented video candidates have been re-scaled and retrieved by the model. Meanwhile, we also observe a slight reduction in the number of over-represented videos. In all, the Sinkhorn Knopp algorithm in the unification module indeed mitigates the over- and under-representation issue in the inference stage.

B.2 Comparison of Different Alignments

As discussed in the main paper, our coarse-to-fine alignment module captures comprehensive cross-modal clues compared to coarse-grained or fine-grained alignment. We provide more visualization results in Figure 6. For the first text query (on the first row of Figure 6), the coarse-grained alignment only captures the scene of “singing” and the fine-grained alignment only focuses on the object “guitar”. For the second text query (on the second row of Figure 6), the coarse-grained alignment only considers the scene information like “driving”, and “video game” while the fine-grained alignment only captures the detail information “motorcycle”. For the last text query (on the last row of Figure 6), the coarse-grained alignment overlooks the detailed information “basketball” and the fine-grained alignment ignores the scene of “crowd” and the action of “run into”. To sum up, the coarse-grained or fine-grained alignment could overlook some crucial cross-modal clues while our coarse-to-fine alignment is capable of capturing both high-level and detailed information and retrieving the correct video candidate.

Appendix C Method Details

In this section, we present more details of UCoFiA. First, we discuss the patch selection module. Then, we present details of the Sinkhorn-Knopp Algorithm that normalizes the similarity matrix for unification.

As discussed in the main paper, due to the high redundancy of patch tokens, inspired by , we propose a patch selection module to choose the top-K salient patches from each frame for patch-word alignment. Here we present the details of the patch selection module.

Then, according to the saliency score UU, we select the indices of KK most salient patches within a video frame ind∈{0,1}Kind\in\{0,1\}^{K}. Through this one-hot vector indind, we extract the top-K salient patch by

C.2 Sinkhorn-Knopp Algorithm

As discussed in the main paper, inspired by , we utilize the Sinkhorn-Knopp algorithm to normalize the similarity scores for each granularity and make sure the marginal similarities (the sum of retrieval similarities between one specific video and all texts) for different videos are almost identical so that each video has a fair chance to be selected. Below, we discuss the algorithm in detail.