Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu, Xiangyang Ji, Li Yuan, Jie Chen

Introduction

Representation learning based on both vision and language has many potential benefits and direct applicability to cross-modal tasks, such as text-video retrieval and video-question answering . Visual-language learning has recently boomed due to the success of contrastive learning , e.g., CLIP , to project the video and text features into a common latent space according to the semantic similarities of video-text pairs. In this manner, cross-modal contrastive learning enables networks to learn discriminative video-language representations.

The cross-modal contrastive approach typically models the cross-modal interaction via solely the global similarity of each modality. Specifically, as shown in Fig. 1a, it only exploits the coarse-grained labels of video-text pairs to learn a global semantic interaction. However, in most cases, we expect to capture fine-grained interpretable information, such as how much cross-modal alignment is helped or hindered by the interaction of a visual entity and a textual phrase. Representation that relies on cross-modal contrastive learning cannot do this in a supervised manner, as manually labeling these interpretable relationships is unavailable, especially on large-scale datasets. This suggests that there might be other learning signals that could complement and improve pure contrastive formulations.

In contrast to prior works , we model cross-modal representation learning as a multivariate cooperative game by formulating video and text as players in a cooperative game, as illustrated in Fig. 1b. Intuitively, if visual representations and textual representations have strong semantic correspondence, they tend to cooperate together and contribute to the cross-modal similarity score. Motivated by this spirit, we consider the set containing multiple representations as a coalition, and propose to quantify the trend of cooperation within a coalition via the game-theoretic interaction index, i.e., Banzhaf Interaction for its simplicity and efficiency. Banzhaf Interaction is one of the most popular concepts in cooperative games . As shown in Fig. 2, it measures the additional benefits brought by the coalition compared with the costs of the lost coalitions of these players with others. When a coalition has high Banzhaf Interaction, it will also have a high contribution to the semantic similarity. Thus, we can use Banzhaf Interaction to value possible correspondence between video frames and text words for sensitive and explainable cross-modal contrast.

To this end, we propose Hierarchical Banzhaf Interaction (HBI). Concretely, we take video frames and text words as players and the cross-modality similarity measurement as the characteristic function in the cooperative game. Then, we use the Banzhaf Interaction to represent the trend of cooperation between any set of features. Besides, to efficiently generate coalitions among game players, we propose an adaptive token merge module to cluster the original video frames (text words). By stacking token merge modules, we achieve hierarchical interaction, i.e., entity-level interactions on the frames and words, action-level interactions on the clips and phrases, and event-level interactions on the segments and paragraphs. In particular, we show that the Banzhaf Interaction index satisfies Symmetry, Dummy, Additivity, and Recursivity axiom in Sec. 3.4. This result implies that the representation learned via Banzhaf Interaction has four properties that the features of the contrastive method do not. We find that explicitly establishing the fine-grained interpretable relationships between video and text brings a sensible improvement to already very strong video-language representation learning results. Experiment results on three text-video retrieval benchmark datasets (MSRVTT , ActivityNet Captions , and DiDeMo ) and the video question answering benchmark dataset (MSRVTT-QA ) show the advantages of the proposed method. The main contributions are as follows:

To the best of our knowledge, we are the first to model video-language learning as a multivariate cooperative game process and propose a novel proxy training objective, which uses Banzhaf interaction to value possible correspondence between video frames and text words for sensitive and explainable cross-modal contrast.

Our method achieves new state-of-the-art performance on text-video retrieval benchmarks of MSRVTT, ActivityNet Captions and DiDeMo, as well as on the video-question answering task on MSRVTT-QA.

More encouragingly, our method can also serve as a visualization tool to promote the understanding of cross-modal interaction, which may have a far-reaching impact on the community.

Related Work

Cooperative Game Theory. The cooperative game theory consists of a set of players with a characteristic function . The characteristic function maps each team of players to a real number which indicates the payoff obtained by all players working together to complete the task. The core of the cooperative game theory is to allocate different payoffs to game individuals fairly and reasonably. Game theory has found many applications in the field of model interpretability , but there is little exploration in cross-modal learning. Banzhaf Interaction is one of the most popular concepts in cooperative games . Recently, LOUPE uses two-player interaction as a vision-language pre-training task. In this paper, we design a new framework of multivariate interaction for video-text representation learning. Besides, our method can be directly co-trained with target task losses for high flexibility.

Visual-Language Learning. Recently, contrastive learning methods show great success in cross-modal tasks , such as text-video retrieval and video-question answering . Text-video retrieval requires the model to map text and video to the same latent space, where the similarity between them can be directly calculated . Video-question answering requires the model to predict an answer using visual information . Due to manually labeling the fine-grained relationships being unavailable, cross-modal contrastive learning cannot capture fine-grained information in a supervised manner. To this end, we model video-text as game players with multivariate cooperative game theory and propose to combine Banzhaf Interaction with cross-modal contrastive learning. In contrast to prior works, we explicitly capture the fine-grained semantic relationships between video frames and text words via Banzhaf Interaction. Then, we use these relationships as additional learning signals to improve pure contrastive learning.

Method

Generally, given a corpus of video-text pairs (v,t)(\bm{v},\bm{t}), cross-modal representation learning aims to learn a video encoder and a text encoder. The problem is formulated as a cross-modality similarity measurement Sv,t\textrm{S}_{\bm{v},\bm{t}} by cross-modal contrastive learning, where the matched video-text pairs are close and the mismatched pairs are away from each other.

To learn fine-grained semantic alignment, the input video v\bm{v} is embedded into frame sequence Vf={vfi}i=1Nv\bm{V}_{f}=\{v^{i}_{f}\}^{N_{v}}_{i=1}, where NvN_{v} is the length of video v\bm{v}. The input text t\bm{t} is embedded into word sequence Tw={twj}j=1Nt\bm{T}_{w}=\{t^{j}_{w}\}^{N_{t}}_{j=1}, where NtN_{t} is the length of text t\bm{t}. Then, the alignment matrix is defined as: A=[aij]Nv×NtA=[a_{ij}]^{N_{v}\times N_{t}}, where aij=(vfi)Ttwj∥vfi∥∥twj∥a_{ij}=\frac{(v^{i}_{f})^{T}t^{j}_{w}}{\|v^{i}_{f}\|\|t^{j}_{w}\|} represents the alignment score between the ithi_{th} video frame and the jthj_{th} text word. For the ithi_{th} video frame, we calculate its maximum alignment score as maxj aij\underset{j}{\textrm{max}}\ a_{ij}. Then, we use the weighted average maximum alignment score over all video frames as the video-to-text similarity. Similarly, we can obtain the text-to-video similarity. The total similarity score can be defined as:

where [ωv0,ωv1,...,ωvNv]=Softmax(MLPv(Vf))[\omega_{v}^{0},\omega_{v}^{1},...,\omega_{v}^{N_{v}}]=\textrm{Softmax}(\textrm{MLP}_{v}(\bm{V}_{f})) and [ωt0,ωt1,...,ωtNt]=Softmax(MLPt(Tw))[\omega_{t}^{0},\omega_{t}^{1},...,\omega_{t}^{N_{t}}]=\textrm{Softmax}(\textrm{MLP}_{t}(\bm{T}_{w})) are the weights of the video frames and text words, respectively. Then the cross-modal contrastive loss can be formulated as:

where BB is the batch size and τ\tau is the temperature hyper-parameter. This loss function maximizes the similarity of positive pairs and minimizes the similarity of negative pairs.

Prior works typically directly apply the cross-modal contrastive loss to optimize the similarity scores Sv,t\textrm{S}_{\bm{v},\bm{t}}. To move a step further, we model video-text as game players with multivariate cooperative game theory to handle the uncertainty during fine-grained semantic interaction with diverse granularity, flexible combination, and vague intensity.

1.2 Banzhaf Interaction

We start by introducing notation and outlining assumptions about the cooperative game theory. Then, we review Banzhaf Interaction for a cooperative game.

The cooperative game theory consists of a set N={1,2,...,n}\mathcal{N}=\{1,2,...,n\} of players with a characteristic function ϕ\phi. The characteristic function ϕ\phi maps each team of players to a real number. This number indicates the payoff obtained by all players working together to complete the task. The core of the cooperative game theory is calculating how much gain is obtained and how to distribute the total gain fairly .

In a cooperative game, some players tend to form a coalition: it may happen that ϕ({i})\phi(\{i\}) and ϕ({j})\phi(\{j\}) are small, and at the same time ϕ({i,j})\phi(\{i,j\}) is large. The Banzhaf Interaction measures the additional benefits brought by the target coalition compared with the costs of the lost coalitions of these players with others. The costs of the lost coalitions can be estimated by each player in the target coalition working individually. For a coalition {i,j}\{i,j\}, we consider [{i,j}][\{i,j\}] as a single hypothetical player, which is the union of the players in {i,j}\{i,j\}. Then, the reduced game is formed by removing the individual players in {i,j}\{i,j\} from the game and adding [{i,j}][\{i,j\}] to the game.

Banzhaf Interaction . Given a coalition {i,j}⊆N\{i,j\}\subseteq\mathcal{N}, the Banzhaf Interaction I([{i,j}])\mathcal{I}([\{i,j\}]) for the player [{i,j}][\{i,j\}] is defined as:

where p(C)=12n−2p(\mathcal{C})=\frac{1}{2^{n-2}} is the likelihood of C\mathcal{C} being sampled. “N∖{i,j}\mathcal{N}\setminus\{i,j\}” denotes removing {i,j}\{i,j\} from N\mathcal{N}.

Intuitively, I([{i,j}])\mathcal{I}([\{i,j\}]) reflects the tendency of interactions inside {i,j}\{i,j\}. The higher value of I([{i,j}])\mathcal{I}([\{i,j\}]) indicates that player ii and player jj cooperate closely with each other.

1.3 Video-Text as Game Players

Given features Vf={vfi}i=1Nv\bm{V}_{f}=\{v^{i}_{f}\}^{N_{v}}_{i=1} and Tw={twj}j=1Nt\bm{T}_{w}=\{t^{j}_{w}\}^{N_{t}}_{j=1}, fine-grained cross-modal learning aims to find semantically matched video-text feature pairs. Specifically, if a video frame and a text word have strong semantic correspondence, then they tend to cooperate with each other and contribute to the fine-grained similarity score. Thus, we can consider N={vfi}i=1Nv∪{twj}j=1Nt\mathcal{N}=\{v^{i}_{f}\}^{N_{v}}_{i=1}\cup\{t^{j}_{w}\}^{N_{t}}_{j=1} as the players in the game.

To achieve the goal of the cooperative game and cross-modal learning to be completely consistent, the characteristic function ϕ\phi should meet all the following criteria: (a) the final score benefits from strongly corresponding semantic pairs {vf+,tw+}\{v^{+}_{f},t^{+}_{w}\}, i.e., ϕ(N)−ϕ(N∖{vf+,tw+}∪{[{vf+,tw+}]})\textless0\phi(\mathcal{N})-\phi(\mathcal{N}\setminus\{v^{+}_{f},t^{+}_{w}\}\cup\{[\{v^{+}_{f},t^{+}_{w}\}]\})\textless 0; (b) the final score is compromised by semantically irrelevant pairs {vf−,tw−}\{v^{-}_{f},t^{-}_{w}\}, i.e., ϕ(N)−ϕ(N∖{vf−,tw−}∪{[{vf−,tw−}]})\textgreater0\phi(\mathcal{N})-\phi(\mathcal{N}\setminus\{v^{-}_{f},t^{-}_{w}\}\cup\{[\{v^{-}_{f},t^{-}_{w}\}]\})\textgreater 0; (c) when there are no players to cooperate, the final score is zero, i.e., ϕ({vfi}i=1Nv)=ϕ({twj}j=1Nt)=ϕ(∅)=0\phi(\{v^{i}_{f}\}^{N_{v}}_{i=1})=\phi(\{t^{j}_{w}\}^{N_{t}}_{j=1})=\phi(\emptyset)=0, where ∅\emptyset denotes the empty set.

Note that anything satisfying the above conditions can be used as the characteristic function ϕ\phi. For simplicity, we use cross-modality similarity measurement S as ϕ\phi. Then, we can use Banzhaf Interaction to value possible correspondence between video frames and text words, and to enhance cross-modal representation learning.

2 Hierarchical Banzhaf Interaction

In the following, we first introduce the simple two-player interaction between a video frame and a text word. Then, we expand the two-player interaction to the multivariate interaction via the token merge module. Fig. 3 illustrates the overall framework of our method.

For a coalition {vfi,twj}\{v^{i}_{f},t^{j}_{w}\}, referring to Eq. 3, we can calculate the Banzhaf Interaction I([{vfi,twj}])\mathcal{I}([\{v^{i}_{f},t^{j}_{w}\}]). Due to the disparity in semantic similarity and interaction index, we design a prediction header to predict the fine-grained relationship Ri,j\mathcal{R}_{i,j} between the ithi_{th} video frame and the jthj_{th} text word. The prediction header consists of a convolutional layer for encoding, a self-attention module for capturing global interaction, and a convolutional layer for decoding. We provide the experiment results of the prediction header with different structures in Tab. 6.

Then, we optimize the Kullback-Leibler (KL) divergence between the I([{vfi,twj}])\mathcal{I}([\{v^{i}_{f},t^{j}_{w}\}]) and Ri,j\mathcal{R}_{i,j}. Concretely, we define the probability distribution of the video-to-text task and the text-to-video task as:

where pi,jI=exp(I([{vfi,twj}]))∑k=1Ntexp(I([{vfi,twk}])),p^i,jI=exp(I([{vfi,twj}]))∑k=1Nvexp(I([{vfk,twj}]))p_{i,j}^{\mathcal{I}}\footnotesize{=}\frac{\textrm{exp}(\mathcal{I}([\{v^{i}_{f},t^{j}_{w}\}]))}{\sum_{k=1}^{N_{t}}\textrm{exp}(\mathcal{I}([\{v^{i}_{f},t^{k}_{w}\}]))},\hat{p}_{i,j}^{\mathcal{I}}\footnotesize{=}\frac{\textrm{exp}(\mathcal{I}([\{v^{i}_{f},t^{j}_{w}\}]))}{\sum_{k=1}^{N_{v}}\textrm{exp}(\mathcal{I}([\{v^{k}_{f},t^{j}_{w}\}]))}. Similarly, the probability distribution Dv2tR\mathcal{D}_{v2t}^{\mathcal{R}} and Dt2vR\mathcal{D}_{t2v}^{\mathcal{R}} are calculated in the same way using Ri,j\mathcal{R}^{i,j}, i.e., Dv2tR=[pi,1R,pi,2R,...,pi,NtR],Dt2vR=[p^1,jR,p^2,jR,...,p^Nv,jR]\mathcal{D}_{v2t}^{\mathcal{R}}=[p_{i,1}^{\mathcal{R}},p_{i,2}^{\mathcal{R}},...,p_{i,N_{t}}^{\mathcal{R}}],\mathcal{D}_{t2v}^{\mathcal{R}}=[\hat{p}_{1,j}^{\mathcal{R}},\hat{p}_{2,j}^{\mathcal{R}},...,\hat{p}_{N_{v},j}^{\mathcal{R}}], where pi,jR=exp(Ri,j)∑k=1Ntexp(Ri,k),p^i,jR=exp(Ri,j)∑k=1Nvexp(Rk,j)p_{i,j}^{\mathcal{R}}\footnotesize{=}\frac{\textrm{exp}(\mathcal{R}_{i,j})}{\sum_{k=1}^{N_{t}}\textrm{exp}(\mathcal{R}_{i,k})},\hat{p}_{i,j}^{\mathcal{R}}\footnotesize{=}\frac{\textrm{exp}(\mathcal{R}_{i,j})}{\sum_{k=1}^{N_{v}}\textrm{exp}(\mathcal{R}_{k,j})}. Finally, the Banzhaf Interaction loss LI\mathcal{L}_{I} is defined as:

The Banzhaf Interaction loss LI\mathcal{L}_{I} brings the probability distributions of the output R\mathcal{R} of the prediction header and Banzhaf Interaction I\mathcal{I} close together to establish fine-grained semantic alignment between video frames and text words. In particular, it can be directly removed during inference, rendering an efficient and semantics-sensitive model.

For multivariate interaction, an intuitive method is to compute Banzhaf Interaction on any candidate set of visual frames and text words directly. However, the number of candidate sets is too large, i.e., 2Nv+Nt2^{N_{v}+N_{t}}. To reduce the number of candidate sets, we cluster the original visual (textual) tokens and compute the Banzhaf Interaction between the merged tokens. By stacking token merge modules, we get cross-modal interaction efficiently at different semantic levels, i.e., entity-level interactions on the frames and words, action-level interactions on the clips and phrases, and event-level interactions on the segments and paragraphs. Fig. 4 illustrates the framework of the token merge module.

Specifically, we utilize DPC-KNN , a k-nearest neighbor-based density peaks clustering algorithm, to cluster the visual (textual) tokens. Starting with the frame-level tokens Vf={vfi}i=1Nv\bm{V}_{f}=\{v^{i}_{f}\}^{N_{v}}_{i=1}, we first use a one-dimensional convolutional layer to enhance the temporal information between tokens. Then, we compute the local density ρi\rho_{i} of each token vfiv^{i}_{f} according to its KK-nearest neighbors:

where KNN(vfi)\textrm{KNN}(v^{i}_{f}) is the KK-nearest neighbors of vfiv^{i}_{f}. After that, we compute the distance index δi\delta_{i} of each token vfiv^{i}_{f}:

Intuitively, ρ\rho denotes the local density of tokens, and δ\delta represents the distance from other high-density tokens.

We consider those tokens with relatively high ρi×δi\rho_{i}\times\delta_{i} as cluster centers, and then assign other tokens to the nearest cluster center according to the Euclidean distances. Inspired by , we use the weighted average tokens of each cluster to represent the corresponding cluster, where the weight W ⁣= ⁣Softmax(MLPw(Vf))W\!=\!\textrm{Softmax}(\textrm{MLP}_{w}(\bm{V}_{f})). Then, we feed the weighted average tokens as queries QQ and the original tokens as keys KK and values VV into an attention module. We treat the output of the attention module as features at a higher semantic level than the entity level, that is, the action-level visual tokens. Similarly, we merge the action-level tokens again to get the event-level tokens. The action-level textual tokens and event-level textual tokens are calculated in the same way.

3 Training Objective

Combining the cross-modal contrastive loss LC\mathcal{L}_{C} and Banzhaf Interaction loss LI\mathcal{L}_{I}, the full objective of semantic alignment can be formulated as L=LC+αLI\mathcal{L}=\mathcal{L}_{C}+\alpha\mathcal{L}_{I}, where α\alpha is the trade-off hyper-parameter. We train the network at three semantic levels, which are shown as follows,

where Le\mathcal{L}^{e}, La\mathcal{L}^{a}, and Lo\mathcal{L}^{o} represent the semantic alignment loss at the entity level, action level, and event level, respectively.

To further improve the generalization ability, we optimize the additional KL divergence between the distribution among different semantic levels. We find that the entity-level similarity Sv,te\textrm{S}^{e}_{\bm{v},\bm{t}} converges first in the training process, so we distill the entity-level similarity to the other two semantic levels. The analyses and experiments are provided in Appendix.

Starting with entity-level similarity Sv,te\textrm{S}^{e}_{\bm{v},\bm{t}} distilling to action-level similarity Sv,ta\textrm{S}^{a}_{\bm{v},\bm{t}}, we first compute the distribution Dv2te\mathcal{D}_{v2t}^{e} and Dt2ve\mathcal{D}_{t2v}^{e} by replacing I([{v,t}])\mathcal{I}([\{v,t\}]) with Sv,te\textrm{S}^{e}_{\bm{v},\bm{t}} in Eq. 4. The distribution Dv2ta\mathcal{D}_{v2t}^{a} and Dt2va\mathcal{D}_{t2v}^{a} are calculated using Sv,ta\textrm{S}^{a}_{\bm{v},\bm{t}}. The LDe2a\mathcal{L}_{D}^{e2a} loss is defined as:

The LDe2o\mathcal{L}_{D}^{e2o} loss from entity-level similarity to event-level similarity is calculated in the same way.

The overall loss is the combination of semantically alignment losses and self-distillation losses, which is defined as:

where β\beta is the trade-off hyper-parameter. We provide the ablation experiments for each part of the loss function in Tab. 7. We find that Banzhaf Interaction loss LI\mathcal{L}_{I} significantly improves the performance, while deep supervision and self-distillation can improve the generalization ability.

4 Theoretical Analysis

Similar to Banzhaf value axioms , the following axioms convey intuitive properties that a cross-modal interaction score should satisfy.

Symmetry states that if changing the value of two coalitions has the same effect on the output under all values of the other variables, then both coalitions should have an identical interaction score. Dummy states that if changing the value of a coalition [C][\mathcal{C}] has no effect on the output under all values of other variables, then the interaction value of [C][\mathcal{C}] should be zero. Additivity states the sum of the interaction scores of the two characteristic functions is equal to the interaction score of the sum of these characteristic functions. Recursivity states that if the interaction is positive, then the interaction score of [{i,j}][\{i,j\}] should be greater than simply the sum of individual values. If the interaction is negative, the interaction score of [{i,j}][\{i,j\}] should be less than the sum.

The Banzhaf Interaction index satisfies Symmetry, Dummy, Additivity and Recursivity axiom.

We refer the reader to Appendix for more detail about Theorem 1. This result implies that the representation learned via Banzhaf Interaction has four properties that the features of the contrastive method do not. Besides, we compare Banzhaf Interaction and cosine similarity in Tab. 1, mainly in three aspects. (1) Global receptive field. In contrast to cosine similarity, which only operates at the element level, Banzhaf Interaction operates at the set level to leverage the global context. (2) Robustness. Cosine similarity fluctuates by visual and language style. In contrast, Banzhaf Interaction measures the relative value of benefit and opportunity cost to be robust to the style deviation. (3) Flexibility. Our framework can use other characteristic functions ϕ\phi besides similarity, which is left for future work to explore. Therefore, Banzhaf Interaction is a promising interaction score to enhance cross-modal representation learning.

Experiments

Datasets. MSRVTT contains 10K YouTube videos, each with 20 text descriptions. We follow the training protocol in and evaluate on the 1K-A testing split . ActivityNet Captions consists of densely annotated temporal segments of 20K YouTube videos. We use the 10K training split to train the model and report the performance on the 5K “val1” split. DiDeMo contains 10K videos annotated 40K text descriptions. We follow the training and evaluation protocol in . MSRVTT-QA is based on the MSRVTT and has 243K VideoQA pairs.

Metrics. We choose Recall at rank K (R@K), Median Rank (MdR), and mean rank (MnR) to evaluate the retrieval performance. We choose answer accuracy to evaluate the video question answering performance.

Implementation Details. Since the calculation of the exact Banzhaf Interaction is an NP-hard problem , existing methods mainly use sampling-based methods to obtain unbiased estimates. To speed up the computation of Banzhaf Interaction for many data instances, we pre-train a tiny model to learn a mapping from a set of input features to a result using MSE loss. The tiny model consists of 2 CNN layers and a self-attention layer. The input is the similarity matrix of video frames and text tokens, and the output is the estimation of Banzhaf Interaction. We refer the reader to Appendix for the details. For text-video retrieval, we utilize the CLIP (ViT-B/32) as the pre-trained model. For video question answering, we use the target vocabulary and train a fully connected layer on top of the final language features to classify the answer. More details are in the Appendix.

2 Comparison with State-of-the-art

In Tab. 2, we show the results of our method on MSRVTT, ActivityNet Captions, and DiDeMo datasets. Our model consistently outperforms the recently proposed state-of-the-art methods on both text-to-video retrieval and video-to-text retrieval tasks. Tab. 6 shows the results of our method for video-question answering. Massive experiments on text-video retrieval and video-question answering tasks demonstrate the superiority and flexibility of our method.

3 Ablation Study

Effect of the prediction header of R\mathcal{R}. To explore the impact of the structure of the prediction header on our method, we compare several popular structures in Tab. 6. We find that the combination of CNN and attention (“CNN+SA”) can capture both local and global interaction, so it is beneficial for predicting the fine-grained relationship.

Ablation about components. As shown in Tab. 7, Banzhaf Interaction boosts the baseline with the improvement up to 0.8% at R@1. Moreover, deep supervision and self-distillation significantly improve the generalization ability. Our full model achieves the best performance and outperforms the baseline by 2.0% at R@1 for text-to-video retrieval. This demonstrates that the three parts are beneficial for aligning videos and texts.

The efficiency of the cluster module. The ablation results are provided in Tab. 7. Nv−N_{v}^{-} and Nt−N_{t}^{-} denote the number of visual and textual clusters, respectively. The first row represents the baseline without the cluster module. We find that large numbers of clusters may make similar tokens classified in different clusters. From Tab. 7, we take the {Nva,Nvo,Nta,Nto}\{N^{a}_{v},N^{o}_{v},N^{a}_{t},N^{o}_{t}\} as {3,2,6,3}\{3,2,6,3\} to get the best performance on the sum of recall at rank {1,5,10}\{1,5,10\} (Rsum).

The efficiency of our method. In Tab. 7, we calculate iteration time and inference time using two Tesla V100 GPUs on MSRVTT dataset. Since the Banzhaf Interaction can be removed during inference, our method only takes additional 1s for processing the test set. This result demonstrates the superiority of our efficient design.

Parameter sensitivity. The parameter α\alpha is the hyper-parameter that trades off LC\mathcal{L}_{C} and LI\mathcal{L}_{I}. We evaluate the scale range setting α∈[0.3,1.7]\alpha\in[0.3,1.7] as shown in Fig. 6a. From Fig. 6a, we adopt α=1.0\alpha=1.0 to achieve the best performance. In Fig. 6b, we show the influence of the hyper-parameter β\beta. We evaluate the scale range setting β∈[0.5,3.5]\beta\in[0.5,3.5]. We find that the model achieves the best performance at β=2.0\beta=2.0, so we set β=2.0\beta=2.0 as default in practice.

4 Qualitative Analysis

To better understand the proposed method, we show the visualization of the hierarchical interaction in Fig. 5. We find that the semantic similarities between coalitions are generally higher than the semantic similarities between individual frames and individual words. For example, the coalition “{two, men, talking, after, a}” has a high semantic similarity with the video coalition representing the men talking action. On the contrary, when these words interact with the corresponding frame as individuals, they show low semantic similarity. Interestingly, the model uses the word “fire” instead of the phrase “one puts out a fire” to understand the video-text pair. This is due to insufficient training data, the model can not understand the low-frequency phrase. The visualization illustrates that the proposed method can be used as a tool for visualizing the cross-modal interaction and help us understand the cross-modal model.

Conclusion

In this paper, we creatively model cross-modal representation learning as a multivariate cooperative game by formulating video and text as players in a cooperative game. Specifically, we propose Hierarchical Banzhaf Interaction (HBI) to value possible correspondence between video frames and text words for sensitive and explainable cross-modal contrast. Although manually labeling the fine-grained relationships between videos and text is unavailable, our method shows a promising alternative to obtaining fine-grained labels based on Banzhaf Interaction. More encouragingly, our method can also serve as a visualization tool to promote the understanding of cross-modal interaction.

Acknowledgements. This work was supported in part by the National Key R&D Program of China (No. 2022ZD0118201), Natural Science Foundation of China (No. 61972217, 32071459, 62176249, 62006133, 62271465), and the Natural Science Foundation of Guangdong Province in China (No. 2019B1515120049).

References

Appendix A Datasets and Implementation Details

MSRVTT. MSRVTT contains 10K YouTube videos, each with 20 text descriptions. We follow the training protocol in and evaluate on text-to-video and video-to-text search tasks on the 1K-A testing split with 1K video or text candidates defined by .

ActivityNet Captions. ActivityNet Captions consists densely annotated temporal segments of 20K YouTube videos. Following , we concatenate descriptions of segments in a video to construct “video-paragraph” for retrieval. We use the 10K training split to finetune the model and report the performance on the 5K “val1” split.

DiDeMo. DiDeMo contains 10K videos annotated 40K text descriptions. We concatenate descriptions of segments in a video to construct “video-paragraph” for retrieval. We follow the training and evaluation protocol in .

MSRVTT-QA. MSRVTT-QA is based on the MSRVTT dataset and has 243K VideoQA pairs.

A.2 Implementation Details

For fair comparisons, we follow common practice to extract the video representations of input videos and the language representations of input texts. In detail, for video representations, we first extract the frames from the video clip as the input sequence of video. Then we use ViT to encode the frame sequence, by exploiting the transformer architecture to model the interactions between image patches. Followed by the CLIP , the output from the [class] token is used as the frame embedding. Finally, we obtain the video representation Vf={vfi}i=1Nv\bm{V}_{f}=\{v^{i}_{f}\}^{N_{v}}_{i=1}. For text representation, we directly use the text encoder of CLIP to acquire the text representation Tw={twj}j=1Nt\bm{T}_{w}=\{t^{j}_{w}\}^{N_{t}}_{j=1}.

The dimension of the feature is 512. The temporal transformer is composed of 4-layer blocks, each including 8 heads and 512 hidden channels. The temporal position embedding and parameters are initialized from the CLIP’s text encoder. We use the Adam optimizer and set the temperature τ\tau to 0.01. The initial learning rate is 1e-7 for text encoder and video encoder and 1e-3 for other modules.

For text-video retrieval, we utilize the CLIP (ViT-B/32) as the pre-trained model. The frame length and caption length are 12 and 24 for MSRVTT. The network is optimized with the batch size of 128 in 5 epochs. We set the caption length to 64 for ActivityNet Captions and DiDeMo.

For video question answering , we use the target vocabulary and train a fully connected layer on top of the final language features to classify the answer. The frame length and question length are 12 and 32 for MSRVTT-QA. The network is optimized with the batch size of 32 in 5 epochs.

Appendix B Proof of Theorem 1

We start by reviewing Banzhaf Values and Banzhaf Interaction for a cooperative game.

where p(C)=12n−1p(\mathcal{C})=\frac{1}{2^{n-1}} is the likelihood of C\mathcal{C} being sampled. “N∖{i}\mathcal{N}\setminus\{i\}” denotes removing {i}\{i\} from N\mathcal{N}.

Banzhaf Interaction. In a cooperative game, some players tend to form a coalition: it may happen that ϕ({i})\phi(\{i\}) and ϕ({j})\phi(\{j\}) are small and at the same time ϕ({i,j})\phi(\{i,j\}) is large. The Banzhaf Interaction measures the additional benefits brought by the coalition compared with the costs of the lost interactions of these players with others. For a coalition {i,j}\{i,j\}, we consider [{i,j}][\{i,j\}] as a single hypothetical player, which is the union of the players in {i,j}\{i,j\}. Then, the reduced game is formed by removing the individual players in {i,j}\{i,j\} from the game and adding [{i,j}][\{i,j\}] to the game.

Banzhaf Interaction . Given a coalition {i,j}⊆N\{i,j\}\subseteq\mathcal{N}, the Banzhaf Interaction I([{i,j}])\mathcal{I}([\{i,j\}]) for the player [{i,j}][\{i,j\}] is defined as:

where p(C)=12n−2p(\mathcal{C})=\frac{1}{2^{n-2}} is the likelihood of C\mathcal{C} being sampled. “N∖{i,j}\mathcal{N}\setminus\{i,j\}” denotes removing {i,j}\{i,j\} from N\mathcal{N}.

Similar to Banzhaf value axioms , the following axioms convey intuitive properties that a cross-modal interaction score should satisfy.

The Banzhaf Interaction index satisfies Symmetry, Dummy, Additivity and Recursivity axiom.

Symmetry states that if changing the value of two coalitions has the same effect on the output under all values of the other variables, then both coalitions should have an identical interaction score.

Proof. We consider C={i,j},C′={i′,j′}\mathcal{C}=\{i,j\},\mathcal{C^{{}^{\prime}}}=\{i^{{}^{\prime}},j^{{}^{\prime}}\} fixed. Let us choose T⊆N\mathcal{T}\subseteq\mathcal{N}, and consider the unanimity game. Clearly, ϕ(T∪{[{i,j}})−ϕ(T∪{[{i′,j′}]})=0,ϕ(T∪i)−ϕ(T∪i′)=0,ϕ(T∪j)−ϕ(T∪j′)=0\phi(\mathcal{T}\cup\{[\{i,j\}\})-\phi(\mathcal{T}\cup\{[\{i^{{}^{\prime}},j^{{}^{\prime}}\}]\})=0,\phi(\mathcal{T}\cup{i})-\phi(\mathcal{T}\cup{i^{{}^{\prime}}})=0,\phi(\mathcal{T}\cup{j})-\phi(\mathcal{T}\cup{j^{{}^{\prime}}})=0. That is, for every T⊆N\mathcal{T}\subseteq\mathcal{N}, C={i,j}\mathcal{C}=\{i,j\} and C′={i′,j′}\mathcal{C^{{}^{\prime}}}=\{i^{{}^{\prime}},j^{{}^{\prime}}\} produce the same benefits. Thus, Banzhaf Interaction satisfies Symmetry axiom, i.e., I([C])=I([C′])\mathcal{I}([\mathcal{C}])=\mathcal{I}([\mathcal{C^{{}^{\prime}}}]).

B.2 Dummy Axiom

Dummy states that if changing the value of a coalition [C][\mathcal{C}] has no effect on the output under all values of other variables, then the interaction value of [C][\mathcal{C}] should be zero.

Proof. We consider C={i,j}\mathcal{C}=\{i,j\} fixed. Let us choose T⊆N\mathcal{T}\subseteq\mathcal{N}, and consider the unanimity game. Clearly, ϕ(S∪{[C]})−ϕ(S)=0,∑i∈Cϕ(S∪i)=0\phi(\mathcal{S}\cup\{[\mathcal{C}]\})-\phi(\mathcal{S})=0,\sum_{i\in\mathcal{C}}\phi(\mathcal{S}\cup{i})=0. For every T⊆N\mathcal{T}\subseteq\mathcal{N}, C={i,j}\mathcal{C}=\{i,j\} has no interaction with any player. Thus, Banzhaf Interaction satisfies Dummy axiom, i.e., I([C])=0\mathcal{I}([\mathcal{C}])=0.

B.3 Additivity Axiom

Additivity states the sum of the interaction scores of the two characteristic functions is equal to the interaction score of the sum of these characteristic functions.

Proof. Let us choose T⊆N\mathcal{T}\subseteq\mathcal{N}, and consider the unanimity game. Clearly, for the characteristic function Φ(∗)=ϕ(∗)+ϕ′(∗)\Phi(*)=\phi(*)+\phi^{{}^{\prime}}(*), Φ(T)=ϕ(T)+ϕ′(T)\Phi(\mathcal{T})=\phi(\mathcal{T})+\phi^{{}^{\prime}}(\mathcal{T}). That is, for every T⊆N\mathcal{T}\subseteq\mathcal{N}, the sum of the scores of the two characteristic functions (ϕ(∗),ϕ′(∗)\phi(*),\phi^{{}^{\prime}}(*)) is equal to the score of the sum of these characteristic functions Φ(∗)\Phi(*). Thus, Banzhaf Interaction satisfies Additivity axiom.

B.4 Recursivity Axiom

We hypothesize that the interaction score should depend on the values of ii when jj is absent, and jj when ii is absent. And somehow, their interaction should also be taken into account. Specifically, Recursivity states that if the interaction is positive, then the interaction score of [{i,j}][\{i,j\}] should be greater than simply the sum of individual values. If the interaction is negative, the interaction score of [{i,j}][\{i,j\}] should be less than the sum.

Proof. We can rewrite Eq. B as I([C])=B([C]∣N∖C∪{[C]})−∑i∈CB(i∣N∖C∪{i})\mathcal{I}([\mathcal{C}])=\mathcal{B}([\mathcal{C}]|\mathcal{N}\setminus\mathcal{C}\cup\{[\mathcal{\mathcal{C}}]\})-\sum_{i\in\mathcal{C}}\mathcal{B}(i|\mathcal{N}\setminus\mathcal{C}\cup\{i\}). Clearly, the above formula is equivalent to Recursivity axiom. Thus, Banzhaf Interaction satisfies Recursivity axiom.

Appendix C Discussions

Since the calculation of the exact Banzhaf Interaction is an NP-hard problem , existing methods mainly use sampling-based methods to obtain unbiased estimates. To speed up the computation of Banzhaf Interaction for many data instances, we pre-train a tiny model to learn a mapping from a set of input features to a result using MSE loss. The tiny model consists of a convolutional layer for encoding features, a self-attention module for capturing global interaction, and a convolutional layer for decoding. The tiny model has 64 hidden channels. The input is the similarity matrix of video frames and text tokens, and the output is the estimation of Banzhaf Interaction.

To explore the impact of the Banzhaf Interaction estimator on our method, we compare the sampling-based method and pre-trained tiny model estimator in Tab. A. Given the costly training time, the ablation study is based on a subset of MSRVTT dataset (3K videos, each with 20 text descriptions). We find that the pre-trained tiny model maintains the estimation accuracy while avoiding intensive computations. The average training time is reduced from 19.79 seconds per iteration to 3.14 seconds per iteration.

C.2 The Structure of the Prediction Header

Due to the disparity in semantic similarity and interaction index, we design a prediction header to predict the fine-grained relationship Ri,j\mathcal{R}_{i,j} between the ithi_{th} video frame and the jthj_{th} text word. To explore the impact of the structure of the prediction header on our method, we compare four popular structures, i.e., “MLP”, “CNN”, “MLP+SA” and “CNN+SA”. Fig. A illustrates the structures.

C.3 Self-Distillation

“MLP” consists of a linear layer with a Relu activation function for encoding features and a linear layer for decoding. The dimension of the hidden channels is 64. “CNN” consists of a convolutional layer with a Relu activation function for encoding features and a convolutional layer for decoding. The dimension of the hidden channels is 64. “MLP+SA” consists of a linear layer with a Relu activation function for encoding features, a self-attention module for capturing global interaction, and a linear layer for decoding. The dimension of the hidden channels is 64. “CNN+SA” consists of a convolutional layer with a Relu activation function for encoding features, a self-attention module for capturing global interaction, and a convolutional layer for decoding. The dimension of the hidden channels is 64.

As shown in Tab. B, we find that the combination of CNN and attention (“CNN+SA”) can capture both local and global interaction, so it is beneficial for predicting the fine-grained relationship between video and text. As a result, we adopt “CNN+SA” to achieve the best performance.

Fig. B shows the performance of each semantic level. We find that the entity level converges first in the training process. This is because higher-level semantic features are merged from lower-level semantic features. When lower-level semantic features do not converge, it is difficult for higher-level semantic features to learn semantic information. Based on this observation, we propose using lower-level semantic features to guide the learning of higher-level semantic features. Thus, we distill the entity-level similarity to the other two semantic levels.

To illustrate the impact of the self-distillation of our method, we conduct ablation experiments on MSRVTT dataset in Tab. C. As we can see, self-distillation improves the generalization ability. Distilling from the entity level to the other two semantic levels achieves the best results. As a result, we distill the entity-level similarity to the other two semantic levels as default in practice.

C.4 Ablation for Video-Question Answering Task

To illustrate the importance of each part of our method for the video-question answering, we conduct ablation experiments on MSRVTT-QA dataset in Tab. D. As we can see, Banzhaf Interaction boosts the baseline with the improvement up to 0.6% at Top1 accuracy. Moreover, deep supervision and self-distillation significantly improve the generalization ability. Self-distillation provides limited improvement for video-question answering compared to text-video retrieval. This is because reasoning relies primarily on high-level semantic features. Therefore, it is difficult for low-level semantic features to guide high-level semantic features. Our full model achieves the best performance and outperforms the baseline by 1.0% at Top1 accuracy.

C.5 The Number of Semantic Levels

To efficiently generate coalitions among game players, we cluster the original visual (textual) tokens and compute the Banzhaf Interaction between the merged tokens. By stacking token merge modules, we get cross-modal interaction efficiently at different semantic levels.

To explore the impact of the number of semantic levels on our method, we conduct ablation experiments on MSRVTT dataset in Tab. E. We find that the performance of the model increases with the number of semantic levels. These results indicate that stacking more token merge modules can provide more coalitions, which enables the model to learn more diverse semantic interaction information. We make a trade-off between the number of semantic levels and computation cost and set the number of semantic levels to 3 in practice.

C.6 Limitations of our Work

The cross-modal contrastive approach typically exploits the coarse-grained labels of video-text pairs to learn a global semantic interaction. To move a step further, we model video-text as game players with multivariate cooperative game theory to handle the uncertainty during fine-grained semantic interaction with diverse granularity, flexible combination, and vague intensity. Therefore, our method inevitably requires more training time costs. Although our method takes less time than TS2-Net during the inference stage (see Tab. 5 in the main paper), more effort could be paid to obtain an efficient structure in the future.

Appendix D Visualizations

We show two retrieval examples from the MSR-VTT testing set for text-to-video retrieval in Fig. C. As shown in Fig. C, our method successfully retrieves the ground-truth video. These results demonstrate that our method can align video and text effectively.

D.2 Video-Question Answering

We show the visualization of the video-question answering in Fig. D. As shown in Fig. D, our method succeeds in getting the ground-truth answer. These results demonstrate that our method can deal with cross-modal inference task effectively.

D.3 Hierarchical Interaction

To better understand the proposed method, we show the visualization of the hierarchical interaction in Fig. E, Fig. F and Fig. G. This experiment shows that our Hierarchical Banzhaf Interaction (HBI) can effectively handle fine-grained semantic interaction with diverse granularity, flexible combination, and vague intensity. More encouragingly, the visualization illustrates that the proposed method can be used as a tool for visualizing the cross-modal interaction and help us understand the cross-modal model.