Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations

Peng Jin, Jinfa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David A. Clifton, Jie Chen

Introduction

Text-video retrieval Yu2017Retrieval , which aims to fetch relevant videos using textual queries or vice versa, is an important yet challenging cross-modal task. The dominating paradigm of text-video retrieval is contrastive learning wu2018unsupervised ; hjelm2018learning ; chen2020a ; he2020momentum ; chen2021an ; chen2020improved , which is a commonly adopted framework for video-and-language representation learning. The core idea of contrastive learning is to pull the textual and visual representations of matched text-video pairs together and push the representations of unmatched text-video pairs apart. In this manner, contrastive learning enables neural networks to learn discriminative video-and-language representations.

However, standard contrastive learning has intrinsic limitations for text-video retrieval tasks, since the success of contrastive learning largely depends on the volume and variety of negative samples chen2020a . Without adequate negative samples, it would be hard to guide the direction of sample learning, and the features could not be contrasted well. As shown in Figure 1a, we randomly select three videos (each video has 20 text captions) from the benchmark dataset xu2016msr and analyze the entire feature space of a recent standard contrastive learning method gabeur2020multi . Here we visualize the feature space with t-SNE maaten2014accelerating . In an ideal feature space, the instances of the same semantic classes should be close to each other. However, as shown in Figure 1a, we notice that the feature space learned by standard contrastive learning fails to preserve inter-modal semantic relatedness, where videos and texts with the same class semantics are still far away. In other words, the semantic-relevant but modal-different instances can not be grouped together in this feature space. To improve the performance of contrastive learning, recent image-text representation learning methods either increase batch sizes or maintain large data memory banks chen2020a ; chen2021an ; misra2020self ; wu2018unsupervised ; he2020momentum . These works target to collect sufficient negative samples for contrastive learning. However, such text-video retrieval approaches are relatively expensive due to the incurred extra computational costs, especially on large-scale datasets sun2019videobert . So an urgent challenge here is to find an efficient method to learn a semantically relevant feature space.

To efficiently bridge the modality gap and group visual and textual representation according to semantics, we propose to learn a low-rank and compact latent feature space. For this purpose, we propose a novel method named Expectation-Maximization Contrastive Learning (EMCL), which uses the same parametric model to abstract and reconstruct both textual and visual representations. In detail, we find a set of reconstruction bases for subspace representation by estimating the parameters in Expectation-Maximization (EM) algorithm dempster1977maximum . In the learned subspace, both video and text features are represented with the distributions over the same sets of hidden variables, which can preserve strong semantic relations across modalities. As shown in Figure 1b, our EMCL effectively learns a semantically related subspace which has smaller intra-class variance and larger inter-class variance compared to the previous contrastive learning method (Figure 1a). We further propose EMCL-Net based on the EMCL and apply EMCL-Net to the task of text-video retrieval. In particular, for better adapting the proposed EMCL into the downstream task, our EMCL-Net introduces a parameter initialization strategy of the EM algorithm. Experimental results on three text-video retrieval benchmark datasets, i.e., MSR-VTT xu2016msr , ActivityNet krishna2017dense , and LSMDC rohrbach2015a , show the advantages of the proposed EMCL. The main contributions are as follows:

We identify the intrinsic limitation of contrastive learning for text-video retrieval, i.e., the cross-modal representation bias could not be fully eliminated via standard contrastive learning approaches.

To alleviate such limitation, we reformulate the contrastive learning for video-and-language representations into an expectation-maximization iteration manner and propose a plug-and-play feature projection module named Expectation-Maximization Contrastive Learning (EMCL), which learns the subspace that aims to become semantic-relevant representation.

Based on our EMCL, we further propose the EMCL-Net, which introduces a parameter initialization strategy for the EM algorithm. Experiments show that our approach achieves state-of-the-art results on three text-video retrieval datasets. More encouragingly, our method can be easily applied to boost the performances of existing approaches either as a jointly training layer or an out-of-the-box inference module with no extra training.

Related Work

Cross-modal learning is widely studied in many areas, including cross-modal retrieval Yu2017Retrieval ; wu2018unsupervised ; hjelm2018learning , transfer learning phung2021learning ; neyshabur2020being , domain adaptation stojanov2021domain ; liang2021pareto , and captioning liu2020prophet in which different modalities/domains come from different distributions. The main challenge of cross-modal learning is to use given vision-and-language pairs to learn common representations shared between modalities liu2019MIA . Most existing works of text-video retrieval chen2020a ; he2020momentum ; gorti2022x map text and video to the same latent space, where the similarity between them can be directly calculated gabeur2022masking ; cao2022visual ; wang2022many ; dong2021dual ; wei2021universal ; wray2021semantic ; chen2021learning ; croitoruteachtext ; yang2021taco ; qi2021semantics . In detail, CE liu2019use introduces a mixture-of-experts method that mixes the features of many different pre-training experts; MMT gabeur2020multi shows a multi-modal transformer with multiple self-attended layers for video embedding. Recently, contrastive learning methods, e.g., CLIP radford2021learning , show great success in advancing the state-of-the-art performances of cross-modal tasks luo2021clip4clip . Contrastive learning methods dosovitskiy2014discriminative ; cao2022locvtp ; he2020momentum ; misra2020self ; chen2020improved ; liu2021CA try to learn data representations from positive and negative pairs, making the representations of positive pairs (usually data augmentations that retain semantic information) have high similarity, and negative pairs (different semantic examples) have low similarity. Inspired by the great success of contrastive learning, we employ CLIP radford2021learning for learning video-and-language representations. In this paper, a positive pair is a matching text-video pair, and a negative pair contains non-matching text and video. However, due to the multi-modal nature and spatial-temporal evolution of video chen2019weakly ; liu2017hierarchical , it is essential to capture the most important semantic concept for the video. Meanwhile, the visualization in Figure 1 also shows that the feature space of video-and-language representations has semantically irrelevant redundancy which leads to non-ideal contrast in the feature space. liang2022mind analyzes the non-optimal feature space with modality gap impacts model’s performance and fairness. Therefore, we propose to conduct contrastive learning in an expectation-maximization iteration manner to eliminate the non-optimal semantically redundant dimensions, acquiring compact representations.

Approach

In this section, we first introduce how to reformulate the Contrastive Learning into an expectation-maximization iteration manner, i.e., Expectation-Maximization Contrastive Learning (EMCL). Then, we introduce how to incorporate the proposed EMCL into the neural network for video-and-language representations, which are used to perform the text-video retrieval task.

Expectation-Maximization Algorithm The Expectation-Maximization (EM) algorithm dempster1977maximum is an iterative optimization strategy, which was originally designed to solve the problem when data are missing in the process of parameter estimation. Briefly, given unobserved hidden variables Z={z1,z2,...,zN}\bm{Z}=\{z_{1},z_{2},...,z_{N}\} and observed data sets X={x1,x2,...,xN}\bm{X}=\{x_{1},x_{2},...,x_{N}\} with NN samples, the goal of EM algorithm is to estimate the maximum likelihood solution of model parameters θ^=arg⁡max⁡∑i=1Nlog⁡∑zip(xi,zi;θ)\hat{\theta}=\arg\max\sum_{i=1}^{N}\log\sum_{z_{i}}p(x_{i},z_{i};\theta). In step E, the EM algorithm calculates the conditional probability expectation Qi(zi)=p(zi∣xi,θ)Q_{i}(z_{i})=p(z_{i}|x_{i},\theta). In the M step, the likelihood function is maximized to get the revised parameter θ^=arg⁡max⁡∑i=1N∑ziQi(zi)log⁡p(xi,zi;θ)Qi(zi)\hat{\theta}=\arg\max\sum_{i=1}^{N}\sum_{z_{i}}Q_{i}(z_{i})\log\frac{p(x_{i},z_{i};\theta)}{Q_{i}(z_{i})}.

Gaussian Mixture Model The Gaussian Mixture Model (GMM) richardson1997on combines multiple single Gaussian models. Assuming that GMM consists of KK Gaussians, given the input X={x1,x2,...,xN}\bm{X}=\{x_{1},x_{2},...,x_{N}\} with hidden variables Y\bm{Y}, the probability of GMM is as follows:

where N(∗∣μk,Σk)\mathcal{N}(*|\mu_{k},\Sigma_{k}) is the probability function of the kthk_{\text{th}} Gaussian and πk\pi_{k} is the prior probability.

Through the E step of the EM algorithm, the estimated value of yn,ky_{n,k} can be obtained by:

In the M step of the EM algorithm, the parameters of the Gaussian mixture model are updated as:

2 Expectation-Maximization Contrastive Learning (EMCL)

In this section, we introduce our EMCL in detail. Representing the features in low-dimensional space is a fundamental method to eliminate the inherent redundancies in features, which is also the goal of our method. We leverage these compact features for contrastive learning to efficiently bridge the gap and improve their performance. The overview of our EMCL is shown in Figure 2 and Algorithm 1.

where P(∗)P(*) is the probability function and p(∗)p(*) is a kernel function. Here we assume that xi,jx_{i,j} can be represented by KK shared Gaussian distributions like GMM richardson1997on . For simplicity, we freeze Σ\bm{\Sigma} and π\bm{\pi}, and only consider the mean λ=μ\bm{\lambda}=\bm{\mu} of the Gaussian model.

There are many choices for the kernel function p(x^,y^)p(\hat{x},\hat{y}), such as linear kernel (x^Ty^+c)(\hat{x}^{T}\hat{y}+c), polynomial kernel (ax^Ty^+c)d(a\hat{x}^{T}\hat{y}+c)^{d}, Gaussian kernel exp(−γ^∥x^−y^∥2)exp(-\hat{\gamma}\|\hat{x}-\hat{y}\|^{2}) and so on. We find that different kernel functions have slightly different impacts on the final results. For easy implementation, we use Gaussian kernel and rewrite it in a form similar to the attention model bahdanau2015neural ; vaswani2017attention . The formula is:

where σ\sigma is a hyper-parameter to adjust the distribution, similar to the mean and covariance in the Gaussian distribution. In this paper, we define x:,jx_{:,j} to be the jthj_{\textrm{th}} column of X\bm{X}, and λ:,k\lambda_{:,k} to be the kthk_{\textrm{th}} column of λ\bm{\lambda}.

Reviewing the EM algorithm dempster1977maximum , it estimates the parameters of the model by continuously executing E steps and M steps. In step E, it calculates the conditional probability expectation with current parameters. In step M, it maximizes the likelihood function to update the parameters.

In step E, referring to the Eq. (2), the estimated value of yj,ky_{j,k} can be obtained by:

In the M step, we update base λi,k\lambda_{i,k} according to Eq. (3), which is formulated as:

By repeatedly iterating step E and step M, our algorithm forces this set of optimal subspaces to represent the original video features and original text features at the same time. After several iterations, we could keep the information that appears in both the video and the text to remove the redundancy.

Feature Reconstruction When reconstructing features, we use Y\bm{Y} and λ\bm{\lambda} to linearly reconstruct the original features. Learning from the meaning of Y\bm{Y} and λ\bm{\lambda} mentioned above, yj,ky_{j,k} indicates whether the kthk_{\text{th}} subspace is selected by x:,jx_{:,j} and λ:,k\lambda_{:,k} represents the base of the kthk_{\text{th}} subspace. The formula for feature reconstruction is:

As shown in Eq. (8), we only use the corresponding base λi,k\lambda_{i,k} (for all k∈[1,K]k\in[1,K]) to reconstruct the ithi_{\text{th}} sample xi,jx_{i,j}. Therefore, we estimate different xi,jx_{i,j} from KK subspaces.

Contrastive Learning The above feature reconstruction method generates the low-rank and compact features stripped of redundancy. Therefore, we can leverage these compact features for contrastive learning to efficiently bridge the modality gap and improve task performance.

3 EMCL for Cross-Modal Learning

In this section, we discuss in detail how to incorporate the EMCL into the neural network for cross-modal downstream tasks. Targeting on the cross-modal tasks, we propose the ECML-Net, which introduces a parameter initialization strategy for the EM algorithm that can establish the connection between batches. At last, we take an important yet challenging cross-modal task, i.e., the text-video retrieval, as an example to illustrate how to employ our EMCL-Net to perform downstream tasks.

where α∈\alpha\in is the momentum. Similar to the average moving method in the BN layer, we don’t update MM in the inference stage. Because the initial value is crucial to the stability of the EM algorithm abdolali2021beyond ; li2019expectation , we need to limit the value range of the initial value. In each iteration, we use the L2 normalization operation to limit the value range of λ\bm{\lambda}.

Training Objective During the training stage, the EMCL module is trained with the neural network. We use MM to initialize the λ\lambda. Then the components YY and λ\lambda are updated in unsupervised way by iteration, as shown in Algorithm 1. Moreover, we update the initial value MM using an average moving method. The goal of the EMCL-Net is to map text and video into a joint representation space to measure the text-video similarity. Following common practice, we use cosine similarity s(t,v)=c^tTc^v∥c^t∥∥c^v∥s(t,v)=\frac{\hat{c}_{t}^{T}\hat{c}_{v}}{\left\|\hat{c}_{t}\right\|\left\|\hat{c}_{v}\right\|} as the similarity measure between text tt and video vv, where c^t\hat{c}_{t} represents the reconstructed text representation of text tt; and c^v\hat{c}_{v} represents the reconstructed video representation of video vv. We train our model with InfoNCE loss van2018representation :

where BB is the batch size and τ\tau is the temperature hyper-parameter. This loss function maximizes the similarity of positive pairs s(ti,vi)s(t_{i},v_{i}) and minimizes the similarity of negative pairs.

During the inference stage, given a set of queries (text/video) and a set of candidates (videos/texts), we use the trained MM to initialize the λ\lambda. Then YY and λ\lambda can be updated in an unsupervised way by iteration. The goal of cross-modal retrieval task is to map the query and candidate into a joint video-and-language representation space to measure the query-candidate similarity s(t,v)s(t,v). In this way, we can output the similarity scores for all the input candidates, and take the candidates with the Top-1/5/10/50 similarity scores with input query as the final prediction.

Experiments

Datasets, Metrics and Implementation Details Datasets. We conduct the experiments on three popular text-video retrieval datasets, i.e., MSR-VTT xu2016msr , ActivityNet Captions krishna2017dense , LSMDC rohrbach2015a , and follow common practice luo2021clip4clip ; cheng2021improving ; wang2022disentangled to pre-process the datasets for fair comparison. In detail, MSR-VTT xu2016msr contains 10,000 videos, each with 20 text descriptions; We follow the 1k-A split liu2019use with 9,000 videos for training and 1,000 for testing. ActivityNet Captions krishna2017dense contains 20,000 videos with multiple sentence descriptions; We report results on the "vall" split (10,009 training, 4,917 testing) as in gabeur2020multi . LSMDC rohrbach2015a contains 118,081 video clips from 202 movies; We follow the split of gabeur2020multi with 1,000 videos for testing.

Metrics. We choose standard retrieval metrics: Recall at K (R@K) and Median Rank (MdR) to evaluate the text-to-video and video-to-text retrieval performance.

Implementation Details. We utilize the CLIP (ViT-B/32) radford2021learning equipped with Temporal Transformer luo2021clip4clip as pre-trained Bi-Encoder (Base Model). Following previous works luo2021clip4clip , the frame length and caption length are 12 and 32 for MSR-VTT and LSMDC. For ActivityNet, a long video retrieval dataset, we set the frame length to 64 and caption length to 64. We follow training schedules from previous works luo2021clip4clip ; cheng2021improving ; wang2022disentangled . Concretely, we use the Adam optimizer kingma2014adam with a linear warmup. The initial learning rate is 1e-7 for text encoder and video encoder and 1e-4 for other modules. We set the temperature τ=0.01\tau=0.01, σ=1\sigma=1, the momentum α=0.9\alpha=0.9, the number of iterations is set to 9 and the parameter KK is set to 32. The network is optimized with the batch size of 128 in 5 epochs. During the training stage, the EMCL module is trained with the neural network. Each time a new batch of features is fed into the model, we use MM to initialize the λ\lambda. Then the components YY and λ\lambda are updated by iteration. Moreover, we update the initial value M using an average moving method. During the inference stage, given a set of queries (text/video) and a set of candidates (videos/texts), we use the trained MM to initialize the λ\lambda. Then the components YY and λ\lambda can be updated by iteration. When the EMCL module is incorporated into trained baselines as an out-of-the-box inference module with no extra training, λ\lambda is randomly initialized. Then the components YY and λ\lambda can be updated in an unsupervised way by iteration.

Comparisons to State-of-the-art Table 2 shows the results of our method on three text-video retrieval datasets. As we can see, our ECML-Net method consistently outperforms the recently proposed state-of-the-art methods on both text-to-video retrieval and video-to-text retrieval tasks across all datasets and metrics.

Generalization Analysis Since the input and output dimensions of the EMCL module are the same, it is model-agnostic and can be applied to features extracted from any language and video encoders. Therefore, we further equip our EMCL with three strong baseline models, i.e., MMT gabeur2020multi , CLIP4Clip luo2021clip4clip , DCR wang2022disentangled , and evaluate the performance of these models on the MSR-VTT datasets. Specifically, we insert the EMCL module at the end of the video-text encoders. Table 2 shows that our EMCL can be applied to successfully boost all baselines either as a jointly training layer or an out-of-the-box inference module with no extra training. Overall, our approach can boost the baselines with the most significant improvement up to 3.5% and 4.2% for text-to-video task and video-to-text task in terms of R@1 metric, respectively. The significant improvements demonstrate the generalization ability of EMCL.

Comparisons to Other Baseline Methods We further compare our method with other baseline methods in Table 4. PCA tipping1999probabilistic is a popular method for finding salient features shared by the two modalities so that the modality gap can be reduced to a certain extent. “Transformer” represents using a shared transformer between two modalities, which can explicitly project the data from two modalities into a shared space. “Fully Connected Layers” represents using two shared fully connected layers between two modalities. “Sparse Autoencoders” is a common sparse autoencoder ng2011sparse , which reduces the average response of the encoding layer and learns the compact representation. Different from these representative dimensionality reduction methods, the motivation of our method is to use contrastive learning to preserve video-text semantic relatedness. By maximizing the joint posterior probability of video and text, we find a semantically related subspace for compact video-text representation. In contrast, PCA and its variants are feature dimension reduction methods, which aim to maximize the variance of projected data and cannot guarantee semantic relationships. Sparse autoencoder reduces the average response of the encoding layer for sparsity, which may potentially discard semantic information. In addition, our method has linear complexity O(BDK)\mathcal{O}(BDK) and does not require additional training. In contrast, the complexity of PCA is O(D3)\mathcal{O}(D^{3}) (D>BD>B and D>KD>K), and the autoencoder requires additional training.

Ablative Analysis Effect of the parameter initialization strategy. As shown in Table 4, with or without parameter initialization strategy, our method can successfully promote the base model. It is worth noting that the EM algorithm dempster1977maximum is sensitive to initial values abdolali2021beyond ; li2019expectation . In other words, the convergence of EM algorithm depends mainly on initial parameters. Random initialization leads to large fluctuations in convergence results. This fatal flaw limits the performance of our module. Fortunately, with the proposed parameter initialization strategy, our EMCL method surpasses the base model by a large margin with 3.5% R@1 and 1.4% R@1 on the text-to-video and video-to-text tasks, respectively, proving the effectiveness of our parameter initialization strategy used for EMCL-Net.

Effect of the number of subspaces. In Figure 3a, we show the effect of the number of subspaces KK. On the one hand, we find that fewer subspaces mean fewer semantic centers, which limits our module’s ability to reconstruct the features. On the other hand, a larger number of subspaces requires more training data, which increases the cost of our model learning. We set the center size K=32K=32 to achieve the best performance in practice.

Effect of the iteration number. In Figure 3b, we show the influence of EM iteration number TT. Overall performance improves slightly before leveling off. We find that the algorithm has converged when the number of iterations is 9, so we set the number of iterations to 9 as default.

Hyper-parameter selection. The parameter σ\sigma is hyper-parameter to adjust the distribution (Eq. 5), similar to the mean and covariance in the Gaussian distribution. We evaluate the scale range setting σ∈[0.3,2.0]\sigma\in[0.3,2.0] as shown in Figure 3c. We find that R@1 is improved from 43.8% to 51.6% when σ=0.7\sigma=0.7 and saturated with σ=1\sigma=1. As a result, we adopt σ=1\sigma=1 to achieve the best performance.

Generalize to other tasks Video captioning. The purpose of video captioning is to describe the content of the video in fluent sentences. “Base modelcap” uses CLIP radford2021learning to extract video features and is trained with cross-entropy loss. To generate higher-quality sentences, we apply EMCL between video features and ground-truth text features. As shown in Table 6, using EMCL brings significant improvements on caption quality, e.g., gaining a relative improvement of 0.4% at METEOR.

Video question answering. Visual question answering requires the model to predict an answer using visual information li2022joint ; li2022toward . We use the target vocabulary for MSRVTT-QA dataset xu2017video , and train a fully connected layer on top of the final language features to classify the answer. “Base modelqa” uses CLIP radford2021learning , a transformer-based dosovitskiy2021an ; li2022locality visual-language pre-training model, to extract video-and-language features and is trained with cross-entropy loss. To learn compact video-and-language representations, we apply EMCL between video features and question features. Table 6 shows that EMCL can be applied to boost video question answering successfully and boost the baseline with an improvement up to 0.8%.

Qualitative Analysis Analysis of EMCL iterative process. As shown in Figure 4a, reducing the intra-class variance makes videos and texts belonging to the same semantic class gather, and increasing the inter-class variance makes those belonging to different semantic classes separate from each other. The two conclusions we get from Figure 4b and Figure 4c are as follows. (1) EMCL module reduces the intra-class variance of the same semantic classes and increases the inter-class variance of different semantic classes. (2) The number KK of subspaces has a great influence on the EMCL module. The effect of the EMCL module is limited if KK is too small, while the intra-class variance increases if KK is too large.

Visualization. To better understand EMCL, we provide the visualization form of both the original representations and the filtered representations. We notice that the EMCL module eliminates the redundant dimensions, and the features reconstructed by our module are very compact in the feature space. As shown in Figure 5, even though features have intense noise in the redundancy dimension, the EMCL module still learns semantic information. The EMCL module forces the differences between classes in subspace to be more obvious than original, which is helpful for the contrastive learning to learn the semantic centers shared by videos and texts.

Conclusion

In this paper, we studied the intrinsic limitation of classic contrastive learning for text-video retrieval. We found that the contrastive method in the entire representation space fails to preserve inter-modal semantic relatedness, which makes the features gather or separate in the subspace which is irrelevant to semantics. To mitigate this effect, we propose to directly learn the subspace that is related to shared semantics and do contrastive learning in it. By learning the subspaces related to semantics, we are able to learn the common semantic center of video and text in the semantic subspace. Further, it is worth noting that our method could be applied for other contrastive learning tasks, which include similar samples containing redundant dimensions or with a limited number of negative samples.

Acknowledgements This work is supported by the Nature Science Foundation of China (No. 61972217, 62081360152, 62006133, 32071459), Guangdong Basic and Applied Basic Research Foundation (No.2019B1515120049) and Guangdong Science and Technology Department (No. 2020B1111340056). Also, this work is supported in part by the National Institute for Health Research (NIHR) Oxford Biomedical Research Centre; an InnoHK Project at the Hong Kong Centre for Cerebro-cardiovascular Health Engineering; and the Pandemic Sciences Institute, University of Oxford, Oxford, UK.

References