TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
Jianwei Yang, Yonatan Bisk, Jianfeng Gao
Introduction
Aligning or grounding language to videos is a challenging topic in the context of vision-language (VL) research as it requires the model to understand contents, dynamics, and causality presented in videos . Inspired by the success of BERT in natural language processing, there is a growing interest in applying transformer-based multi-modal models for video-text alignment and representation learning . These models are typically pretrained on large amounts of noisy video-text pairs using contrastive learning , and then applied in a zero-shot manner or finetuned for various downstream tasks, such as text-video retrieval , video action step localization , video action segmentation , video question answering and video captioning .
In this paper, we present a new variant of contrastive learning, Token-Aware Cascade contrastive learning (TACo) to improve the video-text alignment for both large-scale pretraining and downstream specific tasks. As the name indicates, TACo makes two modifications to the conventional contrastive learning used in video-language domain. The first is the token-aware contrastive loss which is computed by taking into account the syntactic classes of words. This is motivated by the observation that, given a video and its corresponding text, content words, such as nouns and verbs, are more likely than function words to be aligned with (or grounded to) visual contents in the video. Conventional contrastive learning typically compute the loss after aggregating over all the words in the text and frames in the video (loss or in Fig. 1). In contrast, the token-aware contrastive loss is computed using only a subset of words whose syntactic classes belong to a pre-defined set (e.g., nouns and verbs), which forces the grounding of individual words to the video (loss ). For example, we pay particular attention to the words “add”, “tomatos”, “pan” and “stir” in Fig. 1.
The second technique we introduce is a cascade sampling method to find a small set of hard negative examples for training the multi-modal fusion layers. Consider a batch of video-text pairs. For each of the video-text pairs, the ideal case is that we use the remaining negative videos or texts to compute the contrastive loss after multi-modal fusion. However, the cost of computing the contrastive loss quickly becomes prohibitive when it is coupled with multi-modal fusion layers, considering its high complexity where is total number of visual and textual tokens. A conventional way to address this is using random sampling to select a small subset of negative pairs. In this paper, instead of random sampling, we propose a cascade sampling method as shown in the top-right of Fig. 1 to efficiently select a small set of hard negative examples on the fly during training. It leverages the video-text alignment scores computed in and before multi-modal fusion layers, and helps to learn the multi-modal fusion layers more effectively without any extra overhead.
We perform a comprehensive empirical study to validate the effectiveness of TACo in both pretraining and dataset-specific scenarios. We apply TACo and different variants of contrastive losses to train or pretrain and finetune on various downstream tasks including text-video retrieval (YouCook2, MSR-VTT and ActivityNet) , video action step localization (CrossTask) and action segmentation (COIN) . Our results show that TACo improves the text-video retrieval performance over current state-of-the-art across three benchmarks. Furthermore, the learned multi-modal representation and video representation can be effectively transferred to CrossTask and COIN, and achieve better or comparable performance to current state-of-the-art methods.
Related work
Video-language pretraining. Realistic application scenarios around videos have prompted emergence of various video-language tasks, such as text-video retrieval , video question answering , video captioning , etc. Inspired by the success of BERT for large-scale pretraining in language domain , transformers have been employed in the video-language domain as well as image-language domain . Combined with large scale datasets, e.g. Howto100M this approach has proven to be effective on various downstream tasks. Depending on the tasks of interest, some approaches train a multi-modal transformer using a combination of multiple losses including video-text alignment , masked token (words/frames/objects) prediction , and frame order prediction , etc. Some other approaches exploited various contrastive learning techniques to directly optimize the feature space without multi-modal fusion . In most of previous works, these two approaches were explored separately. Very recently, an updated version of used two independent alignment losses before and after multi-modal fusion in a single framework. In this paper, however, these two losses cooperate closely with each other during training in that the earlier stage helps to discover the hard negatives while the multi-modal layers with more capacity help to tackle those hard samples particularly.
Video-text alignment. Aligning videos to text requires the model to understand motion and temporal coherence. Some works have relied on attention mechanisms to extract key information from videos , while others preserve visual information by composing pairwise joint representation using 3D tensors or use multi-level video encoders to separately encode the spatial and temporal cues . These models usually rely on a rank or margin loss to learn the correct alignment for video-text pairs. Another line of work learns fine-grained or hierarchical alignment between videos and texts . In , the authors proposed a fine-grained alignment by extracting the nouns and verbs from action phrase in a sentence and projecting them into a shared space with videos. Alternatively, the authors in extract a hierarchical semantic graph and apply graph reasoning to achieve the alignment at different levels. Similar ideas have been also proposed in the image-text alignment by decomposing the images and texts into sub-tokens . Thus far, it has not been studied how these task-specific architectures can be integrated into large-scale pretraining. In this paper, we are the first to propose a simple yet effective token-aware contrastive loss for fine-grained alignment for pretraining and downstream tasks.
Negative sampling. Key to efficient contrastive training is a good source of negative examples. Most of current approaches use random sampling strategies for training video-text alignment . However, in the domain of image-text retrieval, a few works tried hard negative sampling to choose the hardest negatives for training. In , the authors computed the alignment scores for all image-text pairs in a mini-batch and use the hardest negative sample to compute the marginal loss. However, this strategy can only be applied without multi-modal fusion. In those models which have multi-modal fusion layers for better representations , the authors instead compute the matching score offline and then use it to sample hard negatives for finetuning image-text retrieval model, which however is difficult for large-scale pretraining due to the high computational cost. In this paper, our cascade hard negative mining is particularly designed to address these issues as we efficiently select the hard negative samples online before multi-modal fusion and send them to the fusion layers for computing the loss. As we will show in our experiments, this technique can be seamlessly applied to both large-scale pretraining and downstream tasks.
Method
As depicted in Fig. 1, our model has three components:
Language encoding module . We use pretrained tokenizer and BERT to tokenize the input texts and extract textual features, respectively. Given a raw sentence, we append a “[CLS]” and “[SEP]” to the beginning and end, respectively. At the top, we can obtain a sequence of textual features . We ensure the output feature dimension of video encoder to be identical to that of language encoder. During training, we update the parameters in our language encoder to adapt to the texts in specific domain, e.g., cooking instructions in YouCook2 .
The above three components comprise our video-text alignment model which is then trained with the proposed token-aware cascade contrastive loss. We start with a brief review of conventional contrastive learning and then introduce the proposed technique.
2 Contrastive learning: a revisit
Given a set of video-text pairs , our goal is to learn an optimal scoring function such that paired video and text have higher scores than all the other unmatched pairs . From the probabilistic perspective, aligning to is equivalent to maximizing the conditional probability while minimizing the probability for all negative pairs . According to , can be approximated by:
where is the alignment score between and ; the denominator is a sum over all possible videos, which is a partition function for normalization. Adding cross-entropy loss on , we can then derive the NCE loss :
The denominator in Eq. 2 requires a sum over all videos in a dataset, which is intractable in practice. Therefore, we usually compute the NCE loss on a mini-batch of video-text pairs sampled from the whole dataset. Ideally, we want to learn the parameters of the model to minimize the above NCE loss, such that is maximized over all tuples . A number of previous works used the above formula for contrastive learning . Meanwhile, there are some variants of computing contrastive loss in video-langauge representation learning. For example, omits the denominator and incorporate a margin s.t. in a mini-batch. optimizes binary cross-entropy (BCE) by assigning a positive label (1) and other pairs a negative label (0).
3 TACo: our approach
The way of using contrastive learning in previous works has two issues. First, the loss is computed at sentence-level by taking ‘[CLS]’ token or the maximum over all tokens in a sentence. Clearly, the content words (e.g., nouns, verbs) are more likely to align with the visual contents or concepts in the videos compared with function words (e.g., stop words). Second, the high computational cost in multi-modal fusion layers hinder the usage of large batch of negative samples, which however is essential to contrastive learning . Motivated by these two issues, we introduce TACo, a simple yet effective method to improve the contrastive learning. We elaborate below how these contrastive losses are computed.
Given the video-text pairs in a mini-batch, we first use our video encoder and language encoder to obtain a batch of video features and text features , respectively. Then, we average all tokens of a video clip to get , and take the first ‘[CLS]’ token for each text to get . Based on and , we compute the sentence-level contrastive loss:
where is a scalar temperature parameter. In Eq. 3, the computation is simply a number of dot-products between video and text features. Giving such efficiency, we can use all the negative samples in a mini-batch to compute the loss. Through this, we optimize and so as to project the video and text samples into an aligned feature space.
The ‘[CLS]’ token and average of video tokens in Eq. 3 overlooks the differences across tokens and frames, and thus may not provide the pressure to push individual tokens (e.g., nouns and verbs) to ground on the specific video contents. To encourage correct alignment, in addition to the sentence-level loss, we introduce a token-level contrastive loss:
where is another scalar temperature parameter; is the indices of tokens of interest in -th text, and is the -th token embedding in -th text. measures the similarity between video features and specific token embedding . It first computes the dot-product between and all video tokens , and then take the maximum over scores to get the final alignment score. Through Eq. 4, the model uses individual tokens as anchors to align with video, which is complementary to the sentence-level loss in Eq. 3. Similar to Eq. 3, we can compute this token-level contrastive loss efficiently, and thus use all the negative samples. As a whole, these two losses are used to optimize and in a token-aware manner.
Token of interest. In Eq. 4, we need to decide which tokens should be included in . In this paper, we heuristically select nouns and verbs as the targets considering they are more “concrete” in the videos. In practice, nouns or verbs usually have different discriminativeness even if they are all the same type. For example, “man” is a noun but is less informative than “gymnast”. To reflect this, we further assign different words with different weights by computing their inverse document frequency (idf) . A higher idf means it is more unique across the corpus, and hence will weigh more when computing the token-level contrastive loss. Another practical issue for computing the loss is that the tokens are usually sub-words due to the BERT tokenizer. Hence, for all tokens that belongs to the same word, we will assign the same weights accordingly.
After computing the token-aware contrastive loss, we feed the features from separate modalities to multi-modal fusion layers to enable more interactions between them two. Similar to previous work , we take the feature corresponding to the “[CLS]” in the outputs. We regard this as the summary of two modalities and then compute the contrastive loss:
where is the multi-modal fusion output for “[CLS]” token taking and as inputs; is the parameter in a linear layerfor clarity, we omit the bias term in the formula. Based on Eq. 5, we optimize all parameters in our model in collaboration with Eq. 3 and Eq. 4.
In Eq. 5, a practical challenge is that we can hardly use all negative samples in the mini-batch, due to the high computational and memory cost in the multi-modal fusion. The complexity of self-attention layer makes it intractable to pass all pairs into the multi-modal layers. Previous work solved this by performing random sampling to cut the number of negative samples to . However, randomly choosing negative samples may result in sub-optimal learning since the pairs are scarce. We therefore introduce a cascade sampling strategy to find hard negatives instead of random ones.
Cascade hard negative sampling. To reduce the computational cost in Eq. 5, we choose among all possible video-text pairs a small subset which are most difficult. However, computing the alignment scores for all pairs using Eq. 5 and then select the hard negatives is a “chicken-and-egg” problem. Instead, we propose to use the similarities between all video-text pairs computed in Eq. 3 and Eq. 4 as the guidance. Specifically, for each text-video pair , we take their global similarity computed in Eq. 3 and token-level similarity by aggregating for all tokens of interest in . Then we sum the two similarities as the alignment score for the given pair. For each text, we choose the top aligned negative videos and vice versa. The resulting pairs are then fed into the multi-modal fusion layers. Through this strategy, we can effectively select the difficult negative samples on the fly at no extra cost. Since the multi-modal fusion layers has more capacity (parameters) to distinguish these hard negatives from positive ones, our sampling strategy naturally prompts the cooperation between the three contrastive losses.
Finally, we present a comprehensive comparison to differentiate our model with previous works with respect to the used contrastive learning method in Table 1.
4 Objective
The training objective in our method is finding optimal by minimizing the combination of the above three contrastive losses:
where is the weight of token-level loss (0.5 by default). During inference, we make the prediction by summing the alignment scores from all the three scoring functions.
Experimental setup
In our experiments, we train and evaluate our model on the following established benchmarks:
• YouCook2 consists of 2k videos about routine cooking activities of 89 recipes. Each video contains multiple video clips annotated with text descriptions by human annotators. Following , we train our models on the training split, and report the text-video retrieval performance on around 3.5k validation clips.
• MSR-VTT contains 10k video clips associated with 200k sentences. There are two validation splits used in previous work. In , the training set has 9k clip-text pairs with the remaining 1k pairs for evaluation, which we denote by split1. In , 1k clip-text pairs are sampled from the 3k pairs in test set for evaluation, while the original 7k pairs are used for training. We denote this by split2. We report text-video retrieval results using both splits.
• ActivityNet . It consists of 20K YouTube videos, each of which is associated with multiple human-annotated captions. Following , we concatenate all the captions for a video into a paragraph and evaluate the paragraph-video retrieval on the “val1” split.
• Howto100M . We compare with previous work under the pretraining protocol on Howto100M . It was collected from YouTube and contains over 1.2M narrated videos associated with automatically generated transcripts. Each video contains over 100 clips on average.
To further verify the transferrability or our learned multi-modal representation from Howto100M, we also evaluate the action step localization and action segmentation on CrossTask and COIN , respectively.
2 Settings
Previous work use a variety of different video and language representations which we find significantly affect the final performance. We summarize different choices below:
• Video representations. For 2D CNN, Resnet-152 is used to extract feature map and then globally pooled to 2048-d . For 3D features, commonly used models are I3D , R(2+1)D and S3D . In , the authors further extract objects from the video clips. In , the authors use collaborative experts to extract features from audio, scene, OCR, face, speech, etc.
• Language representations. There are primarily four variants: 1) GoogleNews pretrained word2vec (w2v) used in ; 2) LSTM or Bidirectional LSTM ; 3) pretrained BERT used in and 4) OpenAI-GPT used in .
In this paper, we use a pretrained BERT-base model for language representation as in . For video features, following , we extract 2D CNN features using Resnet-152 (R-152) pretrained on ImageNet . For 3D CNN features, we use I3D (with Resnext-101 backbone) pretrained on Kinetics-400 and S3D pretrained on Howto100M . The off-the-shelf pretrained weights are provided by and . For simplicity, we denote them by I3D-X101 and S3D-HM in the following.
Another discrepancy among different methods is the number of self-attention layers used in the model. In , the authors use 12 multi-modal self-attention layers while 6 video encoder layers and 2 multi-modal fusion layers are used in . Differently, 4 multi-modal self-attention layers are used in . In this paper, for all our ablation studies below, we use 1 and 2 self-attention layers for our video encoder and multi-modal fusion, respectively. To compare with previous work on specific dataset, we use 2 video encoding layers. While pretraining the model with large-scale dataset Howto100M , we increase to 4 video encoding layers for comparable model capacity to previous works . Note that this largest model is still smaller than or on par with the aforementioned methods.
Video representations. We train our model with different video representations as described above and compare it with the baseline model which has identical architecture but merely trained with as depicted in Eq. 5. The baseline contrastive learning method has been adopted in a number of previous works . This comparison can verify the effectiveness of our proposed contrastive learning method considering two models have exactly the same number of parameters. In Table 4.2, we can see our proposed method outperforms baseline across all feature types introduced in Sec. 4.2 on both YouCook2 and MSR-VTT. Note that our model uses exactly the same number of parameters to the baseline model. These consistent improvements demonstrate the effectiveness and generalization ability of our proposed method. As mentioned above, we also observe the text-video retrieval performance significantly depends on the feature types. We can find 3D features (I3D-X101 and S3D-HM) in general outperform 2D feature (R-152), which is expected since 2D feature does not capture the motions in the videos. Among all three feature types, S3D-HM outperforms the other two with large margin, which demonstrates the potential to learn good video representation by pretraining on large-scale noisy dataset (Howto100M ). Because Howto100M mainly contains instructional videos, it is more close to YouCook2 than MSR-VTT, and hence we see more gain on YouCook2. These comparisons indicate video representations matter much to the final performance.
Results on separate datasets. We separately show the comparisons on YouCook2, MSR-VTT and ActivityNet in Table 5.1.1, 5.1.1 and 5.1.1. For a fair comparison with previous works, we use the same or similar features as listed in the tables. As we can see, our method outperforms all previous work across all datasets. These results validates its effectiveness to learn video-text alignment. Note that previous works either use a variety of loss functions or a collection of multiple features . In contrast, we achieve the best performance using a simpler contrastive learning pipeline with smaller model size. This supports our earlier claim on the efficiency. Comparing the numbers in Table 4.2, Table 5.1.1 and Table 5.1.1, we can find our model achieves better performance with the same video features when using deeper video encoder (2 layers v.s. 1 layer).
Zero-shot and finetuned performance. In Table 5.1.2, we show the comparisons across different models pretrained on Howto100M. In the upper part of the table, we compare the zero-shot performance on YouCook2 and MSR-VTT. We do not evaluate on ActivityNet since it has different number of input video tokens compared with the pretrained model and thus is not directly compatible to the pretrained model. As we can see, TACo outperforms previous works significantly on YouCook2 and slightly on MSR-VTT. Since YouCook2 has closer domain gap to Howto100M than MSR-VTT, the improvement brought by large-scale pretraining is more significant. However, on MSR-VTT, our model still outperforms MIL-NCE which uses the same video features. In Fig. 2, we show the zero-shot performance on YouCook2 and MSR-VTT when pretraining our models with different contrastive losses as listed in Table 3. Accordingly, it shows our proposed contrastive losses gradually improve the performance, and combining all techniques achieves the best performance. Based on the pretrained model, we further finetune it on specific datasets. In our experiments, we use two feature S3D-HM and R-152+S3D-HM, to compare with the methods with the same/similar settings. As we can see, our model using S3D-HM outperforms UniVL using the same feature but more video encoder layers (6). Different from zero-shot results, we observe more improvement on MSR-VTT than YouCook2 after finetuning. This implies that finetuning on specific datasets can compensate the domain gap to the pretraining datasets. To compare with the methods using features extracted from collaborative experts , we enrich our video representation by adding 2D R-152 feature, which achieves better performance on MSR-VTT, and better Recall@1 and Median Rank on ActivityNet. Note that this combination hurts the performance on YouCook2, and we witnessed a similar trend for models without pretraining in Table 4.2. Finally, comparing with the results without pretraining in Table 5.1.1, 5.1.1 and 5.1.1, we clearly find large-scale pretraining and finetuning brings substantial improvements consistently.
Following , we evaluate action step localization performance on CrossTask dataset . It covers 18 tasks and each video contains multiple video segments annotated with action steps and natural language descriptions. Similar to , we use our model to compute the similarity between each frame and the action step descriptions, which results in a score matrix. Using the official algorithm provided by , we can find the optimal frame-wise order of action steps for a video. By comparing it with the ground-truth annotations, we compute the recall for each task and then do the average. According to the results in Table 5.1.2, our model achieves the best performance compared with previous works. This indicates that our model can learn good video-language representations.
We further evaluate our pretrained model on action segmentation task on COIN dataset, following . Unlike the above task, action segmentation does not rely on texts, and thus can be used to evaluate the learned video representation. As shown in Table 5.1.2, our method significantly outperforms MIL-NCE and ActBert, and achieves comparable performance to UniVL. This indicates that our model is also a good video representation learner.
Conclusion
In this paper, we introduced TACo, a simple yet effective contrastive learning method for learning video-text alignment. It is aimed at addressing two existing issues in current contrastive learning pipelines: missing fine-grained alignment and inefficient sampling for multi-modal fusion. Without introducing any extra parameters, our method achieved promising results on three text-video retrieval benchmarks under various evaluation protocols. We further demonstrated the learned representations can be effectively transferred to other tasks such as action step localization and segmentation. Based on all these encouraging results, we believe TACo is a good alternative to conventional contrastive learning pipeline.
References
Appendix A Tokens of interest
We extract tokens of interest (T.O.I) using the pos-tagger provided by Spacy . In Table 10, we show the statistics of tokens for three datasets. For each token that is tagged at VERB or NOUN, we compute the inverse document frequency (idf) by:
where is the full set of corpus, which are the captions in the training set for a dataset; the denominator counts the number of captions which contain a specific token. Based on Eq. 7, we can compute the idf for each token of interest. The smaller the idf, the more frequent it appears in the corpus. We do not compute the tf term since usually a token only appears once in a single sentence. The full list of tokens and corresponding idfs can be found in Fig. 4. For a given sentence, we first assign the computed idfs to its nouns and verbs and then normalize the idfs, which are then used to weigh the token-level contrastive losses.
In this part, we investigate the contributions of three contrastive losses used in our model. After we train the video-text alignment model using all three losses, we report the performance using separate alignment scores in Table 11. For reference, the top two rows are the performance for using early stage only and later stage only contrastive learning to train the model. The bottom four rows are the separate performance at different stages for our model. As we can see, combining three contrastive losses during training can boost the performance for both early and later stage (row 3 v.s. row 1, row 5 v.s. row 2). This indicates that the three losses are synergistic to each other for a better video-text alignment. On the other hand, the early stage alignment achieves better performance than other two (token-level and later stage), while the fused score is the best. We suspect that this is because early stage alignment is trained with all text-video pairs at sentence-level. In contrast, token-level contrast focuses on single tokens and the multi-modal fusion layers merely see a small part of hard text-video pairs.
The proposed cascade sampling helps the later stage contrastive learning to focus on hard negative samples. As shown in our main submission, adding cascade sampling will improve the performance. We suspect this is because cascade sampling helps learn a better later stage alignment. To verify this, we compare the later stage alignment across three different settings: 1) merely applying later stage contrastive loss; 2) combine early state and later stage contrastive losses and 3) using cascade sampling for later stage contrastive loss. We report the results on YouCook2 in Table 12. Here, note that we only use the later stage alignment scores for evaluating the performance. As we can see, combining early stage and later stage together slightly improves the performance. This is probably because early stage contrastive loss helps to learn a better video and language encoder, from which the multi-modal module takes better representations for cross-modal fusion. After applying the cascade sampling for the later stage contrastive loss, the performance is further improved. Since our cascade sampling strategy can send more difficult samples to the later stage, the cross-modal fusion layers can learn more discriminative representations for video-text alignment. These results validate that the hard negative mining through cascade sampling indeed helps to improve the later-stage text-video alignment, and hence the final performance.
In our main paper, we noticed the number of video encoder layers affects the final performance. To have a more comprehensive study, we use R-152 and S3D-HM as the 2D and 3D features and train the video-text alignment model on YouCook2 with different video encoder layers. As shown in Table 13, using more video encoder layers can significantly boost the text-video retrieval performance. Particularly, when no video encoder layers are used, the model can hardy capture the long-range temporal dynamics, and thus performs poorly. Once we add one video encoder layer, the performance improves significantly. With the increase of encoder layers, the performance is further improved, which is reasonable since more video encoder layers can encode more complicated video contents and dynamics.
Finally, we attempt to compare the model sizes and computational costs for different methods. Unfortunately, all previous methods did not report FLOPs and only MMT discussed #params. However, the results in Table 13 imply that bigger model can usually achieve better performance. Therefore, it is necessary to have a comparison of model size and computational cost between our model and those from other methods. For other methods which do not report the numbers, we estimate them based on the descriptions in the original paper. Table 14 summarizes the comparisons and also reports the #params and FLOPs (all underlined numbers are estimated based on the descriptions in original papers). As shown, our largest model has comparable size and FLOPs to others.
We visualize the text-video retrieval results by varying the weights for the token-level alignment scores during testing. In Fig. 3, we show two text-video retrieval examples on YouCook (top) and MSR-VTT (bottom). From top to bottom, the five rows in each block correspond to the top five retrieved results from the whole test set. As we can see, when we gradually increase the weight for the token-level alignment score, there are more related videos appearing in the top five candidates. For YouCook2, when we set the weight equal to 0.0, the third and fifth video are not well-aligned with the query since they are both not about “tomato”. When we increase the weight to 0.1, we can observe the the fourth video moves to the third place. After we increase the weight to 0.5, we can see all top-5 videos are about cutting tomato. Similarly, for MSR-VTT, we can see the last three videos are not about “two people talking on a table”. When we increase the weight to 0.1, the fifth video is replaced with a more matched video. Keeping increase the weight to 0.5, we can obtain the top 5 videos all about “two people talking with each other on a table”. These visualizations demonstrate the efficacy of our proposed token-level contrastive learning.