LocVTP: Video-Text Pre-training for Temporal Localization

Meng Cao, Tianyu Yang, Junwu Weng, Can Zhang, Jue Wang, Yuexian Zou

Introduction

Video-Text Pre-training (VTP) has attracted increasing attention with the aim to learn generic and transferable joint video-language (VL) representations. Compared to the conventional separate pre-training on each single modality, e.g., video features are pre-trained under the action recognition datasets (Kinetics , Sport1M ), VTP has several advantages: 1) It leverages large-scale unlabeled narrated video data with automatically generated corresponding text data for video-text correspondence pre-training. 2) It tries to map different modality features into a shared latent space, which reduces the difficulties of the cross-modal feature interaction. Thanks to these advantages, VTP has significantly improved the performance of many downstream VL tasks. For example, as illustrated in , the video retrieval performance using features pre-trained with the VTP method MIL-NCE is much higher than that using separately pre-trained way (cf. Fig. 1(a) (left)).

Despite their encouraging performance, we find that most current VTP methods are applicable to limited downstream tasks, i.e., they focus on retrieval-based tasks which require video-level predictions, e.g., video retrieval , video captioning , and video question answering . In contrast, there exists another mainstream localization-based tasks which expect more fine-grained clip-level or frame-level predictions, e.g., temporal grounding , action segmentation , action step localization (cf. Fig. 1(b)). Unfortunately, through experiments, we find their poor generalization abilities on this type of downstream tasks. For example, on temporal grounding, even pre-trained with a much larger dataset HowTo100M , the VTP method MIL-NCE still performs worse than the separately pre-trained counterpart (cf. Fig 1(a) (right)).

In this paper, we analyze that this poor transfer ability on localization-based tasks is due to the absence of two indispensable characteristics: 1) Fine-grained alignment: We contend that the alignment should be conducted on more fine-grained clip-word level instead of the coarse-grained video-sentenceHere we use “sentence” to represent the whole paired text for each video, such as the ASR in HowTo100M or query language in ActivityNet Caption . level. As the temporal grounding example shown in Fig. 2, a given query sentence may contain multiple actions (e.g., “hit the golf ball” (qs1q^{s1}) and “bend down to pick up the ball” (qs2q^{s2})). Thus, aligning each action (or words) to the corresponding clips (i.e., vt1v^{t1} and vt2v^{t2}) will help to obtain more detailed and accurate feature representations. 2) Temporal relation reasoning: We hope the clip features of a certain action can also perceive other actions in the same video. For example, for a typical golf video, action qs2q^{s2} (“bend down to pick up the ball”) always occurs shortly after action qs1q^{s1} (“hit the golf ball”). Thus, incorporating such temporal relationship into VTP can help to improve the temporal awareness of video features.

Based on these observations, we propose a novel video-text pre-training framework for localization tasks, dubbed as LocVTP. By considering both above-mentioned characteristics, LocVTP achieves state-of-the-art performance not only on the widely studied retrieval-based tasks, but also on the less-focused localization-based tasks. Specifically, for fine-grained alignment, we extend the coarse-grained contrastive training with video-sentence alignment to a fine-grained one with clip-word alignment. Since there are no clip-word correspondence annotations in existing large-scale datasets, we utilize the latent space established by the coarse-grained contrastive learning to estimate the clip-word similarity, and then select the clip-word pairs with high similarities as positive samples. To further illustrate this, as shown in Fig. 2 (right), suppose {vt1,qs1}\{\bm{v}^{t1},\bm{q}^{s1}\} and {vt2,qs2}\{\bm{v}^{t2},\bm{q}^{s2}\} are two matched clip-word feature pairs. Semantic embeddings in each pair are mapped to be close to each other, i.e., vt1↔qs1\bm{v}^{t1}\leftrightarrow\bm{q}^{s1}, vt2↔qs2\bm{v}^{t2}\leftrightarrow\bm{q}^{s2}. For temporal relation reasoning, we propose a new pretext task called context warping. Here we use Fig. 2 (right) for illustration. Context warping is designed to generate a new temporally relevant clip features zt1\bm{z}^{t1}, which imitates vt1\bm{v}^{t1}, conditioned on another clip vt2\bm{v}^{t2} and the relative distance t2−t1t2-t1 in time, i.e., zt1=warp⁡(vt2,t2−t1)\bm{z}^{t1}=\operatorname{warp}(\bm{v}^{t2},t2-t1). The predicted relevant clip feature zt1\bm{z}^{t1} is enforced to maintain the original established cross-modal correspondence unchanged, i.e., zt1↔qs1\bm{z}^{t1}\leftrightarrow\bm{q}^{s1}. In this manner, we simulate the contextual reasoning process and enhance the temporal awareness of video features.

We conduct extensive experiments on four downstream tasks (i.e., video retrieval, temporal grounding, action step localization, and action segmentation) across six datasets. The results on both retrieval-based and localization-based tasks demonstrate the superiority and the generalization ability of our LocVTP.

In summary, we make three contributions in this paper:

We propose a localization-oriented video-text pre-training framework, LocVTP, which benefits both retrieval-based and the less-explored localization-based downstream tasks.

We pinpoint two crucial designs in LocVTP, i.e., fine-grained video-text alignment and temporal relation reasoning.

Experimental results show that our LocVTP significantly outperforms previous state-of-the-art methods when transferred to various downstream tasks.

Related Work

Video-Text Pre-training (VTP). With the release of the large-scale instructional dataset HowTo100M, VTP has spurred significant interest in the community. Overall, the mainstream methods can be broadly classified into two classes: 1) Generative methods: Several methods try to extend BERT to the cross-modal domain, i.e., they accept both visual and textual tokens as input and perform the masked-token prediction task. 2) Discriminative methods. These methods learn representations by differentiating input samples using objectives such as the metric loss or contrastive loss . ClipBert enables affordable pre-training from sparsely sampled frames. Frozen adapts the recent ViT as the visual encoder and is flexible to be trained on both image and video datasets. T2VLAD and FCA also perform the fine-grained interactions between video clips and phrases. However, both of them resort to additional overload, e.g., k-means cluster or graph auto-encoder. In contrast, our LocVTP explicitly models the clip-word matching with a more light-weighted similarity comparison manner.

Pre-training for localization tasks. Compared to the retrieval tasks which only require only video-level predictions, localization tasks are essentially different since they need dense clip-level or frame-level predictions and thus the pre-training for these tasks is more challenging. In the pure video domain, this gap has been noticed and several pre-training works tailored for action localization have been proposed. BSP synthesizes temporal boundaries using existing action recognition datasets and conducts boundary type classification to generate localization-friendly features. TSP trains video encoders to be temporally sensitive by predicting the foreground clip label and classifying whether a clip is inside or outside the action. As for the video-language domain, our LocVTP is the first pre-training framework designed for localization tasks. Besides, compared to TSP and BSP which require label information for supervised pre-training, our LocVTP can directly learn from narrated videos.

Approach

Three types of contrastive methods are then performed to learn cross-modal features: 1) The coarse-grained contrastive loss builds the video-sentence level alignment; 2) A correspondence discovery strategy is proposed to build clip-word relations, based on which the fine-grained contrastive loss is applied; 3) Temporal aware contrastive loss with the context warping pretext task is proposed to encode temporal information into video representations.

2 Coarse-grained Contrastive Learning

where q‾i,i∈[1,N]\overline{\bm{q}}_{i},i\in[1,N], is the sentence feature for other samples within the batch. NN denotes the batch size and τ\tau is the temperature parameter. The coarse-grained contrastive loss Lc\mathcal{L}_{c} serves as a base loss to conduct video-sentence level constraint and induces a basic latent space where the detailed cross-modal matching is achieved. Though usually coarse and noisy, this latent space encodes prior for fine-grained clip-word correspondence discovery. In Section 4.6, we design and analyze three potential ways to use this cross-modal matching prior.

3 Fine-grained Contrastive Learning

Beyond the coarse-grained video-sentence alignment, we propose to conduct contrastive learning in a fine-grained manner, i.e., clip-word matching. We contend that introducing such alignment learning into the pre-training stage could narrow down its gap with downstream localization tasks and calibrate the pre-trained feature to be more temporally aware.

Clip-word correspondence discovery. Before performing fine-grained contrastive learning, we firstly need to estimate the clip-word correspondences from video-sentence pairs. Thanks to the priors well established by the coarse-grained contrastive learning, we compute the cosine similarities between the video clips and their corresponding caption words in the pre-built latent space and choose the most similar KK words as the correspondence for each video clip. Note that we select multiple positive words rather than simply pick one with the highest similarity because individual words may have vague meanings while sense-groupA group or sequence of words conveying a particular meaning or idea in linguistics. conveys more precise information (cf. Section 4.7).

Given the video sentence pair {v,q}\left\{\bm{v},\bm{q}\right\}, for the encoded ttht^{th} video clip vt\bm{v}^{t}, we compute its cosine similarities with the sths^{th} word embedding qs\bm{q}^{s} and apply the topk⁡\operatorname{topk} operation to select the most matched KK ones. Following , these KK selected items are average pooled to form the final positive sample:

Fine-grained contrastive loss. With the selected clip-word correspondence as positive pairs, we perform fine-grained representation learning following the cross-modal InfoNCE loss (cf. Figure 4(a)). The negative samples are taken from the other words within the batch. Therefore, the fine-grained contrastive loss is defined as follows.

where qis\bm{q}_{i}^{s} is the sths^{th} word feature of the ithi^{th} sentence qi\bm{q}_{i}.

4 Temporal aware Contrastive Learning

𝑡\bm{q}_{+}^{t} is the pooled positive word features. zt\bm{z}^{t} is the warped feature. We only present positive samples and omit negative ones. Compared with the video-level retrieval task, which favors temporal invariant features , the clip-level localization task prefers temporal aware video embeddings. Specifically, correlated actions in the same video should perceive each other. This characteristic is however not embodied in the aforementioned contrastive learning.

Context warping head. To alleviate this, we set up a context-warping operation to enforce the video clip to perceive the context. For the video clip vt\bm{v}^{t} in a matched clip-word pair {vt,q+t}\{\bm{v}^{t},\bm{q}^{t}_{+}\} (cf. Section 3.3), we warp its contextual video clip with δ\delta temporal distance, i.e., vt+δ\bm{v}^{t+\delta}, to “reconstruct” itself. To supervise this warping process, we set up a temporal aware contrastive loss to maintain the established correspondence. Specifically, we propose a context warping head g(⋅)g(\cdot) to instantiate this warping process, by taking the context clip feature vt+δ\bm{v}^{t+\delta} and temporal distance δ\delta as input.

Temporal aware contrastive loss. Through the context warping head, the warped feature zt\bm{z}^{t} should mimic the reference feature vt\bm{v}^{t}. Since vt\bm{v}^{t} has the clip-word alignment with q+t\bm{q}^{t}_{+}, such correspondence should be preserved between the warped feature zt\bm{z}^{t} and q+t\bm{q}^{t}_{+} (cf. Fig. 4(b)).

This process enforces video features to learn the ability of temporally reasoning, thus leading to more localization-friendly video features.

Integrating the above constraints, our final loss function is as follows.

where λc\lambda_{c}, λf\lambda_{f}, and λt\lambda_{t} balance the focus on different constraints during training.

Experiments

Datasets. We pre-trained our model on three public datasets: 1) HowTo100M . It consists of more than 1.2M videos accompanied with ASR-generated speech transcription. The provided transcription is used to create video-sentence pairs separated by each timestamp. 2) WebVid-2M . It contains about 2.5M well-aligned web video-text pairs. 3) Google Conceptual Captions . It contains 3.3M image and description pairs harvested from the web.

Encoders. Following , we adopted ViT-B/16 with space-time attention as the video encoder. The spatial attention weights in the transformer were initialized with ImageNet-21k pre-trained weights while the temporal attention weights were set to zero. We chose a lightweight DistilBERT as the language encoder. Following , the language encoder was initialized with the weights pre-trained on English Wikipedia and Toronto Book Corpus.

Implementation Details. For the video in each video-sentence pair, we sampled 8 clips of 16 frames equidistantly and fed them to the video encoder to obtain clip-level features. All frames were resized to 224×224224\times 224. For downstream transfer, we extracted video features with the well-trained model in a dense manner, i.e., every 16 consecutive frames were grouped to compute one clip feature.

2 Transfer Results on Video Retrieval

Datasets. We evaluate our LocVTP on the widely-used benchmark MSR-VTT dataset . It is composed of 10K YouTube videos (9K for training and 1K for test). We report results on the train/test splits introduced in .

Results. 1) As can be seen, we achieve state-of-the-art performance under both sets of data, i.e., HowTo100M and CC3M+WV2M. Specifically, when pre-trained on CC3M+WV2M, LocVTP outperforms Frozen by an absolute lift of 4.8% on R@5. 2) It should be pointed out that although using RGB data only, our LocVTP achieves better performance than the methods using multi-modal expert features including motion, face, and speech, e.g., MMT . 3) The recent work CLIP provides a stronger vision encoder and we also evaluate the performance based on it. It is shown that the CLIP’s weights greatly improve the performance of LocVTP with R@5 achieving 72.8%, surpassing top-performing CLIP-based methods. 4) Our LocVTP also outperforms previous methods under the zero-shot setting, showing its generalization ability.

3 Transfer Results on Temporal Grounding

Settings. We validate the performance of pre-trained representations on temporal grounding, which aims to localize actions corresponding to the sentence from an untrimmed video. Specifically, we re-train the mainstream temporal grounding method 2D-TAN We choose 2D-TAN since it is relatively simple without too many dataset-specific parameters, which can fairly verify the effectiveness of pre-training features. Results on more advanced baselines are available in the supplementary material. by only replacing the original input features with pre-trained ones. For ease of feature extraction, we choose representative VTP methods with publicly-available codes for comparisons.

Datasets and Metrics. 1) ActivityNet Captions (ANet) . It contains 20K untrimmed videos with 100K descriptions. By convention, we use 37,417 video-query pairs for training, 17,505 pairs for validation, and 17,031 pairs for testing. 2) Charades-STA . Following the official split, 12,408 video-query pairs are used for training, and 3,720 pairs for testing. 3) TACoS . It has 10,146 video-query pairs for training, 4,589 pairs for validation, and 4,083 pairs for testing.

Following prior works, we adopt “R@n, IoU@m” (abbreviated as RnmR^{m}_{n}) as the metric, Specifically, RnmR^{m}_{n} is defined as the percentage of at least one of top-n retrieved moments having IoU with the ground-truth moment larger than mm.

Results. 1) As shown in Table 2, even trained with a much larger dataset, the current popular video-text pre-training frameworks achieve inferior performance compared to the separately pre-trained one. For example, Frozen reaches 43.3% at R10.5R_{1}^{0.5} on ANet Captions, which is 1.1% absolute value lower than the separately pre-trained counterpart. 2) Either pre-trained on HowTo100M or CC + WV, our LocVTP outperforms both video-text pre-training methods by a large margin on all three datasets. For example, pre-trained on HowTo100M, LocVTP surpasses the separately pre-trained method by 3.8% on R10.5R_{1}^{0.5} of ANet Captions. 3) For more fair comparisons, we sample a subset of HowTo100M by selecting the same training sample as Kinetics (300K training pairs), denoted as HT‡ in Table 2. Although using noisy ASR captions, the results demonstrates that under the same training data volume, our LocVTP still shows better performance compared to the separately pre-trained method. This manifests that our performance improvement is brought by the sound architecture design rather than just the use of the large-scale dataset.

4 Transfer Results on Action Step Localization

Settings. In action step localization, each video belongs to a task and is annotated with multiple action steps described with short natural languages. The goal is to align each frame with the correct step in the text form. Following , we take as the downstream localization method. Specifically, we compute the similarity between each frame and the action step descriptions in feature space to find the optimal frame-wise order of action steps for a video.

Datasets and Metrics. We experiment on the instructional video dataset CrossTask , which includes 83 tasks and 4.7K videos. Each task is described with an ordered list of steps with manual natural language descriptions. We perform the same evaluation protocol as in by reporting the average recall (CTR).

Results. Table 3 reports the action step localization performance on CrossTask dataset. Our LocVTP pre-trained feature achieves state-of-the-art performance with CTR reaching 51.7%, surpassing the previous method VideoClip by 4.4%. Our competitive performance demonstrates that LocVTP features can effectively perceive detailed action steps.

5 Transfer Results on Action Segmentation

Settings. We assess our LocVTP on action segmentation, which aims to predict the action label frame-wisely for each video frame. It is a pure vision task without the use of the text encoder. Following , we encode the input video frames with the well-trained video encoder and apply a linear classifier upon the features to predict action labels.

Datasets and Metrics. We conduct experiments on the widely used COIN dataset and the frame-wise accuracy (FA) is taken as the evaluation metric.

Results. As shown in Table 3, our LocVTP achieves state-of-the-art performance with FA reaching 72.9%. This further demonstrates the superiority of our feature in localization tasks even in the absence of language guidance.

Training Strategy. Coarse-grained contrastive alignment loss Lc\mathcal{L}_{c} provides a basic cross-modal matching prior and we introduce three potential ways to use it: 1) multi-stage training: first perform coarse-grained training and then use the trained model to initialize other stages. 2) warm-up training: decrease λc\lambda_{c} exponentially from 1 to 0 throughout the training process. 3) weighted training: set λc\lambda_{c} to a constant value. Here we set λc=0.5\lambda_{c}=0.5. As shown in Table 4(a), we find the weighted training strategy achieves the best performance and warm-up training is slightly behind. Multi-stage training is the least effective one.

Loss Component. We present the loss component ablations in Table 4(b). As shown, both fine-grained loss Lf\mathcal{L}_{f} and temporal aware loss Lt\mathcal{L}_{t} are crucial. For example, compared to the full version (exp.#1), removing Lf\mathcal{L}_{f} and Lt\mathcal{L}_{t} brings about 1.4% and 1.5% performance degradation on the R10.5R_{1}^{0.5} metric, respectively.

More downstream temporal grounding baselines. We take another temporal grounding method CSMGAN as the downstream baseline. As shown in Table. 4(c), our LocVTP pre-trained feature consistently benefits this more advanced baseline.

7 Ablations on Fine-grained Contrastive Loss

Correspondence Discovery Strategies. We experiment four potential strategies to extract cross-modal correspondences: 1) random: randomly select KK words for each clip; 2) 2d-topk: select the most similar K×TK\times T clip-word pairs; 3) word-topk: select the most similar KK clips for each word; 4) clip-topk: select the most similar KK words for each clip, namely the method illustrated in Section 3.3. As indicated in Table. 5(a), the random and 2d-topk matching strategies are the two worst options. For the word-topk matching, it is also sub-optimal, which can be attributed to the possibility of introducing words without concrete meanings (e.g., articles or pronouns) into matched pairs.

Number of Selected Pairs KK. We further ablate the hyper-parameter KK used in the clip-topk strategy. Table 5(b) shows that the performance saturates at K=3K=3 and slightly decreases for K=4K=4. We conjecture that this may be because too few words have vague meanings while too large KK value leads to the inability to establish accurate correspondences.

8 Ablations on Temporal aware Contrastive Loss

Context Projection Head Components. In Eq. (4), the warped feature is generated based on both the direction sgn⁡(δ)\operatorname{sgn}(\delta) and distance ∣δ∣\lvert\delta\rvert. Here we investigate eliminating either of them to see the difference. We observe in Table. 5(c) that removing either component decreases the performance, which indicates that both the direction and distance of bias δ\delta are crucial for feature warping.

Maximum Bias Distance δmax\delta_{max}. Here we ablate different values for δmax\delta_{max}. From Table 5(d), we can see that δmax=4\delta_{max}=4 achieves the best performance. This may be because that small bias makes the model unable to perceive enough context, while a large bias makes contextual reasoning too difficult.

Intra-modal v.s. Cross-modal Constraint. In Section. 3.4, given the matched clip-word pair {vt,q+t}\{\bm{v}^{t},\bm{q}^{t}_{+}\} and the warped feature zt\bm{z}^{t}, we force the cross-modal supervision, i.e., zt↔q+t\bm{z}^{t}\leftrightarrow\bm{q}^{t}_{+}. Here, we apply the temporal aware contrastive loss Lt\mathcal{L}_{t} in a intra-modal manner which regards zt\bm{z}^{t} and vt\bm{v}^{t} as positive pairs, i.e., zt↔vt\bm{z}^{t}\leftrightarrow\bm{v}^{t}. The results in Table 5(e) show that our adopted cross-modal mode outperforms the intra-modal one.

Temporal Sensitivity Analysis. As a sanity check, we devise two proxy tasks to evaluate the temporal sensitivity of pre-trained video features. As shown in Fig. 5(a), nn equidistantly sampled clips from one video are fed into the frozen video backbone to extract their corresponding features. Two linear classifiers are trained to perform two tasks: order prediction and distance estimation. The first task predicts the temporal index while the second one estimates the temporal distance of two clips. The results in Table 5(f) show that our LocVTP with temporal aware loss Lt\mathcal{L}_{t} outperforms the variant without it as well as two typical VTP methods (i.e., UniVL and MIL-NCE), which shows that Lt\mathcal{L}_{t} clearly contributes to the localization ability.

9 VisualizationMore visualizations are left in the supplementary materials.

Cross-modal Correspondence Visualizations. Fig. 5(b) shows two framesHere we use “frame” to indicate the center frame of a video snippet. and their corresponding similarity scores with caption words. The top K highest scored words are marked with red (K=3K=3). Frame #1 and frame #2 have similar appearance views yet correspond to different action processes. Our method pinpoints the subtle differences and accurately finds the most relevant words.

UMAP Visualizations. As shown in Fig. 6, we provide UMAP visualizations for fused multi-modal features, which are generated by multiplying the extracted video feature by one query feature. With the temporal aware loss Lt\mathcal{L}_{t}, our LocVTP shows more separable distributions compared with LocVTP w/o Lt\mathcal{L}_{t}, manifesting that Lt\mathcal{L}_{t} helps distinguish action-of-interest from background.

Similarity Distribution Visualizations. In Eq.(4), context projection head warps contextual clip vt+δ\bm{v}^{t+\delta} to the reference one vt\bm{v}^{t}. Here we collect 10K paired training samples and compute three sets of cosine similarities: reference similarity (vt,q+t)\small{(\bm{v}^{t},\bm{q}_{+}^{t})}, bias similarity (vt+δ,q+t)\small{(\bm{v}^{t+\delta},\bm{q}_{+}^{t})}, and projection similarity (zt,q+t)\small{(\bm{z}^{t},\bm{q}_{+}^{t})}. Fig. 5(c) plots the histogram of these similarities. We can see that the distribution of projection similarity is close to that of reference similarity while far away from that of bias similarity. This demonstrates that our context projection head can effectively warp contextual features conditioned on the temporal information.

Conclusions

In this paper, we propose LocVTP, the first video-text pre-training framework for temporal localization tasks. Specifically, we apply cross-modal contrastive learning at both coarse-grained video-sentence and fine-grained clip-word levels. Besides, we propose a context warping pretext task and a temporal aware contrastive loss to enhance the temporal awareness of video features. Experimental results show that LocVTP achieves state-of-the-art performance when transferred to both retrieval-based and localization-based downstream tasks.

Acknowledgements. This paper was partially supported by NSFC (No: 621760 08) and Shenzhen Science & Technology Research Program (No: GXWD20201231 165807007-20200814115301001).

References