Query-Dependent Video Representation for Moment Retrieval and Highlight Detection

WonJun Moon, Sangeek Hyun, SangUk Park, Dongchan Park, Jae-Pil Heo

Introduction

Along with the advance of digital devices and platforms, video is now one of the most desired data types for consumers . Although the large information capacity of videos might be beneficial in many aspects, e.g., informative and entertaining, inspecting the videos is time-consuming, so that it is hard to capture the desired moments .

Indeed, the need to retrieve user-requested or highlight moments within videos is greatly raised. Numerous research efforts were put into the search for the requested moments in the video and summarizing the video highlights . Recently, Moment-DETR further spotlighted the topic by proposing a QVHighlights dataset that enables the model to perform both tasks, retrieving the moments with their highlight-ness, simultaneously.

When describing the moment, one of the most favored types of query is the natural language sentence (text). While early methods utilized convolution networks , recent approaches have shown that deploying the attention mechanism of transformer architecture is more effective to fuse the text query into the video representation. For example, Moment-DETR introduced the transformer architecture which processes both text and video tokens as input by modifying the detection transformer (DETR), and UMT proposed transformer architectures to take multi-modal sources, e.g., video and audio. Also, they utilized the text queries in the transformer decoder. Although they brought breakthroughs in the field of MR/HD with seminal architectures, they overlooked the role of the text query. To validate our claim, we investigate the Moment-DETR in terms of the impact of text query in MR/HD (Fig.1). Given the video clips with a relevant positive query and an irrelevant negative query, we observe that the baseline often neglects the given text query when estimating the query-relevance scores, i.e., saliency scores, for each video clip.

To this end, we propose Query-Dependent DETR (QD-DETR) that produces query-dependent video representation. Our key focus is to ensure that the model’s prediction for each clip is highly dependent on the query. First, to fully utilize the contextual information in the query, we revise the transformer encoder to be equipped with cross-attention layers at the very first layers. By inserting a video as the query and a text as the key and value of the cross-attention layers, our encoder enforces the engagement of the text query in extracting video representation. Then, in order to not only inject a lot of textual information into the video feature but also make it fully exploited, we leverage the negative video-query pairs generated by mixing the original pairs. Specifically, the model is learned to suppress the saliency scores of such negative (irrelevant) pairs. Our expectation is the increased contribution of the text query in prediction since the videos will be sometimes required to yield high saliency scores and sometimes low ones depending on whether the text query is relevant or not. Lastly, to apply the dynamic criterion to mark highlights for each instance, we deploy a saliency token to represent the entire video and utilize it as an input-adaptive saliency criterion. With all components combined, our QD-DETR produces query-dependent video representation by integrating source and query modalities. This further allows the use of positional queries in the transformer decoder. Overall, our superior performances over the existing approaches validate the significance of the role of text query for MR/HD.

Related Work

MR is the task of localizing the moment relevant to the given text description. Popular approaches are modeling the cross-modal interaction between text query-video pair or understanding the context of the temporal relation among video clips . On the other hand, TVT exploited the additional data, i.e., subtitle, to capture the moment, and FVMR enhanced the model in terms of inference speed for efficient MR.

Different from the MR, HD aims to measure the clip-wise importance level of the given video . Due to its popularity and applicability, HD can be divided into several branches. From the perspective of annotation, we can categorize HD into supervised, weakly supervised, and unsupervised HD. Supervised HD utilizes fine-grained highlight scores, which are very expensive to collect and annotate . On the other hand, weakly supervised HD learns to detect segments as highlights with video event labels, and finally, unsupervised HD does not require any annotations. Also, while the task is often implemented only with the video, there are works to employ the extra data modalities. Generally, multi-modality was taken into account by using the natural language query to find the desired thumbnail and using additional sources, i.e., audio, to predict highlights .

Although the MR/HD share a common objective to localize or discover the desired part of the given video, they have been studied separately. To handle these tasks at once, Moment-DETR proposed the QVHighlights dataset, which contains a human-written text query and its corresponding moment with clip-level saliency labels. They also introduced the modified version of detection transformer (DETR ) to localize the query-relevant moments and their saliency scores. Following them, UMT focused on processing multi-modal data by utilizing both video and audio features. Different from recent works deploying transformer architectures, here we concentrate on producing a query-dependent representation with the transformer.

2 Detection Transformers

DETR , an end-to-end object detector based on vision transformers, is one of the very recent works that utilize the transformer architectures for computer vision . Although DETR suffered from slow convergence, it simplifies the prediction process by eliminating the need for anchor generation and non-maximum suppression. Since then, along with the advance in DETR , DETR-like architectures have been popular in downstream tasks in both the image and video domains . Some of these works focused on analyzing the role of the decoder query and discovered that using the positional information speeds up the training and also enhances the detection performance . On the other hand, there are trials to extend the application of DETR on multi-modal data , especially dealing with the query from different modalities, i.e., text, for detection (or retrieval). They generally handle the multi-modal data by simply forwarding them together to the transformer. In this paper, we also focus on handling the multi-modal data based on DETR-like architecture. However, different from the aforementioned techniques, we concentrate on the query-dependency of the prediction results.

Query-Dependent DETR

Moment retrieval and highlight detection have the common objective to find preferred moments with the text query. Given a video of LL clips and a text query with NN words, we denote their representations as {v1,v2,...,vL}\{v_{1},v_{2},...,v_{L}\} and {t1,t2,...,tN}\{t_{1},t_{2},...,t_{N}\} extracted by frozen video and text encoders, respectively. With these representations, the main objective is to localize the center coordinate mcm_{c} and width mσm_{\sigma} within the video and rank the highlight score (saliency score) {s1,s2,...,sL}\{s_{1},s_{2},...,s_{L}\} for each clip. A straightforward approach to utilize transfomer for the MR is to make a moment-wise prediction as a set of clips , or generating the moment according to the clip-wise predictions . To exploit the multi-modal information, e.g., video and text query, they either simply concatenated the features across the modalities or inserted the texts to form the moment query to the transformer decoder. However, we claim that the relationship between the video and text query should be carefully considered rather than a simple concatenation since MR/HD requires every video clip to be conditionally assessed with the text queries.

Our overall architecture is described in Fig. 2, following the design of concrete baseline, Moment-DETR . Given a video and query representation extracted from fixed backbones, QD-DETR first transforms the video representation to be query-dependent using cross-attention layers. To further enhance the query-awareness of video representations, we incorporate irrelevant video-query pairs with a low saliency for the learning objective. Then, along with the transformer encoder-decoder architectures, the saliency token is defined that turns into an adaptive saliency predictor when attended by the specific video instance.

In this subsection, we use italic letters to represent query, key, and value of the cross-attention layers. The key objective of the encoder for MR/HD is to produce clip-wise representations equipped with information regarding the degree of query-relevance since these features are directly used for retrieving the query-matched moments and predicting clip-wise saliency scores. However, the encoding process of existing works may not ensure the query conditioning on every clip. For example, Moment-DETR naively concatenated the video with the query for input to the self-attention layers, which may result in an insignificant role of the query if the high similarities among the video clips overwhelm the contribution of the text query. On the other hand, UMT utilizes the text query only for the synthesis of a moment query in the transformer decoder so thus resulting video representations are not associated with the text query.

To take the textual contexts into every video clip representation, we deploy cross-attention layers between the source and the query modalities at the very first layers of the encoder. This ensures the consistent contribution of the query, thereby extracting query-dependent video representation. In detail, whereas the query for cross-attention layers is prepared by projecting the video clips as Qv=[ pq(v1),...,pq(vL) ]Q_{v}=[~{}p_{q}(v_{1}),...,p_{q}(v_{L})~{}], the key and value are computed with the query text features as Kt=[ pk(t1),...,pk(tN) ]K_{t}=[~{}p_{k}(t_{1}),...,p_{k}(t_{N})~{}] and Vt=[ pv(t1),...,pv(tN) ]V_{t}=[~{}p_{v}(t_{1}),...,p_{v}(t_{N})~{}]. pq(⋅)p_{q}(\cdot), pk(⋅)p_{k}(\cdot), and pv(⋅)p_{v}(\cdot) are projection layers for query, key, and value. Then, the cross-attention layer operates as follows:

where dd is the dimension of the projected key, value, and query. Since the softmax scores are distributed only over the query elements, video clips are expressed with the weighted sum of the text queries in proportion to the similarity to texts. Attention scores are then projected through MLP and integrated into the original video representations as the typical transformer layers. For the rest of the paper, we define the query-dependent video tokens, i.e., the output of cross-attention layers, as X={xv1,xv2,...,xvL}X=\{x_{v}^{1},x_{v}^{2},...,x_{v}^{L}\}.

2 Learning from Negative Relationship

While the cross-attention layers explicitly fuse the video and query features for intermediate video clip representations to engage the query information in an architectural way, we argue that given video-text pairs lack diversity to learn the general relationship. For instance, many consecutive clips in a single video often share similar appearances, and the similarity to a specific query will not be highly distinguishable, thereby, the text query may not much affect the prediction.

Thus, we consider the relationships between irrelevant pairs of videos inspired by many recognition practices that learn discriminative features across different categories. To implement such relationships, we define given training video-query pairs as positive pairs and mix the video and query from different pairs to construct negative pairs. Fig. 3 illustrates the ways to augment such negative pairs and utilize them with positive pairs in training. While the video clips in positive pairs are trained to yield segmented saliency scores according to the query-relevance, irrelevant negative video-query pairs are enforced to have the lowest saliency scores. Formally, the loss function for suppressing the saliency of negative pairs xvnegx_{v}^{\text{neg}} are expressed as follows:

where S(⋅)S(\cdot) is the saliency score predictor. This training scheme can also prevent the model from predicting the moments and highlights solely based on the inter-relationship among video clips without consideration of the query-relevance since the same video instance should be predicted differently depending on whether the positive or negative query is given.

3 Input-Adaptive Saliency Predictor

Naive implementation for saliency predictor S(⋅)S(\cdot) would be stacking one or more fully-connected layers. However, such a general head provides identical criteria for the saliency prediction of every video-query pair, neglecting the diverse nature of video and natural language query pairs. This violates our key idea to extract query-dependent video representation.

Thus, we define the saliency token xsx_{s} to be utilized as an input-adaptive saliency predictor. Briefly, the saliency token is a randomly initialized learnable vector that becomes an input-adaptive predictor when added to the sequence of encoded video tokens and projected through the transformer encoder. To illustrate, as shown in Fig. 2, we first concatenate the saliency token with the query-dependent video tokens XX. We process these tokens to the transformer encoder which makes the saliency token to be re-organized with the input-dependent contexts. Consequently, saliency and video tokens are projected by a corresponding single fully-connected layer with weights, wsw_{s} and wvw_{v}, respectively, where their scaled-dot product becomes the saliency scores. Formally, saliency score S(xvi)S(x_{v}^{i}) is computed as follows:

where dd is the channel dimension of projected tokens.

4 Decoder and Objectives

Recently, understanding the role of the query in the detection transformer is being spotlighted . It is verified that designing the query with the positional information helps not only for acceleration of training but also for enhancing accuracy. Yet, it is hard to directly employ these studies in tasks handling multi-modal data, e.g., MR/HD, since multi-modal data often have different definitions of position; the position can be understood as time in the video and word order in the text.

On the contrary, our architectural design eliminates the need to feed the text query to the decoder since the query information is already taken into the video representations. To this end, we modify the 2D dynamic anchor boxes to represent 1D moments in the video. Specifically, we utilize the center coordinate mcm_{c} and the duration mσm_{\sigma} of the moments to design the queries. Similarly to the previous way in the image domain, we pool the features around the center coordinate and modulate the cross-attention map with the moment duration. Then, the coordinates and durations are layer-wisely revised.

Loss Functions.

Training objectives for QD-DETR include loss functions for MR/HD, respectively. First, objective functions for MR, in which the key focus is to locate the desired moments, are adopted from the baseline . Moment retrieval loss LmrL_{\text{mr}} measures the discrepancy between the GT moment and the predicted counterpart. It consists of a L1L1 loss and a generalized IoU loss LgIoU(⋅)L_{\text{gIoU}}(\cdot) from previous work with minor modification to localize temporal moments. Additionally, the cross-entropy loss is used to classify the predicted moments as y^\hat{y} either to foreground and background by LCE=−∑y∈Yylog⁡(y^)L_{\text{CE}}=-\sum_{y\in Y}y\log(\hat{y}) where {fg,bg}⊂Y\{\text{fg},\text{bg}\}\subset Y. Thus, LmrL_{\text{mr}} is defined as follows:

where mm and m^\hat{m} are ground-truth moment and its correspond prediction containing center coordinate mcm_{c} and duration mσm_{\sigma}. Also, λ∗\lambda_{*} are hyperparameters for balancing the losses.

Loss functions for HD are to estimate the saliency score. It comprises two components; margin ranking loss LmarginL_{\text{margin}} and rank-aware contrastive loss LcontL_{\text{cont}}. Following , the margin rank loss operates with two pairs of high-rank and low-rank clips. To be specific, the high-rank clips are ensured to retain higher saliency scores than both the low-rank clips within the GT moment and the negative clips outside the GT moment. In short, LmarginL_{\text{margin}} is defined as:

where Δ\Delta is the margin, S(⋅)S(\cdot) is the saliency score estimator, and xhighx^{\text{high}} and xlowx^{\text{low}} are video tokens from two pairs of high and low-rank clips, respectively. In addition to margin loss which only indirectly guides the saliency predictor, we employ rank-aware contrastive loss to learn the precisely segmented saliency levels with the contrastive loss. Given the maximum rank value RR, each clip in the mini-batch has a saliency score lower than RR. Then, we iterate the batch for RR times, each time utilizing the samples with higher saliency scores than the iteration index (r∈{0,1,...,R−1}r\in\{0,1,...,R-1\}) to build the positive set XrposX^{\text{pos}}_{r}. Samples with a lower rank than the iteration index are included in the negative set XrnegX^{\text{neg}}_{r}. Then, the rank-aware contrastive loss LcontL_{\text{cont}} is defined as:

where τ\tau is a temperature scaling parameter. Note that, XrnegX^{\text{neg}}_{r} also include all clips in negative pairs xvnegx_{v}^{\text{neg}} defined in Sec. 3.2. Finally with margin loss and rank-aware contrastive loss, LhlL_{\text{hl}} and total loss function LtotalL_{\text{total}} are defined as follows:

Evaluation

We compare QD-DETR against baselines in MR and HD throughout Tab. 4, Tab. 3, and Tab. 4.1. Our experiments with multi-modal sources, i.e., video with audio, are implemented by simply concatenating the video and audio along the channel axis. Throughout the tables, we use bolds to denote the best scores.

In Tab. 4, the task is to jointly learn and predict MR/HD. As observed, our QD-DETR outperforms state-of-the-art (SOTA) approaches with all evaluation metrics. Among methods utilizing the video source, QD-DETR shows a dramatic increase with stricter metrics with high IOU; it outperforms previous SOTA by large margins up to 36% in R1@0.7 and mAP@0.75. On the other hand, QD-DETR with video and audio sources boosts 11.84% on average of the metrics for MR compared to the SOTA method employing the multi-modal source data. These results verify the importance of emphasizing the source (video-only or video+audio) descriptive contexts in the text queries.

To investigate the effectiveness of each component in our work, we conduct an extensive ablation study in Tab. 4.2. Note that, CATE and DAM denote cross-attentive transformer encoder and dynamic anchor moments, respectively. Rows (b) to (e) show the effectiveness of each component compared to the baseline (a). To explain, whereas (e) only boosts the MR performances since it only affects the transformer decoder, (b), (c), and (d) are especially beneficial for both MR/HD tasks since they are focused on query-dependent video representations ; (b) ensures the contributions of text query in the video representation, (c) fully exploits the contexts of the text query, and (d) provides input-adaptive saliency predictor instead of MLP. Moreover, while our components are verified that they are all complementary to others, DAM’s effectiveness is especially dependent on the usage of CATE (compare between {(a, e)} and {(b, f), (i, j)}). We claim that this is because DAM exploits the position information of the input tokens to capture the corresponding moments. However, without CATE, input tokens are a mixture of multi-modal tokens, thereby providing confusing position information.

To provide in-depth examinations of each component, we inspect the difference between the positive and the negative saliency scores in Fig. 4. Since the role of text query is trivial in our baseline, each distribution significantly overlies on top of the other. Then, as we add CATE and negative pair learning, we observe a consistent decrease in overlapped areas and a larger gap between the average saliency scores of positive and negative histograms. Also, we believe that the widely-distributed histogram of saliency scores for the ’CATE+Neg.pair’ is due to using an identical criterion for saliency prediction for diverse video-query representations. By employing an input-adaptive saliency predictor, we notice that scores for positive queries are in almost optimal shape.

In addition, some might ask whether CATE benefits the training because of additional encoder layers. To answer this, we conduct another ablation study in Tab. 4.2. Briefly, since our transformer encoder utilizes 2 cross-attention layers and 2 self-attention layers, we conduct comparisons against the transformer encoder composed of 4 self-attention layers (SATE). First, we compare CATE and SATE on Moment-DETR; by comparing the results in 2nd2^{\text{nd}} and 3rd3^{\text{rd}} rows, we find that CATE is much more beneficial than SATE even with the same number of layers. Furthermore, the last two rows show comparisons within QD-DETR architecture that has the same tendency. These results clearly demonstrate that the improvements from CATE are mainly from emphasizing the role of the text query rather than additional layers.

In this subsection, we study how the query-dependent video representation sensitively reacts to the change in the contexts of the text query. In Fig. 5, the measured saliency scores according to the video-query relevance are visualized. We found that the more the query is relevant to the video clips, the higher the saliency scores retained for the query. For instance, whereas the negative query that is totally irrelevant to the video instance has the lowest scores, the scores for semi-positive reside between the positive and the negative ones. Also, we find that QD-DETR sometimes provides a more precise moment prediction than a given ground-truth moment, as can be seen with the temporal box bounded by the dotted lines. We believe that the tendency of a bit higher saliency scores at non-relevant clips for a positive query is due to the information mixing in the self-attention layers.

As elaborated in the paper, we aim to highlight the role of the text query in retrieving the relevant moments and estimating their accordance level with the given text query. Likewise, the proposed components expect a given query to maintain a meaningful context. If not, and noisy text queries are provided, i.e., mismatched or irrelevant ground truth texts, the training may not be effective as reported.

Conclusion

Although the advent of transformer architecture has been powerful for MR/HD, investigation of the role of text query has been lacking in such architectures. Therefore, we focused on studying the role of the text query. As we found that the textual information is not fully exploited in expressing the video representations, we designed the cross-attentive transformer encoder and proposed a negative-pair training scheme. Cross-attentive encoder assures the query’s contributions while extracting video representation, and negative-pair training enforces the model to learn the relationship between query and video by preventing solving the problems without consideration of the query. Finally, to preserve the diversity of query-dependent video representation, we defined the saliency token to be an input-adaptive saliency predictor. Extensive experiments validated the strength of QD-DETR with superior performances.

Acknowledgements. This work was supported in part by MSIT/IITP (No. 2022-0-00680, 2019-0-00421, 2020-0-01821, 2021-0-02068), and MSIT&KNPA/KIPoT (Police Lab 2.0, No. 210121M06).

In this section, we elaborate on the implementation details and hyperparameters used for experiments in the main manuscript. To unify configurations across all experiments, our encoder composes of 4 layers of transformer block (2 cross-attention layers and 2 self-attention layers) whereas there are only 2 layers in the decoder (For HD dataset, i.e., TVSum, we only use encoding layers). We set the hidden dimension of transformers as 256, and use the Adam optimizer with a weight decay of 1e-4. Besides, we set the temperature of a scaling parameter τ\tau for contrastive loss as 0.5 for all experiments. Loss balancing parameters are λmargin=1\lambda_{\text{margin}}=1, λcont=1\lambda_{\text{cont}}=1, λL1=10\lambda_{L1}=10, λgIoU=1\lambda_{\text{gIoU}}=1, λCE=4\lambda_{\text{CE}}=4 and λneg=1\lambda_{\text{neg}}=1, unless otherwise mentioned. Additionally, we use the PANN model trained on AudioSet to extract audio features1 for experiments with the audio modality.

Other configurations are described as follows:

QVHighlight. We use video features extracted from both pretrained SlowFast (SF) and CLIP encoder , and text embeddings from CLIP, following the Moment-DETR. We train QD-DETR for 200 epochs with a batch size of 32 and a learning rate of 1e-4.

Charades-STA. We utilize official VGG features with GloVe text embedding. To compare with additional baselines, we also test our model on pretrained C3D , SlowFast and CLIP for video features with CLIP text embedding. Specifically, we utilize pre-extracted features provided by other baselines repositories: UMThttps://github.com/TencentARC/UMT, VSLNethttps://github.com/IsaacChanghau/VSLNet and Moment-DETRhttps://github.com/jayleicn/moment_detr. We train ours for 100 epochs with a batch size of 8 and a learning rate of 1e-4.

TVSum. I3D features pretrained on Kinetics-400 are utilized as a visual one, and CLIP features are used for the text embedding. Following the most recent work , we train our model for 2000 epochs with a learning rate of 1e-3. The batch size is set to 4.

As discussed in the limitation, the performance of QD-DETR may depend on the quality of provided ground truth text descriptions. Yet, this does not imply the QD-DETR’s vulnerability against commonly used meaningless words in text descriptions. As we think the queries with longer lengths may have a higher chance of including noisy texts, we divide the validation set into 3 groups each with long-, medium-, and short-length queries, and report the query-length-wise performances of QD-DETR in Sec. 7. As shown, QD-DETR works well regardless of the query length, showing [36.7, 28.0, 26.3%] and [7.3, 11.8, 11.1%] improvements in mAP each for MR and HD with [Short, Medium, Long] queries. This study implies that while irrelevant (wrong) text descriptions for video contexts can degrade the effectiveness of QD-DETR, QD-DETR is robust against meaningless words that are commonly present in text queries.