Cross-Modal Interaction Networks for Query-Based Moment Retrieval in Videos

Zhu Zhang, Zhijie Lin, Zhou Zhao, Zhenxin Xiao

Introduction

Multimedia information retrieval is an important topic in information retrieval systems. Recently, query-based video retrieval (Lin et al., 2014; Xu et al., 2015; Otani et al., 2016) has been well-studied, which searches the most relevant video from large collections according to a given natural language query. However, in practical applications, the untrimmed videos often contain multiple complex events that evolve over time, where a large part of video contents are irrelevant to the query and only a small clip satisfies the query description. Thus, as a natural extension, the query-based moment retrieval aims to automatically localize the start and end boundaries of the target moment semantically corresponding to the given query within an untrimmed video. Different from retrieving an entire video, moment retrieval offers more fine-grained temporal localization in a long video, which avoids manually searching for the moment of interests.

However, localizing the precise moment in a continuous, complicated video is more challenging than simply selecting a video from pre-defined candidate sets. As shown in Figure 1, the query “A man throws the ball and hits the boy in the face before landing in a cup” describes two successive actions, corresponding to complex object interactions within the video. Hence, the accurate retrieval of the target moment requires sufficient understanding of both video and query contents by cross-modal interactions.

Most existing moment retrieval works (Gao et al., 2017; Hendricks et al., 2018, 2017; Liu et al., 2018a, b; Chen et al., 2018; Xu et al., 2019) only focus on one aspect of this emerging task, such as the query representation learning (Liu et al., 2018b), video context modeling (Gao et al., 2017; Liu et al., 2018a) and cross-modal fusion (Chen et al., 2018; Xu et al., 2019), thus fail to develop a comprehensive system to further improve the performance of query-based moment retrieval. In this paper, we consider multiple crucial factors for high-quality moment retrieval.

Firstly, the query description often contains causal temporal actions, thus it is fundamental and crucial to learn fine-grained query representations. Existing works generally adopt widely-used recurrent neural networks, such as GRU networks, to model natural language queries. However, these approaches ignore the syntactic structure of queries. As shown in Figure 1, the syntactic relation implies dependencies of word pairs, helpful for query semantic understanding. Recently, the graph convolution networks (GCN) have been proposed to model the graph structure (Kipf and Welling, 2017; Velickovic et al., 2018), including the visual relationship graph (Yao et al., 2018) and syntactic dependency graph (Marcheggiani and Titov, 2017). Inspired by these works, we develop a syntactic GCN to exploit the syntactic structure of queries. Concretely, we first build the syntactic dependency graph as shown in Figure 1, and then pass information along the dependency edges to learn syntactic-aware query representations. In detail, we consider the direction and label of dependency edges to adequately incorporate syntactic clues.

Secondly, the target moments of complex queries generally contain object interactions over a long time interval, thus there exist long-range semantic dependencies in video context. That is, each frame is not only relevant to adjacent frames, but also associated with distant ones. Existing approaches often apply RNN-based temporal modeling (Chen et al., 2018), or propose R-C3D networks to learn spatio-temporal representations from raw video streams (Xu et al., 2019). Although these methods are able to absorb contextual information for each frame, they still fail to build direct interactions between distant frames. To eliminate the local restrictions, we propose a multi-head self-attention mechanism (Vaswani et al., 2017) to capture long-range semantic dependencies from video context. The self-attention method can develop the frame-to-frame interaction at arbitrary positions and the multi-head setting ensures the sufficient understanding of complicated dependencies.

Thirdly, query-based moment retrieval requires the comprehensive reasoning of video and query contents, thus the cross-modal interaction is necessary for high-quality retrieval. Early approaches (Gao et al., 2017; Hendricks et al., 2017, 2018) ignore this factor and only simply combine the query and moment features for correlation estimations. Although recent methods (Liu et al., 2018a, b; Chen et al., 2018; Xu et al., 2019) have developed a cross-modal interaction by widely-used attention mechanism, they still remain in the rough one-stage interaction, for example, highlighting the crucial context information of moments by the guidance of queries (Liu et al., 2018a). Different from previous works, we adopt a multi-stage cross-modal interaction method to further exploit the potential relation of video and query contents. Specifically, we first adopt a normal attention method to aggregate syntactic-aware query representations for each frame, then apply a cross gate (Feng et al., 2018) to emphasize crucial contents and weaken inessential parts, and next develop the low-rank bilinear fusion to learn a cross-modal semantic representation.

In summary, the key contributions of this work are four-fold:

We design a novel cross-modal interaction networks for query-based moment retrieval, which is a comprehensive system to consider multiple crucial factors of this challenging task: (1) the syntactic structure of natural language queries; (2) long-range semantic dependencies in video context and (3) the sufficient cross-modal interaction.

We propose the syntactic GCN to leverage the syntactic structure of queries for fine-grained representation learning, and adopt a multi-head self-attention method to capture long-range semantic dependencies from video context.

We employ a multi-stage cross-modal interaction to further exploit the potential relation of video and query contents, where an attentive aggregation method extracts relevant syntactic-aware query representations for each frame, a cross gate emphasizes crucial contents and a low-rank bilinear fusion method learn cross-modal semantic representations.

The proposed CMIN method achieves the state-of-the-art performance on ActivityCaptions and TACoS datasets.

The rest of this paper is organized as follows. We briefly review some related works in Section 2. In Section 3, we introduce our proposed method. We then present a variety of experimental results in Section 4. Finally, Section 5 concludes this paper.

Related Work

In this section, we briefly review some related works on image/video retrieval, temporal action localization and query-based moment retrieval.

Given a set of candidate images/videos and a natural language query, image/video retrieval aims to select the image/video that matches this query. Karpathy et al. (Karpathy and Fei-Fei, 2015) propose a deep visual-semantic alignment (DVSA) model for image retrieval, which uses the BiLSTM to encode query features and R-CNN detector (Girshick et al., 2014) to extract object representations. Sun et al. (Sun et al., 2015) advise an automatic visual concept discovery algorithm to boost the performance of image retrieval. Moreover, Hu et al. (Hu et al., 2016) and Mao et al (Mao et al., 2016) regard this problem as natural language object retrieval. As for video retrieval, some methods (Otani et al., 2016; Xu et al., 2015) incorporate deep video-language embeddings to boost retrieval performance, similar to the image-language embedding approach (Socher et al., 2014). And Lin et al. (Lin et al., 2014) first parse the query descriptions into a semantic graph and then match them to visual concepts in videos. Different from these works, query-based moment retrieval aims to localize a moment within an untrimmed video, which is more challenging than simply selecting a video from pre-defined candidate sets.

2. Temporal Action Localization

Temporal action localization is a challenging task to localize action instances in an untrimmed video. Shou et al. (Shou et al., 2016) develop three segment-based 3D ConvNets with localization loss to explicitly explore the temporal overlap in videos. Singh et al. (Singh et al., 2016) propose a multi-stream bi-directional RNN with two additional streams on motion and appearance to achieve the fine-grained action detection.

To leverage the context structure of actions, Zhao et al. (Zhao et al., 2017) advise a structured segment network to model the structure of action instances by a structured temporal pyramid. And Chao et al. (Chao et al., 2018) boost the action localization performance by imitating the Faster RCNN object detection framework (Ren et al., 2015). Although these works have achieved promising performance, they still are limited to a pre-defined list of actions. And query-based moment retrieval tackles this problem by introducing the natural language query.

3. Query-Based Moment Retrieval

Query-based moment retrieval is to detect the target moment depicting the given natural language query in an untrimmed video. Early works study this task in constrained settings, including the fixed spatial prepositions (Tellex and Roy, 2009; Lin et al., 2014), instruction videos (Alayrac et al., 2016; Sener et al., 2015; Song et al., 2016) and ordering constraint (Bojanowski et al., 2015; Tapaswi et al., 2015). Recently, unconstrained query-based moment retrieval has attracted a lot of attention (Hendricks et al., 2017; Gao et al., 2017; Liu et al., 2018a; Hendricks et al., 2018; Liu et al., 2018b; Xu et al., 2019; Chen et al., 2018). These methods are mainly based on a sliding window framework, which first samples candidate moments and then ranks these moments. Hendricks et al. (Hendricks et al., 2017) propose a moment context network to integrate global and local video features for natural language retrieval, and the subsequent work (Hendricks et al., 2018) considers the temporal language by explicitly modeling the context structure of videos. Gao et al. (Gao et al., 2017) develop a cross-modal temporal regression localizer to estimate the alignment scores of candidate moments and textual query, and then adjust the boundaries of high-score moments. With the development of attention mechanism in the field of vision and language interaction (Anderson et al., 2018; Zhao et al., 2018), Liu et al. (Liu et al., 2018a) advise a memory attention to emphasize the visual features and simultaneously utilize the context information. And similar attention strategy (Liu et al., 2018b) is designed to highlight the crucial part of query contents. From the holistic view, Chen et al. (Chen et al., 2018) capture the evolving fine-grained frame-by-word interactions between video and query. Xu et al. (Xu et al., 2019) introduce a multi-level model to integrate visual and textual features earlier and further re-generate queries as an auxiliary task.

Unlike these previous methods, we propose a novel cross-modal interaction network to consider three critical factors for query-based moment retrieval, including the syntactic structure of natural language queries, long-range semantic dependencies in video context and the sufficient cross-modal interaction.

Cross-Modal Interaction Networks

As Figure 2 illustrates, our cross-modal interaction networks consist of four components: 1) the syntactic GCN module leverages the syntactic structure to enhance the query representation learning; 2) the multi-head self-attention module captures long-range semantic dependencies from video context; 3) the multi-stage cross-modal interaction module aggregates syntactic-aware query representations for each frame, emphasizes crucial contents and learns cross-modal semantic representations; 4) the moment retrieval module finally localizes the boundaries of target moments.

We present a video as a sequence of frames v={vi}i=1n∈V{\bf v}=\{{\bf v}_{i}\}_{i=1}^{n}\in V, where vi{\bf v}_{i} is the feature of the ii-th frame and nn is the frame number of the video. Each video is associated with a natural language query, denoted by q={qi}i=1m∈Q{\bf q}=\{{\bf q}_{i}\}_{i=1}^{m}\in Q, where qi{\bf q}_{i} is the feature of the ii-th word and mm is the word number of the query. The query description corresponds to a target moment in the untrimmed video and we denote the start and end boundaries of the target moment by τ=(s,e)∈A{\bf\tau}=(s,e)\in A. Thus, given the training set {V,Q,A}\{V,Q,A\}, our goal is to learn the cross-modal interaction networks to predict the boundary τ^=(s^,e^){\hat{\tau}}=({\hat{s}},{\hat{e}}) of the most relevant moment during inference.

2. Syntactic GCN Module

In this section, we introduce the syntactic GCN module based on a syntactic dependency graph. By passing information along the dependency edges between relevant words, we learn syntactic-aware query representations for subsequent cross-modal interactions.

We first extract word features for the query using a pre-trained Glove word2vec embedding (Pennington et al., 2014), denoted by q=(q1,q2,…,qm){\bf q}=({\bf q}_{1},{\bf q}_{2},\ldots,{\bf q}_{m}), where qi{\bf q}_{i} is the feature of the ii-th word. After that, we develop a bi-directional GRU networks (BiGRU) to learn the query semantic representations. The BiGRU networks incorporate contextual information for each word by combining the forward and backward GRU (Chung et al., 2014) . Sepcifically, we input the sequence of word features to the BiGRU networks, and obtain the contextual representation of each word, given by

where GRUqf{\rm GRU}^{f}_{q} and GRUqb{\rm GRU}^{b}_{q} represent the forward and backward GRU networks, respectively. And the contextual representation hiq{\bf h}^{q}_{i} is the concatenation of the forward and backward hidden state at the ii-th step. Thus, we get the query semantic representations hq=(h1q,h2q,…,hmq){\bf h}^{q}=({\bf h}^{q}_{1},{\bf h}^{q}_{2},\ldots,{\bf h}^{q}_{m}).

Although the BiGRU networks have encoded temporal context of word sequences, they still ignore the syntactic information of natural language, which implies underlying dependencies between word pairs. So we then advise the syntactic graph convolution networks to leverages the syntactic dependencies for better query understanding. We first build the syntactic dependency graph by an NLP toolkit, where each word is regarded as a node and each dependency relation is presented as a directed edge. Formally, we denote a query by a graph G=(V,E)\mathcal{G=(V,E)}, where the node set V\mathcal{V} contains all words and edge set E\mathcal{E} contains all directed syntactic dependencies of word pairs. Note that we add the self-loop for each node into the edge set. As the dependency relations have different types, the directed edges also correspond to different labels, where the self-loop is given a unique label. The original GCN regard the syntactic dependency graph as an undirected graph, denoted by

where Wg{\bf W}^{g} is the transformation matrix, bg{\bf b}^{g} is the bisa vector and ReLU{\rm ReLU} is the rectified linear unit. The N(i){\mathcal{N}}(i) represents the set of nodes with a dependency edge to node ii or from node ii (including self-loop). And the hjq{\bf h}_{j}^{q} is the original representation of node jj from the preceding modeling.

Although the original GCN enhances the word semantic representation by aggregating the clues from its neighbors, it fails to leverage the direction and label information of edges. Thus, we consider a syntactic GCN to exploit the directional and labeled dependency edges between nodes, given by

where dir(i,j){dir(i,j)} indicates the direction of edge (i,j)(i,j): 1) a dependency edge from node ii to jj; 2) a dependency edge from node jj to ii; 3) a self-loop if i=ji=j. Since there is no reason to assume the information transmits only along the syntactic dependency arcs (Marcheggiani and Titov, 2017), we also allow the information to transmit in the opposite direction of directed dependencies here. The Figure 3 describes the three types of information passing, which correspond to three transformation matrix W1g{\bf W}^{g}_{1},W2g{\bf W}^{g}_{2} and W3g{\bf W}^{g}_{3}, respectively. On the other hand, the lab(i,j)lab(i,j) represents the label of edge (i,j)(i,j) to select a distinct bias vector for each type of dependencies. Next, we employ a residual connection (He et al., 2016) to keep the original representation of each node, given by

Furthermore, we stack a multi-layer syntactic GCN to adequately explore the syntactic structure as follows.

By the syntactic GCN with ll layers, we obtain the syntactic-aware query representation ol=(o1l,o2l,…,oml){\bf o}^{l}=({\bf o}^{l}_{1},{\bf o}^{l}_{2},\ldots,{\bf o}^{l}_{m}).

3. Multi-Head Self-Attention Module

In this section, we present the multi-head self-attention module to capture long-range semantic dependencies from video context. By the self-attention method, each frame is able to interact not only with adjacent frames but also with distant ones. And the multi-head setting is beneficial to sufficiently understand the complicated dependencies.

We first extract frame features from the untrimmed video by a pre-trained 3D-ConvNet (Tran et al., 2015), denoted by v=(v1,v2,…,vn){\bf v}=({\bf v}_{1},{\bf v}_{2},\ldots,{\bf v}_{n}). The vi{\bf v}_{i} is the visual feature of the ii-th frame. We then introduce the multi-head self-attention based on the scaled dot-product attention, which originally is proposed in the field of machine translation (Vaswani et al., 2017).

where the Softmax{\rm Softmax} operation is performed on every row. The values are aggregated for each query according to the dot-product score between the query and the corresponding key of values.

Multi-head attention. The multi-head attention consists of HH paralleled scaled dot-product attention layers. For each independent attention layer, the input queries, keys and values are linearly projected to dkd_{k}, dkd_{k} and dvd_{v} dimenstions. Concretely, the result of multi-head attention is given by

where a residual connection is applies similar to syntactic GCN. Here we establish frame-to-frame correlation among video sequences and the multi-head setting allows the attention operation to aggregate information from different representation subspaces. But the temporal modeling is still critical for video semantic understanding, we cannot only depend on the multi-head self-attention and ignore contextual representation learning. Hence, we next employ another BiGRU to learn the self-attentive video semantic representations hv=(h1v,h2v,…,hnv){\bf h}^{v}=({\bf h}^{v}_{1},{\bf h}^{v}_{2},\ldots,{\bf h}^{v}_{n}).

4. Multi-Stage Cross-Modal Interaction Module

In this section, we introduce the cross-modal interaction module to exploit the potential relations of video and query contents, which consists of the attentive aggregation, cross-gated interaction and low-rank bilinear fusion.

where W1m{\bf W}_{1}^{m}, W2m{\bf W}_{2}^{m} are parameter matrices, bm{\bf b}^{m} is the bias vector and the w⊤{\bf w}^{\top} is the row vector. We then apply the softmax operation for each row of MM, given by

where MijrowM^{row}_{ij} represents the correlation of the ii-th frame and jj-th word. Next, we extract the crucial query clues for each frame based on MrowM^{row}, given by

where the his{\bf h}^{s}_{i} represents the aggregated query representation relevant to the ii-th frame.

Cross-gated interaction. With the aggregated query representation his{\bf h}^{s}_{i} and frame semantic representation hiv{\bf h}^{v}_{i}, we then apply a cross gate (Feng et al., 2018) to emphasize crucial contents and weaken inessential parts. In the cross gate method, the gate of query representation depends on the frame representation, and meanwhile the frame representation is also gated by its corresponding query representation, denoted by

where Wv{\bf W}^{v}, Ws{\bf W}^{s} are parameter matrices, bv{\bf b}^{v} and bs{\bf b}^{s} are the bias vectors, σ\sigma is the sigmoid function, and ⊙\odot represents element-wise multiplication. If the aggregated query representation his{\bf h}^{s}_{i} is irrelevant to the frame semantic representation hiv{\bf h}^{v}_{i}, both the two representations are filtered to decrease their influences on subsequent networks. On the contrary, the cross gate can further enhance the effects of relevant frame-query pairs.

Bilinear fusion. After the attentive aggregation and cross gate, we propose a low-rank bilinear fusion method (Kim et al., 2017) to further exploit the cross-modal interaction between h~is{\bf\widetilde{h}}^{s}_{i} and h~iv{\bf\widetilde{h}}^{v}_{i}. The original bilinear fusion method is written by

where fijf_{ij} represents the jj-th dimension of the bilinear output at the time step ii, and the fi{\bf f}_{i} is the fusion result of h~is{\bf\widetilde{h}}^{s}_{i} and h~iv{\bf\widetilde{h}}^{v}_{i}. the Wjf{\bf W}^{f}_{j}, bjf{\bf b}^{f}_{j} are the parameter matrix and the bias vectors for the jj-th dimension. We can note that the original bilinear fusion method requires too many parameters and suffers from the heavy computation cost. Thus, we replace it with the low-rank version (Kim et al., 2017), given by

where the fi{\bf f}_{i} is the biliner fusion result at the time step ii.

Eventually, by the attentive aggregation, cross-gated interaction and low-rank bilinear fusion, we obtain the cross-modal semantic representations for each frame, denoted by f=(f1,f2,…,fn){\bf f}=({\bf f}_{1},{\bf f}_{2},\ldots,{\bf f}_{n}).

5. Moment Retrieval Module

In this section, we present the moment retrieval module to simultaneously score a set of candidate moments with multi-scale windows at each time step, and further adopt a temporal boundaries regression mechanism to adjust the moment boundaries.

By the cross-modal interaction module, we get the the cross-modal semantic representations f=(f1,f2,…,fn){\bf f}=({\bf f}_{1},{\bf f}_{2},\ldots,{\bf f}_{n}). To absorb the contextual evidences, we likewise develop another BiGRU networks to learn the final semantic representations hf=(h1f,h2f,…,hnf){\bf h}^{f}=({\bf h}^{f}_{1},{\bf h}^{f}_{2},\ldots,{\bf h}^{f}_{n}). We then pre-define a set of candidate moments with multi-scale windows at each time step ii, denoted by Ci={(s^ij,e^ij)}j=1kC_{i}=\{({\hat{s}}_{ij},{\hat{e}}_{ij})\}_{j=1}^{k}, where (s^ij,e^ij)=(i−wj/2,i+wj/2)({\hat{s}}_{ij},{\hat{e}}_{ij})=(i-w_{j}/2,i+w_{j}/2) are the start and end boundaries of the jj-th candidate moment at time ii, wjw_{j} is the width of jj-th moment and kk is the number of moments. Note that we set the fixed window width wjw_{j} for jj-th candidate moment at every time step. Thus, we can simultaneously produce the confidence scores for these moments at time ii by a fully connected layer with sigmoid nonlinearity, given by

Alignment loss. We first adopt an alignment loss to make the moment aligned to the target moment have high confidence scores and the misaligned moment have low confidence scores. Formally, we first compute the IoU (i.e. Intersection over Union) score IoUijIoU_{ij} of each candidate moment Cij=(s^ij,e^ij)C_{ij}=({\hat{s}}_{ij},{\hat{e}}_{ij}) with the target moment (s,e)(s,e). If the IoU score of a candidate is less than a clearing threshold λ\lambda, we reset it to 0. Next, we calculate the alignment loss by

where we consider all candidate moments during alignment training, and apply the concrete IoU score rather than set 0 or 1 according to a threshold value for every candidate. This setting is helpful for distinguishing high-score candidates.

Regression loss. As these multi-scale temporal windows have fixed widths, our candidate moments are restricted to discrete boundaries. To go beyond this limitation, we apply a boundary regression mechanism to adjust the temporal boundaries of high-score moments. Concretely, we fine-tune the localization offsets of high-score moments by a regression loss. First, we define a set ChC_{h} of high-score moments which IoUIoU scores are larger than a high-score threshold γ\gamma, then compute the start and end offset values for those high-score moments as follows:

where (s,e)(s,e) are the boundaries of the target moment, and (s^,e^)({\hat{s}},{\hat{e}}) are the boundaries of a high-score moment in ChC_{h}. Thus, the (δs,δe)(\delta_{s},\delta_{e}) denote its ground truth offsets and the predicted offsets (δ^s,δ^e)({\hat{\delta}}_{s},{\hat{\delta}}_{e}) are given by preceding fully connected layer. Next, we design the regression loss as follows:

where ChC_{h} is the set of high-score moments, NN is the size of ChC_{h} and RR represents the smooth L1 function.

With the alignment loss and regression loss, we eventually propose a multi-task loss to train the cross-modal interaction networks in an end-to-end manner, denoted by

where α\alpha is a hyper-parameter to control the balance of two losses.

During inference, we simply choose the candidate moment with the highest confidence score. If we need to select multiple moments (i.e. Top K), we first rank all candidates according to their confidence scores and adopt a non-maximum suppression (NMS) to select moments in order.

Experiments

We first introduce two public datasets for query-based moment retrieval.

ActivityCaption (Krishna et al., 2017): The ActivityCaption dataset is originally developed for the task of dense video caption, which contains 20,000 untrimmed videos and each video includes multiple natural language descriptions with temporal annotations. The video contents of this dataset are diverse and open. For query-based moment retrieval, each description is regarded as a query and corresponds to a target moment. Since the caption annotations of test data of ActivityCaption are not publically available, we take the val_1 as the validation set and val_2 as test data. The details of the ActivityCaption dataset are summarized in Table 1.

TACoS (Regneri et al., 2013): The TACoS dataset is developed onMPII Compositive (Rohrbach et al., 2012) and only contains 127 videos. But each video of TACoS has a large amount of temporal textual annotations. The contents of TACoS are limited to cooking scenes, thus lack the diversity. Moreover, the videos of TACoS are longer but the target moments are shorter than ActivityCaption, which make the query-based moment retrieval harder. The details of this dataset are also summarized in Table 1.

2. Implementation Details

In this section, we introduce some implementation details of our CMIN method, including the data preprocessing and model setting.

Data Preprocessing. We first resize every frame of videos to 112 ×\times 112 and extract the visual features by a pre-trained 3D-ConvNet (Tran et al., 2015). Specifically, we define continuous 16 frames as a unit and each unit overlaps 8 frames with adjacent units. We then input the units to the pre-trained 3D-ConvNet and obtain 4,096 dimension features for each unit. We next reduce the dimensionality of features from 4,096 to 500 using PCA, which is helpful for decreasing model parameters. These 500-d features are used as the frame features of our CMIN. Since some videos are overlong, we uniformly downsample their feature sequences to 200.

For natural language queries, we first extract the syntactic dependency graph using the library of NLTK (Bird and Loper, 2004) and employ the pre-trained Glove word2vec (Pennington et al., 2014) to extract the embedding features for each word token. The dimension of word features is 300.

Model Setting. In our CMIN, we sample kk candidate moments with multi-scale windows at each time step. Concretely, we set 77 window widths of $fortheActivityCaptiondatasetandfor the ActivityCaption dataset and4windowwidthsofwindow widths offorTACoS.Thus,wehave1,400samplesforeachvideoonActivityCaptionand800samplesonTACoS.Notethatwecutoffcandidateexamplesthatarebeyondtheboundariesofvideos.Wethensettheclearingthresholdfor TACoS. Thus, we have 1,400 samples for each video on ActivityCaption and 800 samples on TACoS. Note that we cut off candidate examples that are beyond the boundaries of videos. We then set the clearing threshold\lambdato0.3,thehigh−scorethresholdto 0.3, the high-score threshold\gammato0.7,andthebalancehyper−parameterto 0.7, and the balance hyper-parameter\alpha$ to 0.001. Moreover, the dimension of the hidden state of BiGRU networks is set to 512 (256 for one direction). The dimensions of the linear matrice in the multi-head self-attention and bilinear fusion are also set to 512. During training, we adopt an adam optimizer (Duchi et al., 2011) to minimize the multi-task loss and the learning rate is set to 0.001. We employ a mini-batch method and the batch size is 128.

3. Evaluation Criteria

To measure the retrieval performance of our CMIN and baselines, we adopt the “R@n, IoU=m” as evaluation criteria, which are proposed in (Gao et al., 2017). Concretely, we first calculate the IoU (i.e. Intersection over Union) between the selected moment and ground-truth moment, and the “R@n, IoU=m” means the percentage of at least one of top-n selected moments having IoU larger than m. The metric is on the query level, so the overall performance is the average among all the queries, denoted by R(n,m)=1Nq∑i=1Nqr(n,m,qi)R(n,m)=\frac{1}{N_{q}}\sum_{i=1}^{N_{q}}r(n,m,q_{i}), where the r(n,m,qi)r(n,m,q_{i}) represents whether one of the top-n selected moments of the query qiq_{i} has IoU>mIoU>m, and NqN_{q} is the total number of testing queries.

4. Performance Comparisons

We compare our proposed CMIN method with some existing state-of-the-art methods to verify the effectiveness.

MCN (Hendricks et al., 2017): The MCN method adopts a moment context network to integrate local and global moment features for query-based moment retrieval.

VSA-RNN and VSA-STV (Gao et al., 2017): The two methods are the extensions of the DVSA model (Karpathy and Fei-Fei, 2015). They both simply transforms the visual features of candidate moments and query features into a common space, and then estimate the correlation scores to select the most relevant one. The VSA-RNN applies a LSTM network to encode queries and VSA-STV adopts the off-the-shelf skip-thought (Kiros et al., 2015) feature extractor.

CTRL (Gao et al., 2017): The CTRL method proposes a cross-modal temporal regression localizer to estimate the alignment scores of candidate moments and textual queries by leveraging contextual contents of these moments, and then adjust the start and end boundaries of high-score moments.

AMRN (Liu et al., 2018a): The AMRN method emphasizes the visual moment features by attentive contextual contents and develops a cross-modal feature representation.

QSPN (Xu et al., 2019): The QSPN method introduces a multi-level model to integrate visual and textual features earlier with an attention mechanism, learn spatio-temporal visual representations and further re-generate queries as the auxiliary task.

The former three approaches only focus on the visual features within each moment and ignore the contextual information. And the latter three approaches incorporate the contextual evidence to improve the retrieval performance, where the AMRN utilizes the attention mechanism to filter irrelevant context and QSPN further develops an early interaction strategy for cross-modal features.

Table 2 and Table 3 show the overall performance evaluation results of our CMIN and all baselines on ActivityCaption and TACoS datasets, respectively. We choose the evaluation criteria “R@n, IoU=m” with n∈{1,5}n\in\{1,5\}, m∈{0.3,0.5,0.7}m\in\{0.3,0.5,0.7\} for ActivityCaption and n∈{1,5}n\in\{1,5\} , m∈{0.1,0.3,0.5}m\in\{0.1,0.3,0.5\} for TACoS. Note that we report the baseline performance based on either their original paper or our implementation by selecting the higher one. The experimental results reveal a number of interesting points:

While modeling moment features, the MCN applies a mean-pooling operation to aggregate all features of video sequences as context of the current moment, which may introduce noises into the moment representations and degrade the retrieval accuracy. Thus, the MCN achieves the worst performance on all criteria.

The context based methods CTRL, ACRN, QSPN and CMIN outperform the simple SVA-STV and VSA-RNN, which suggest the context modeling is crucial for high-quality moment retrieval. And the performance of the VSA-STA is slightly better than VSA-RNN, demonstrating the skip-thought feature extractor is helpful for query understanding.

The QSPN adopts an attention mechanism to fuse the visual and textual features, which develop the early interactions between moments and queries. The fact that the QSPN achieves better performance than CTRL and ACRN verifies the cross-modal interaction is critical for query-based moment retrieval.

On all the criteria of two datasets, the CMIN not only outperforms all previous state-of-the-art baselines, but also achieves tremendous improvements, especially on ActivityCaption. These results verify the effectiveness of our syntactic GCN, multi-head self-attention and multi-stage cross-modal interaction.

Moreover, we can find that the overall experimental results on TACoS are lower than ActivityCaption, and meanwhile, the CMIN can only achieve a smaller improvement on TACoS. That is because the videos are longer and the target moments are shorter on TACoS as shown in Table 1, which increase the number of potential candidate moments and make this task harder. Moreover, the invariant cooking scenes and shorter query description may also improve the retrieval difficulty. Since invariant scenes require better discrimination ability and the short query is hard to describe a moment clearly.

5. Ablation Study

To prove the contribution of each component of our CMIN method, we next conduct some ablation studies on the syntactic GCN, multi-head self-attention and multi-stage cross-modal interaction. Concretely, we discard one component at a time to generate an ablation model as follows.

CMIN(w/o. GCN): We first remove the syntactic GCN layer from the query representation learning and take the query representations from BiGRU networks as the input of multi-stage cross-modal interaction module.

CMIN(w/o. SA): We then discard the multi-head self-attention from the video representation learning to validate the importance of long-range semantic dependency modeling.

CMIN(w/o. CG): We next remove the cross gate in the multi-stage cross-modal interaction module, directly applying the bilinear fusion for frame representations and aggregation query representations.

CMIN(w/o. BF): We finally replace the low-rank bilinear fusion method with a simple concatenation for query and video features.

The ablation results on ActivityCaption and TACoS datasets are shown in Table 4 and Table 5, respectively. By analyzing the ablation results, we can find several interesting points:

The CMIN(full) outperforms all ablation models on both ActivityCaption and TACoS datasets, which demonstrates the syntactic GCN, multi-head self-attention, cross gate and low-rank bilinear fusion are all helpful for query-based moment retrieval.

The CMIN(w/o. GCN) achieves the worst performance on AcitivityCaption, and also have poor results on TACoS, indicating the utilization of syntactic structure is critical for query semantic understanding and subsequent modeling.

All ablation models still yield better results than all baselines. This fact demonstrates that our comprehensive retrieval framework is suitable for this task and the excellent performance does not only depend on a key component.

Moreover, for the syntactic GCN module, the number of stacked layers is a crucial hyper-parameter. Therefore, We further explore the effect of this hyper-parameter by varying the number of layers from 1 to 5. Figure 4 and Figure 5 shows the impact of layer number on ActivityCaption and TACoS datasets. Here we select “R@1,IoU=0.3” and “R@1,IoU=0.5” as evaluation criteria. From the tables, we note that the CMIN achieves the best performance while the number of layers is set to 2, and stacking too many or too few layers will both affect the performance of query-based moment retrieval. Because only one syntactic GCN layer cannot sufficiently leverage the syntactic dependencies of natural language queries and too many syntactic GCN layers will result in over-smoothing, that is, each word representation converges to the same value.

6. Qualitative Analysis

To qualitatively validate the effectiveness of the CMIN method, we display several typical examples of query-based moment retrieval. Figure 6 and Figure 7 show the retrieval results of the CMIN method and the best baseline QSPN on ActivityCaption and TACoS datasets, respectively. We can find that natural language queries are very diverse and often contain successive temporal actions. By intuitive comparison, the CMIN can retrieve more accurate boundaries of target moments than QSPN. Moreover, the retrieval precision on TACoS is lower than ActivityCaption, which is consistent with previous qualitative evaluations.

Furthermore, as the fundamental component of our multi-stage cross-modal interaction module, the video-to-query attentive aggregation builds a bridge between video and query information. Thus, we demonstrate how the attentive aggregation mechanism works to further understand the interaction process. As shown in Figure 8, the video-to-query attention results are visualized using a thermodynamic diagram, where the darker color means the higher correlation of the pair of frame and word representations. We note that each frame can attend the semantically related words and ignore these irrelevant words. For example, the word “line” has the highest attention score over the query for the fourth frame. This suggests the attentive aggregation strategy effectively establishes the relationship between visual and textual information, and is helpful for high-quality moment retrieval.

Conclusion

In this paper, we propose a novel cross-modal interaction network for query-based moment retrieval, which considers three critical factors of this task, including the syntactic structure of natural language queries, long-range semantic dependencies in video context and the fine-grained cross-modal interaction. Specifically, we advise a syntactic GCN to leverage the syntactic structure of queries for fine-grained representation learning, then propose a multi-head self-attention to capture long-range semantic dependencies from video context, and employ a multi-stage cross-modal interaction to explore the potential relations of video and query contents. The extensive experiments on ActivityCaption and TACoS datasets demonstrate the effectiveness of our proposed method.

References