Context-aware Biaffine Localizing Network for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, Yulai Xie

Introduction

Video understanding is a fundamental task in computer vision and has drawn increasing attention over the last years due to its various applications in video event detection , video summarization , video captioning and temporal action localization , etc. Recently, temporal sentence grounding (TSG) has been proposed as an important yet challenging task. This task requires automatically determining the start and end timestamps of a target segment in an untrimmed video that contains an activity semantically corresponding to a given sentence description, as shown in Figure 1. It is substantially more challenging as it needs to not only model the complex multi-modal interactions among vision and language features, but also capture complicated context information for their semantics alignment.

Most previous methods tackle TSG task following a multi-modal matching architecture, which generates multiple candidate proposals of different time intervals and ranks them according to their similarities with the sentence query. These methods severely rely on the quality of proposals, and break the intrinsic temporal structure and global context of videos. Recently, several works directly regress the temporal locations of the target segment. Specifically, they either regress the start/end timestamps based on the entire video representation , or predict at each frame to determine whether this frame is a start or end boundary . However, in these methods, start and end features are never jointly considered. Given a video with two target segments that have the same starting action (open the door) but end with different actions (go in, go out), predicting the start point independently may lead to timestamp confusion. Moreover, the ending point also tends to be inaccurate if it is predicted conditioned on the wrong start time.

Different from the aforementioned frameworks, we address the TSG from a new perspective: we reformulate this task by scoring all pairs of start and end indices simultaneously with a biaffine mechanism which interacts characteristics of each pair of possible start and end frames. This biaffine-based architecture is inspired by the dependency parsing task in natural language processing, in which the system predicts a dependency head for each child token and assigns a relation to the head-child pairs. However, there are two main obstacles when the biaffine mechanism used in dependency parsing is applied to TSG task. First, different from adjacent words in a sentence which carry different meaning, video is continuous and the adjacent frames naturally contain visually similar appearance. As shown in Figure 1, the adjacent frames near the segment boundary possess the similar semantics on “woman” and “pulls”. Thus, it is difficult to distinguish the specific boundary from adjacent frames without referring to these adjacent features as local context. Second, causalities between word pairs are usually indirect and can be far apart, but in videos, events in different intervals are directly correlated and rely on the whole contents to reason the precise semantics. For example, in Figure 1, without perceiving “a second time” from a global perspective, the first and second time of “pull on a rowing machine” possess similar semantic boundaries but totally different temporal indices. Therefore, the causal relations between the video events which act as global context are essential for understanding the segment.

Based on the above considerations, we propose a novel Context-aware Biaffine Localizing Network (CBLN), for temporal sentence grounding. Specifically, we develop a multi-context biaffine localization (MCBL) module which aggregates both local and global contexts to enrich the information of each frame representation. For each frame, we span the entire video features with different window sizes to get multi-scale video events as global contexts, and extract different numbers of adjacent frame features as multi-scale local contexts. The multi-scale local and global contexts are then inter-modulated to produce more adaptive contexts, which serve as the input representation for further biaffine localization by concatenating with frame-wise feature. At last, we obtain the output scores from the biaffine model to identify the similarities of all possible start-end pairs according to the semantics of sentence query. Besides, to provide fine-grained query-guided video representation for above biaffine localization, we also develop a multi-modal self attention (MMSA) module to sufficiently capture dependencies among video frames under the guidance of the sentence description. By jointly learning the overall model, our CBLN is able to localize query in video effectively.

Our main contributions are summarized as follows:

From a new perspective, we adopt biaffine mechanism to the TSG task. The biaffine-based architecture simultaneously scores all possible pairs of start and end frames for segment localization. Compared to previous methods, it gets rid of complicated proposal design and interacts both start-end timestamps effectively.

To alleviate the limitation of the biaffine localization, we further develop a multi-context biaffine localization module which utilizes multi-scale local and global contexts to enrich frame representations.

We conduct extensive experiments to validate the effectiveness of our proposed CBLN on three datasets (ActivityNet Captions, TACoS, and Charades-STA), and show that it significantly outperforms the state-of-the-arts by a large margin.

Related Work

Temporal Sentence Grounding. Temporal sentence grounding (TSG) is a new task introduced recently that aims to retrieve video segments using language queries. The early works employ a multi-modal matching architecture that first generates segment proposals, and then ranks them according to the similarity between proposals and the query to select the best matching one. Some of them propose to apply the sliding windows to generate proposals and subsequently integrate the query with segment representations via a matrix operation. Instead of using the sliding windows, latest works directly integrate sentence information with each fine-grained video clip unit, and predict the scores of candidate segments by gradually merging the fusion feature sequence over time. Although those methods achieve promising performances, they are severely limited by the quality of proposals.

To overcome above drawback, recent works directly regress the temporal locations of the target segment. Yuan et al. propose a co-attention based network to regress the start and end boundaries of the target segment. To improve the grounding with dense supervisions, Zeng et al. regress the distances from each frame to the target start (end) frame. Chen et al. propose a graph based bottom-up framework to capture multi-level semantics and encode the plentiful scene relationships.

Different from aforementioned two types of methods, we give a new solution to address the TSG problem. Specifically, we regard each frame as start or end frame to build all possible candidate segments and score all of them simultaneously with a biaffine mechanism . After that, we obtain the output scores for all segments and choose the best one corresponding to the highest value as the grounding result. Although 2D-TAN also scores all possible segments among the video, it directly max-pools the contained frame features to represent corresponding segment and lose discriminative frame-wise feature. Thus, compared to our detailed cross-modal interaction, this method fails to perform fine-grained interaction. Meanwhile, it captures the proposal-wise relations with convolution layers. Instead, our CBLN incorporates explicit local and global context to capture fine-grained frame-wise relations.

Biaffine based Dependency Parsing. Biaffine mechanism is widely used in dependency parsing which aims to build up a syntactic dependency tree for a given sentence. This task needs to capture all possible relations between the word pairs. Dozat et al. are the first work to learn the long-dependency from a head word to a modifier word with a relation label by proposing biaffine. They utilize biaffine operation as a scoring algorithm to determine the syntactic of a phrase from one word to another. Yu et al. further adapt biaffine to Named Entity Recognition by reformulating this task as the task of identifying start and end indices, as well as assigning a category to the span by these pairs. By treating input video as text passage, the biaffine mechanism is also applicable to TSG task in principle, because TSG aims to determine whether the segment from one frame to another is a full segment containing the activities described by the sentence query. However, the biaffine is not able to capture the adjacent contexts or correlate the global events, thus may lose local and global details about the scene meaning. In this paper, we enhance the biaffine mechanism by aggregating local-global information.

Proposed Method

Given an untrimmed video V\mathcal{V} and a sentence query Q\mathcal{Q}, we represent the video as V={vt}t=1T\mathcal{V}=\{v_{t}\}^{T}_{t=1} frame-by-frame, where vtv_{t} is the tt-th frame and TT is the number of total frames. Similarly, the query with NN words is denoted as Q={qn}n=1N\mathcal{Q}=\{q_{n}\}^{N}_{n=1} word-by-word. Temporal sentence grounding (TSG) aims to localize a segment (τs,τe)(\tau_{s},\tau_{e}) starting at timestamp τs\tau_{s} and ending at timestamp τe\tau_{e} in video V\mathcal{V}, which corresponds to the same semantic as query Q\mathcal{Q}.

The key to our Context-aware Biaffine Localizing Network (CBLN) is that we score all pairs of start and end frames simultaneously with a biaffine mechanism by interacting the characteristics of the start-end pairs. As shown in Figure 2, we first utilize two encoders to extract both video and query features, and then introduce a multi-modal self attention (MMSA) to generate the fine-grained video features for localization. Subsequently, we exploit biaffine localization module to score the start and end frame pairs of all possible segments, as shown in Figure 2 (a). In this way, we can get the scores for all candidate segments and choose the best one as the target segment. By enriching the query-guided video representation with local-global contexts, in Figure 2 (b), we further propose a multi-context biaffine localization (MCBL) module for more precise grounding.

2 Multi-Context Biaffine Localization

Biaffine mechanism is widely used in dependency parsing to assign scores to all possible spans in a sentence, where each span is defined by a pair of start and end words. As TSG can also be reformulated as the task of identifying the start and end frames of a specific segment among a given video, it is appropriate to adapt biaffine mechanism to this task for scoring all possible segment candidates. In our CBLN, before interacting the features of start and end frames using the biaffine mechanism, we first propose a multi-modal self attention (illustrated in Section 3.3) to generate fine-grained query-guided video representation F^={f^t}t=1T\widehat{\bm{F}}=\{\widehat{\bm{f}}_{t}\}_{t=1}^{T}, and then exploit a BiLSTM layer to further aggregate its sequential contexts as:

where F~={f~t}t=1T\widetilde{\bm{F}}=\{\widetilde{\bm{f}}_{t}\}_{t=1}^{T}. Then, as the contexts of the start and end frames are different, we apply two separate linear layers to generate separate hidden representations for each pair of start and end frames:

where Um\bm{U}^{m} and Wm\bm{W}^{m} are learnable parameters, bm\bm{b}^{m} is bias, and ⊕\oplus denotes element-wise addition, σ\sigma is the sigmoid function. After a sigmoid layer, Mp\bm{M}_{p} is the score of segment pp, which indicates the probability of pp matched to the query. Experiments in next section demonstrate that biaffine mechanism achieves a superior performance in TSG task.

2.2 Biaffine Localization with Local-Global Contexts

Although biaffine localization module has a strong ability to address TSG, it measures each segment by only considering the features of its start and end frames. As a result, the context of segments can only be drawn from a limited extent of previous BiLSTM layer. To enrich the context information of the start and end frames, we integrate start/end representation with: 1) Local contexts from adjacent frames. Since the adjacent frames near the segment boundary present similar visual appearance to the start/end frames, we aggregate the local contexts to distinguish them for more accurate boundaries grounding. 2) Global contexts from the entire video. We also extract long-range contexts to reason the temporal relations between different events among the video. Moreover, as shown in Figure 2 (b), the multi-scale local and global contexts are further aggregated. Finally, we concatenate each aggregated local-global contexts with corresponding start/end frame features for parallel multi-context biaffine localization.

Local-Global contexts. Given a start/end frame tt, we define two types of context features: “local” features Rtl\bm{R}^{l}_{t} extracted from adjacent frames, and “global” features Rtg\bm{R}^{g}_{t} spanned from the entire video. For “local” features, since the local contexts cover several adjacent frames around the frame tt, we directly utilize a window of size KlK^{l} on frame tt to extract features as follows:

For “global” features, as the global contexts refer to snippet-level features max-pooled from the long-range video frames, we define the spanning features Rtg\bm{R}^{g}_{t} as a feature bank of snippet features with a window of size KgK^{g} as:

Local-Global aggregation. Since both local and global contexts need to be adaptive to the boundary frame at each position, we modulate them to generate local-guided global and global-guided local contexts. In detail, we aggregate all scales of global features with each scale of local feature separately. In Figure 3, we demonstrate the details of multi-scale local-global aggregation. Given global feature of multiple scales and local feature of one specific scale, we first re-weight each local-global pair (Rtl,Rtg)(\bm{R}^{l}_{t},\bm{R}^{g}_{t}) by:

where NLBlock(⋅)\text{NLBlock}(\cdot) denotes a modified non-local block added with a layer normalization and dropout layer . The first NLBlock is utilized to capture temporal relations between the pooled events. Since feature (Rtg)′(\bm{R}^{g}_{t})^{\prime} loses specific information on the interested frames, another NLBlock is designed to further model adaptive local contexts by reasoning the adjacent features with global events. At last, we concatenate multi-level global/local contexts and project them through a linear layer to compute the final local-guided/global-guided global/local context for frame tt.

3 Multi-Modal Self Attention

Building interaction between video and query is a crucial step to provide detailed query-guided video representation for localization. To achieve this goal, recent works interact each word with each frame by a co-attention mechanism. However, these methods only focus on frame-wise cross-modal matching and lack the interaction over long-range video frames under the query guidance, which is essential for consecutive semantics understanding. To learn a better query-guided video representation forming the input of our biaffine-based localization, as depicted in Figure 4, we first concatenate word feature to each frame feature in a video and feed them into a multi-modal self attention module for long-range dependencies capturing. Specifically, given a word qnq_{n}, we first construct a joint multi-modal feature by concatenating single word feature qn\bm{q}_{n} to the whole video features V={vt}t=1T\bm{V}=\{\bm{v}_{t}\}^{T}_{t=1} as:

The multi-modal self attention module then takes the multi-modal features Fn\bm{F}_{n} as input, and produces a set of query, key and value pair by linear transformations as lntq=fntWq\bm{l}^{q}_{nt}=\bm{f}_{nt}\bm{W}^{q}, lntk=fntWk\bm{l}^{k}_{nt}=\bm{f}_{nt}\bm{W}^{k} and lntv=fntWv\bm{l}^{v}_{nt}=\bm{f}_{nt}\bm{W}^{v} at each frame tt, where Wq,Wk,Wv\bm{W}^{q},\bm{W}^{k},\bm{W}^{v} are parameters to be learned. We compute the multi-modal self attentive feature l^ntv\widehat{\bm{l}}^{v}_{nt}by:

α(nt,nt′)\alpha_{(nt,nt^{\prime})} is the weight coefficient computed by a softmax function, and takes into account of the correlation between (n,t)(n,t) and (n,t′)(n,t^{\prime}) which consists of same word but different frames. Next, we transform l^ntv\widehat{\bm{l}}^{v}_{nt} back to the same dimension as fnt\bm{f}_{nt} via a linear layer and add it element-wise with fnt\bm{f}_{nt} to form a residual connection :

4 Training Details

To train our CBLN, we utilize the scaled Intersection over Union (IoU) values as the supervision signal. Specifically, we compute the IoU score opo_{p} of each segment pp with the ground truth, and scale opo_{p} as the supervision signal by:

where Mp\bm{M}_{p} is the score of the segment pp in the output M\bm{M}.

Experiments

ActivityNet Captions. ActivityNet Captions contains 20000 untrimmed videos with 100000 descriptions from YouTube. The videos are 2 minutes on average, and the annotated video clips have much larger variation, ranging from several seconds to over 3 minutes. Following public split, we use 37417, 17505, and 17031 sentence-video pairs for training, validation, and testing respectively.

TACoS. TACoS is widely used on TSG task and contain 127 videos. The videos from TACoS are collected from cooking scenarios, thus lacking the diversity. They are around 7 minutes on average. We use the same split as , which includes 10146, 4589, 4083 query-segment pairs for training, validation and testing.

Charades-STA. Charades-STA is built on the Charades dataset , which focuses on indoor activities. In total, the video length on the Charades-STA dataset is 30 seconds on average, and there are 12408 and 3720 moment-query pairs in the training and testing sets, respectively.

Evaluation. Following previous works , we adopt “R@n, IoU=m” as our evaluation metrics. The “R@n, IoU=m” is defined as the percentage of at least one of top-n selected moments having IoU larger than m.

2 Implementation Details

For video encoding, we apply C3D to encode the videos on all three datasets, and also extract the I3D and VGG features on Charades-STA dataset. Since some videos are overlong, we set the length of video feature sequences to 200 for ActivityNet Captions and TACoS datasets, 64 for Charades-STA dataset, respectively. As for sentence encoding, we utilize Glove word2vec to embed each word to 300 dimension features. The hidden state dimensions of bi-directional GRU and BiLSTM are set to 512. We train our model with an Adam optimizer with leaning rate 8×10−48\times 10^{-4}, 3×10−43\times 10^{-4}, 4×10−44\times 10^{-4} for ActivityNet Captions, TACoS, and Charades-STA datasets, respectively. The batch size is set to 64. More model details can be found in our supplementary material.

3 Comparisons with state-of-the-arts

Comparisons on ActivityNet Captions. We compare our CBLN with the state-of-the-art methods on the ActivityNet Captions dataset in Table 1. We follow the previous methods to use C3D features for fair comparisons. Particularly, our model outperforms the previously best method DRN by 3.24% and 13.11% absolute improvement in terms of R@1, IoU=0.7 and R@5, IoU=0.7, respectively. Compared to the method 2D-TAN , we also outperform them by 6.89%, 3.61%, 1.06%, 3.38%, 2.19% and 1.45% in terms of all metrics, respectively.

Comparisons on TACoS. We compare our CBLN with the state of-the-art methods with the same C3D features in Table 2. On TACoS dataset, the cooking activities take place in the same kitchen scene with slightly varied cooking objects, thus showing the challenging nature of this dataset. Despite its difficulty, our model still reaches highest scores in terms of both R@1 and R@5 when IoU=0.5, and outperforms both 2D-TAN and DRN by a great margin.

Comparisons on Charades-STA. Table 3 reports the grounding results of various methods. Our CBLN reaches the highest results over all evaluation metrics. Specifically, when using the same VGG features, compared to the previously best method 2D-TAN , our model brings the absolute improvement of 3.86%, 1.19%, 9.06% and 5.34% on all metrics, respectively. For fair comparisons with GDP and LGI , we also perform experiments with same features (i.e., C3D and I3D) reported in their papers. It is obvious that our model still performs better. All these results again verify the effectiveness of our model.

4 Ablation Studies

In this section, we will perform in-depth ablation studies to evaluate the effect of each component in our CBLN on ActivityNet Captions dataset.

Main ablation studies. First of all, to investigate the effectiveness of biaffine localization module in this paper, we build up three baseline models with different grounding heads. For a fair comparison, we keep video/query encoders and cross-modal attention mechanism consistent in all baselines. Baseline (regression) directly regresses temporal boundary , while Baseline (proposal-match) designs pre-defined proposals to match the query . Table 4 shows that the Baseline* (biaffine) with our proposed biaffine mechanism significantly outperforms the other two baselines. It demonstrates that the biaffine mechanism is more suitable for segment localization in TSG task as it can learn detailed start-end frames interaction and get rid of handcrafted proposals compared to previous methods.

Next, to investigate the contribution of the proposed multi-modal self attention (MMSA) module and multi-context biaffine localization (MCBL) module, we also implement three variants of our model as shown in Table 5. Compared to the baseline*, MMSA captures more fine-grained query-video interactions and outperforms it by 1.48% and 3.77% in R@1, and R@5,IoU=0.7, respectively. Moreover, MCBL brings the highest improvement on both two metrics (i.e. 3.74% and 7.19%), which demonstrates its effectiveness of aggregating local-global contexts.

Analysis on local-global contexts generation. As shown in Table 6, we conduct the investigation on the impact of different local-global contexts modeling in multi-context biaffine localization (MCBL) module. First, for local context extraction, we need to preserve the details of each adjacent frame, thus the pooling strategy may lose some discriminative information contained in each frame features. The results also show that the concatenation performs better than both mean- and max-pooling. For global context spanning, it is difficult to work with all the raw features. Therefore, we select representative features by sub-sampling or pooling. In our experiments, we can find that max-pooling is superior to both random sampling and mean-pooling, as random sampling may lose discriminative frame feature and mean-pooling smooths away salient features that are otherwise preserved by max-pooling.

Besides, we also show the influence on different scales of both local KlK^{l} and global KgK^{g} window sizes. With more kinds of different scales, the model usually performs better than individual scale. For local scale KlK^{l}, we choose Kg={1,3,5}K^{g}=\{1,3,5\}. As for global scale KgK^{g}, the variant with four kinds of scales {1,2,4,8}\{1,2,4,8\} achieves the best result but only performs marginally better than the three-scales one {1,2,4}\{1,2,4\} at the expense of significantly larger cost of GPU memory. Thus, we choose Kg={1,2,4}K^{g}=\{1,2,4\} for global contexts in our experiments.

Analysis on local-global contexts aggregation. To interact information of both local and global contexts, we apply stacked non-local blocks to combine them in a learnable way in the local-global aggregation module of MCBL. As shown in Table 7, we can observe that when removing all non-local blocks and separately passing global and local contexts to latter computation, there is a performance drop of 2.55% and 2.53% in R@1, and R@5, IoU=0.7, respectively. When we replace non-local block with concatenation and a linear layer, there is also a drop of 1.73% and 1.67%. These results demonstrate that the stacked non-local blocks are effective for local-global contexts re-weighting. We also evaluate the performance when we replace the final linear layer (for multiple local/global combination, as shown in Figure 3) with max-pooling, it drops 0.74% and 0.63%.

Analysis on the fusion strategies. To better fuse the multiple information obtained by the outputs of both multi-modal self attention (MMSA) module and multi-context biaffine localization (MCBL) module, we compare different fusion operations as shown in Table 8. We can find that mean-pooling achieves the best performance in both two modules.

Analysis on model Complexity. To investigate the complexity of our model, we give an in-depth study in terms of speed and memory as shown in Table 9. Though our full model is not more efficient than CMIN, it outperforms CMIN with a large margin. Our w/o MCBL method still achieves better performance with similar computational cost (compared to CMIN). Compared to 2D-TAN, our full model performs better and much more efficient.

5 Qualitative Results

Figure 5 shows some qualitative results from three datasets. Our multi-context biaffine localization module can provide more contextual details about the segment, thus achieves better grounding results.

Conclusion

In this paper, we have proposed a novel context-aware biaffine localizing network, called CBLN, for temporal sentence grounding. The key to CBLN is that we reformulate this task from a new perspective for scoring all pairs of start and end indices simultaneously by a biaffine mechanism. To enrich the feature representation of each start/end frame, we additionally integrate them with multi-scale local and global contexts. A multi-modal self attention module is also developed to generate fine-grained query-guided video representation for such biaffine-based strategy. The experiments on three public datasets demonstrate the performance of CBLN, which brings significant improvements over the state-of-the-art methods.

Acknowledgements. This work was supported in part by the National Natural Science Foundation of China (No.61972448, No.61902347), and the Zhejiang Provincial Natural Science Foundation (No. LQ19F020002).

References