Unifying Event Detection and Captioning as Sequence Generation via Pre-Training
Qi Zhang, Yuqing Song, Qin Jin
Introduction
Dense Video Captioning (DVC) , as one of the important tasks in video understanding, aims to localize and describe multiple events in untrimmed videos. The early mainstream approaches normally decompose the DVC task into two sub-tasks, event detection and event captioning, and tackle the two sub-tasks separately. However, an obvious limitation of these methods is that they ignore the association of the two sub-tasks which can benefit from each other.
To address this limitation, some recent works explore to enhance the inter-task association between event detection and event captioning. For example, Deng et al. propose a ”top-down” framework which connects visual and language information to localize events or connects visual and event information to generate captions at different stages to enforce the interaction between the two sub-tasks. Wang et al. propose to share the same intermediate features and jointly optimize the two sub-tasks. Although the mutual promotion between the two sub-tasks has been witnessed, it is not trivial to design inter-task interactions for event detection and event captioning due to the large differences in their task specific solutions, which makes them hard to fully benefit from each other.
Besides, the previous methods detect the event boundaries solely based on the video information or the hybrid knowledge of video and text, which ignores the temporal relationships between multiple events. It easily leads to redundant event detection, either producing a large number of event candidates, or having excessive overlap between events. As shown in Figure 1, due to the lack of consideration of the temporal relationship between events, the traditional event detection methods usually detect redundant events with a high degree of overlap. In fact, there are strong temporal dependencies between events in the video. For example, in the instructional videos, events usually occur one after another, explaining each operation step.
To address the above mentioned two limitations, in this paper, we propose to define event detection as a sequence generation task, which generates the event sequence one by one conditioned on previously detected events. It can fully exploit the previous events as context to avoid generating redundant events, and unify the event detection and event captioning sub-tasks into the same framework to explore inter-task interactions in a simpler but highly effective way. Specifically, we define the event as a new modality and propose a unified video-event-text pre-training and fine-tuning framework based on the transformer architecture for both the event detection and the event captioning. We employ two pre-training tasks including Masked Language Modeling (MLM) and Masked Video Feature Regression (MVFR) to learn the video-text representation for event captioning, and propose a novel pre-training task called Masked Event Feature Modeling (MEFM) to learn the video-event representation for event detection. Since the two sub-tasks share the same model architecture and parameters, we can alternately train the unified model with the three pre-training tasks, which is a simple, natural and effective method to strengthen the association between two sub-tasks. Experimental results on the ActivityNet Captions dataset show that even pre-trained on the same dataset without any extra data, our model significantly outperforms the state-of-the-art methods on both event detection and event captioning. In particular, benefiting from the new event detection framework, our model generates more diverse events similar to human annotations. When pre-training the model on extra out-of-domain captioning data, our model achieves more additional gains, which demonstrates the effectiveness of our pre-training and fine-tuning framework.
The main contributions of this work are three-fold:
We transform the event detection sub-task into a sequence generation problem with temporal dependencies modeling between events, which makes our model directly generate an appropriate and low redundancy sequence of events for untrimmed videos.
With the task format unification, we propose a unified video-event-text pre-training model with three pre-training tasks to effectively enhance the inter-task association between event detection and event captioning.
Our model achieves new state-of-the-art dense video captioning results on the ActivityNet Captions dataset, and can be further promoted by pre-training on more out-of-domain video caption data.
Related Works
Dense Video Captioning. The dense video captioning (DVC) task is first proposed by Krishna et al. to localize and describe rich events in long untrimmed videos. A two-stage framework is proposed, which first produces a large amount of event candidates via an event detection module, and then generates descriptions for each event. Some following works separately improve the performance of two modules by producing less redundant event proposals or introducing multi-modal features to enrich event representation . However, handling event detection and captioning independently without any association has obvious limitations. Therefore, some other following works focus on enhancing the inter-task association between event detection module and event captioning module to improve the performance of DVC. Specifically, Li et al. adopt a new descriptiveness regression component to build a bridge between event detection and captioning modules, which measures the descriptive complexity of each event proposal, and adjusts the event proposal boundaries. Deng et al. reverse the predominant “detect-then-describe” fashion and propose a three-stage top-down framework, which first generates coarse-grained sentences, then aligns each sentence with event fragments, and finally refines the caption quality by a Dual-Path Cross Attention module. Wang et al. extend the “Object Query” manner in DETR to the DVC task and decode the intermediate features and event query to produce event proposal and description simultaneously.
Although promising DVC results have been achieved in these methods, they fail to explore the temporal relationship between events for event detection, which leads to highly redundant event proposals. Furthermore, the “propose-then-select” detection manner based on a large amount of event candidates is also computationally complex. In this work, we convert the event detection task into a sequence generation problem, and unify the event detection and captioning into a unified framework, which makes the inter-task association more natural and effective.
Pre-training for V+L Tasks. The pre-training and fine-tuning paradigm has demonstrated strong potentials in V+L cross-modal tasks recently. Such as the Oscar model, first pre-trained on large-scale image-text pairs to learn a joint representation for vision and language, and then fine-tuned on several downstream tasks, achieves the state-of-the-art results on both vision-and-language understanding tasks (e.g., image-text retrieval , visual question answering ) and generation tasks (e.g., image captioning , novel object captioning ). However, due to the coupling complexity of event detection and captioning, none of them have verified the effectiveness of multi-modal pre-training on the dense video captioning task.
Method
Our unified model for dense video captioning following the pre-training and fine-tuning paradigm is illustrated in Figure 2. We unify the event detection and event captioning in one framework and treat the video, events and captions as three independent modalities. Three pre-training tasks are proposed for the cross-modal representation learning, including Masked Language Modeling (MLM), Masked Video Feature Regression (MVFR), and Masked Event Feature Modeling (MEFM). We first alternately pre-train the model with the three pre-training tasks and then fine-tune it for event detection and event captioning respectively.
Video Representation. Given an input video, we follow previous works to use C3D and TSN to extract the raw video features in order to have a fair comparison in experimental evaluations. In addition, we also introduce semantic concept features (called CPT) to enrich the visual representation. Specifically, we employ a bidirectional LSTM as a multi-class multi-label video classifier to predict concept labels for each frame, which are the nouns and verbs extracted from corresponding annotated captions. To enable the model to predict distinguishable fine-grained concept labels for different frames, we optimize the model with an additional event boundary prediction objective, which requires the model to predict whether the current frame is the start, middle or the end frame of an event. The hidden state from the last LSTM layer is used as the concept feature for each video frame. Finally, we concatenate the concept feature with visual appearance features from C3D or TSN to represent the video as a sequence of frame-level features = , where is the number of frames in the video. We employ a fully-connected layer to embed the video feature sequence into the same dimensionality as the word embedding, and add a learnable modality embedding to represent the video modality.
Event Representation. For an event with timestamp in the video, we convert it into a -dimensional binary feature vector, which is expressed as , where is the total number of video frames. The value is set as 1 if the -th frame is included in the event interval, otherwise it is set as 0. In this way, we represent all the events in the video as a sequence of , where is the number of events in the current video. Similar to the video representation, we employ a fully-connected layer to embed the event sequence into the same dimensionality as the video and text embeddings, and add a learnable modality embedding to represent the event modality. Furthermore, we add positional embedding to reserve the temporal order information of events in the video.
Text Representation. For the caption text, we represent each caption with a sequence of word embeddings , where is the total number of words. We further add the positional embedding to keep the sequential information, and add a learnable modality embedding to represent the text modality.
Multimodal Transformer. Our model is based on a multi-stream architecture as illustrated in Figure 2, where three independent transformer encoders are first applied on each modality for the intra-modality learning, and then a cross transformer encoder is employed to capture the inter-context information across different modalities. We define the number of cross encoder layers as and the number of layers in the three single-modal encoders including the modality of video, event and text as , and respectively. The hidden size of all the transformer layers is denoted as , and the number of self-attention heads is denoted as .
2 Proxy Tasks
Masked Language Modeling (MLM). To enable our model the ability for caption generation, we adopt MLM as one of the pre-training tasks like other Vision-and-Language (V+L) pre-training models . The MLM task takes , and as inputs, and predict the masked words in according to the corresponding video content of the current event timestamps . Since all the events in are input to the model while the caption is only corresponding to the current event, we restrict the attentions for other events , so that the model can focus on the current event captioning. Specifically, we mask the cross-modal attention weights in the cross encoder for other events to make the caption generation ignores other events. Similar to the BERT , we randomly mask out the words in with a probability of 15%, and replace the masked ones with a special token [MASK] 80% of the time, with another random word 10% of the time and the original word 10% of the time. The goal of this task is to predict the masked words based on the context information from the whole video content, the current event boundary , and the surrounding captioning words by minimizing the negative log-likelihood as follows:
where denotes all trainable parameters, denotes the whole training set, and denotes the masked words in .
Masked Video Feature Regression (MVFR). In contrast to the MLM task which predicts the masked caption words according to the video content, we also introduce the MVFR task to reconstruct video features based on the description. The input for MVFR task is exactly the same as the MLM task. Suppose is the description for the event , we randomly choose some video frames in for prediction. Specifically, we randomly mask 15% of the features from , where and are the start and end video frame in respectively. Each masked feature is replaced by a special feature vector [MASK], which is an all-zeros vector with the same dimensionality as the original video feature . The hidden state of from the cross transformer encoder is input to a FC layer to predict the original video feature denoted as , according to the remaining video features and the event description . We adopt the L2 regression loss to reduce the distance between and as follows:
where is the dimension of video features. We share the FC prediction layer with the video feature embedding layer.
Masked Event Feature Modeling (MEFM). To predict event boundaries in a sequence generation manner, we propose a new pre-training task called Masked Event Feature Modeling (MEFM). Unlike above two pre-training tasks, the MEFM task takes , as inputs. We randomly mask event embeddings in with a probability of 15%, and replace the masked one with a special feature vector [MASK], which is an all-zeros vector. The model is required to predict the event boundaries of according to the surrounding events in and the video content. We apply a FC layer on the output of cross transformer encoder as a -dimensional binary classifier, where is the length of video frame sequences. The training loss can be defined as follows:
where indicates whether the i-th video frame is included in the masked event interval, and denotes the predicted probability.
3 Pre-Training and Fine-Turning for DVC Task
As shown in Figure 2, we pre-train the model for event detection and event captioning in a unified framework. However, the pre-training tasks of MLM and MVFR for event captioning take three modalities (video, event, text) as input, while the MEFM task for event detection takes two modalities without the text modality as input, making it problematic to optimize the three objectives jointly. Therefore, we train the two sub-tasks alternatively. Specifically, we divide the batch data into two categories: the batch with three modalities input (for MLM, MVFR) named as , and the batch with two modalities input (for MEFM) named as . We introduce a hyper-parameter to represent the probability of choosing batch , and thus (-) denotes the probability of choosing batch . When the batch is fed to the model, the training loss is defined as , while when batch is fed to the model, the loss is defined as .
After pre-training, we fine-tune the model for the event detection sub-task (called ED) and event captioning sub-task (called EC) respectively. The two downstream tasks are similar since they both follow the auto-regressive manner for generation. Therefore, we fine-tune the model in a similar way for the two sub-tasks. Specifically, we adapt our bi-directional pre-trained model to a uni-directional generator by constraining the self-attention mask of the text/event sequence to avoid seeing future items. Similar to the MLM and MEFM pre-training tasks, we randomly choose 15% of word/event features and replace them with the [MASK] token/special all-zeros vector for prediction. Note that the event detection task is to generate event sequences with the whole video as input, while the event captioning task generates captions for the corresponding event according to the whole video and event embedding . Therefore, the fine-tuning objective for ED and EC tasks can be expressed as follows:
In the inference phase, we follow the “detect-then-describe” pipeline. At the stage of event prediction, we first input the whole video frame sequence and a special “start event” vector to the model. Then, we start to generate the event sequence one by one via feeding a [MASK] vector and sampling the predicted event feature from the -dimensional binary classification layer. The predicted event feature vector is then used to replace the previous [MASK] vector, and a new [MASK] vector is fed to the model for the next event generation until the special “end event” vector is predicted. Finally, following the rule that the first 1-value in the event vector is regarded as the start time of an event in the video, and the last 1-value in the event vector is regarded as the end time of the event in the video, we translate the event vector into the event timestamp format. After predicting the event sequence for the untrimmed video, we generate corresponding captions for each predicted event. We first input the whole video, the current event embedding and the start [SOS] token to the model. Then, we follow the same auto-regressive sequence generation process to generate the captions until the end [EOS] token is predicted.
Experiments
Dataset. We conduct experiments on the benchmark ActivityNet Captions dataset , which contains 19,994 videos with an average of 3.65 event proposals per video. We follow the official split with 10009/4925/5044 videos for training, validation, and test. Furthermore, to demonstrate the ability of our model to benefit from more out-of-domain captioning data, we further evaluate our model pre-trained with other non-dense video captioning datasets, including MSRVTT , TGIF and VATEX datasets.
Implementation details. For fair comparisons with the state-of-the-art methods, we use the same video features as PDVC , including the C3D features provided by PDVC and TSN features provided by MT . The max length of video frames is set as 100. We set the layer number of independent encoders , the layer number of cross encoder , the hidden size , and the head number . When pre-training the model for event detection, the hyper-parameter for choosing different batches is set as 1/3 on both C3D and TSN features. When pre-training the model for event captioning, it is set as 1/2 on C3D features and 3/4 on TSN features. We use Adam as the optimizer and train all the model parameters from scratch.
Metrics. We evaluate our method from three aspects. (1) To evaluate the performance of event detection, we use the evaluation tool provided by the 2018 ActivityNet Captions Challengehttps://github.com/ranjaykrishna/densevid_eval (called Evaluator2018), which computes the average precision and average recall against the ground-truth events across temporal IoU (tIoU) at [0.3, 0.5, 0.7, 0.9]. Moreover, we also compute the self-tIoU between the detected events to evaluate the event diversity. (2) To purely evaluate the event captioning ability, we report the captioning results based on the ground-truth events with classic captioning metrics, including BLEU , METEOR and CIDEr . We also use the Evaluator2018 to compute these metrics and report the results with tIoU threshold of 0.9. (3) To evaluate the performance of dense video captioning, the captioning performance based on the generated events, we use SODAhttps://github.com/fujiso/SODA as the evaluation metric, which is a new evaluation metric proposed for dense video captioning to overcome some of the limitations of previous metrics.
To verify that the general video captioning metrics are not appropriate enough for the dense video captioning evaluation, we carefully design an experiment to compare the classic evaluation metrics provided in Evaluator2018 with the newly proposed SODA metric. We first compute the scores of captions for events in different intervals of the video, and observe that the first event of the video usually gets higher captioning scores than other locations. Based on such observation, we propose four simple operations, including Increase, Reduce, Exchange and Extreme, to modify our dense video captioning results, and then show the variations of scores computed by different evaluation metrics. Given a submission file with the best results of our model, the Increase operation randomly copies the first event and its corresponding caption in the submission file with 40% probability for each video. The Reduce operation randomly removes the -th (where ) event and its corresponding caption in the submission file with 15% probability. The Exchange operation is the combination of Increase and Reduce operations, and the Extreme operation removes the -th (where ) event and its corresponding caption with 100% probability. We run the four operations with different random seeds for three times and report the average results in Table 1. Although we intuitively expect that the above four operations should adversely affect the dense video captioning results, the scores evaluated by the Evaluator2018 surprisingly increase significantly from 7.33 to 8.52 on the METEOR metric. On the contrary, SODA correctly reflects the captioning quality with the score constantly decreased. It is because the Evaluator2018 fails to penalize redundant and non-recalled events. Therefore, simply repeating the first event caption with 40% probability can significantly improve the BLEU@4 and METEOR scores by 8% and 2% respectively, although the results actually do not have any substantial improvements. Therefore, we consider SODA as the main evaluation metric for dense video captioning. In this work, we use two versions of SODA to evaluate our DVC performance. The SODAold is commonly used in previous methods, which computes the score based on two references independently and reports the averaged score. The SODAmr however computes the score based on multiple references simultaneously, which are more accurate.
2 Comparison with State-of-the-art Methods
We compare our model with both types of baseline methods, including the ones that tackle event detection and captioning separately, such as MFT , SDVC and ECHR , and the state-of-the-art models that exploit interactions between the two sub-tasks, such as DCE , TDA-CG , DVC , MT , PDVC and SGR .
Event detection results. As shown in Table 2, our proposed event sequence generation model outperforms previous event detection methods by a large margin. Although pre-trained on the same dataset without any extra data, our model achieves significant improvements over previous best results at most of tIoU thresholds, with the average recall score improved from 55.58% to 59.00% and the average precision score improved from 58.07% to 60.32% . Furthermore, our model generates 2.94 events per video on average with the self-tIoU of 0.02, which shows a greater improvement on reducing the event redundancy than traditional event detection methods. It is also much closer to the ground-truth events which have an average self-tIoU of 0.05. In addition to achieving much better results on the event detection accuracy and diversity, our model is also more efficient than previous methods. The PDVC model needs to train an additional predictor for event prediction and SDVC first adopts an extra model (SST ) to extract 1000 event proposals and obtains M candidate proposals with Non-Maximum Suppression. Then, SDVC selects the final proposals from the M candidates in an auto-regressive manner. On the contrary, our model directly generates an appropriate number of events from the raw input video features, which is one-stage without error accumulation and with much less computational cost.
Event captioning results. Table 3 reports the dense video captioning results of different models on the ActivityNet Captions validation set. To have a fair comparison with the state-of-the-art methods, we train our model with C3D and TSN visual features respectively. When inferring based on the ground-truth event proposals, our model trained by C3D features improves the BLUE@4 from 2.64 to 2.67, the METEOR from 10.58 to 11.01, and the CIDEr from 47.26 to 52.42. Similar improvement trends are obtained with the TSN features on METEOR and CIDEr as well. It demonstrates the effectiveness of our model for event captioning. When inferring based on the generated proposals, we follow the “detect-then-describe” pipeline, which first predicts the event sequence and then describes each event clip. We evaluate our model with the more appropriate evaluation metric SODA for dense video captioning (SODA ), and report two scores, including SODAold and SODAmr. As shown in Table 3, our model with C3D features achieves the best SODAold score and SODAmr score, outperforming all previous methods. The performance of our model with more advanced TSN features surpasses all previous methods as well.
3 Ablation Studies
In this section, we ablate our model to demonstrate the effectiveness of different components for event captioning and event detection respectively.
Table 4 shows the ablation results of our model for the event captioning. To purely analyse the caption generation qualities, we use the ground-truth video events to avoid the impact of event detection quality. As shown in Table 4, the model pre-trained only with MLM task (Row 2) has outperformed the non-pretrain baseline (Row 1) even without any extra data for pre-training. When combining with other pre-training tasks, including MVFR and MVFR+MEFM (Row 3 and 4), our model achieves significant improvements on the captioning metrics, which demonstrates the effectiveness of the proposed pre-training tasks. Specifically, the improvement brought by MEFM task is the most significant, which improves the BLEU@4 from 2.38 to 2.59, the METEOR from 10.48 to 10.94 and the CIDEr from 51.08 to 52.13. It shows that the event captioning task can benefit from the event detection task, and demonstrates the advantages of our proposed unified framework for the inter-task interaction.
The same trend can also be found in the Table 5. The pre-trained model significantly outperforms the non-pretrain baseline in Row 1. Adding the MLM and MVFR tasks in the pre-training stage greatly improves the event detection results, with the average recall improved from 56.49 to 58.32 and the average precision improved from 58.12 to 59.95. It shows that the event detection task can also benefit from the event captioning task. Besides different pre-training tasks, enhancing the video representation with semantic concept features also brings additional gains for both the event captioning and event detection tasks, as shown in Row 5 vs. Row 4 of the Table 4 and Table 5. The surprising performance improvement by the semantic concept features (short for CPT) in event detection (improved from 58.32 to 59.00 on average recall and 59.95 to 60.32 on average precision) further demonstrates that comprehensive semantic understanding of video frames can greatly help event boundary detection. This is also the reason why the event detection and event captioning can help each other, because they are both based on the semantic understanding of videos. Finally, combining all the components together in our model achieves the best performance for both event detection and event captioning (Row 5).
We also ablate the influence of the hyper-parameter in our model in Figure 3. We observe that enhancing inter-task association in different degrees all outperform the model without inter-task interaction (where ) for both the event captioning and event detection. With gradually increasing from 1/2 to 1, the METEOR score continues to decline, while the CIDEr and BLUE@4 scores fluctuate slightly. We observe that modulating hyper-parameter to 1/2 obtains the best performance for event captioning. While when pre-training for event detection, the best performance is achieved when hyper-parameter is set as 1/3.
4 Pre-training with Out-of-domain Data
Due to the advantage of pre-training and fine-tuning framework, our model can benefit from more out-of-domain data besides the ActivityNet dataset. Specifically, we pre-train the model with conventional non-dense video captioning datasets including MSRVTT , TGIF and VATEX , which results in 676K additional video-caption pairs in total. Since the conventional video captioning datasets do not contain event annotations, we pre-train the model only with the MLM and MVFR tasks on these out-of-domain data, and fine-tune it on the ActivityNet Captions training set to adapt to the dense video captioning task. Figure 4 shows the results of our model pre-trained with extra data. Compared with the model pre-trained only on the ActivityNet dataset, the additional out-of-domain data brings significant additional gains on all the captioning metrics, e.g., improving BLEU@4, METEOR and CIDEr by more than 0.4, 0.8 and 4 points respectively. It demonstrates the merit of our unified pre-training based model framework, which enables the dense video captioning task to benefit from conventional non-dense video captioning datasets.
Conclusion
In this work, we define the event detection task as a sequence generation problem to fully exploit the temporal relationship between events for more accurate and diverse event detection. Benefiting from the unification of event detection and event captioning sub-tasks, we propose a unified dense video captioning model based on pre-training and fine-tuning framework. We design a new “event” modality and propose three pre-training tasks to interact the event detection and event captioning sub-tasks naturally. Experimental results on the ActivityNet Captions dataset show that our model significantly outperforms the state-of-the-art methods on both event detection and event captioning, and achieves additional gains when leveraging more out-of-domain data.
Acknowledgement. This work was partially supported by National Key R&D Program of China (No. 2020AAA0108600) and National Natural Science Foundation of China (No. 62072462).