Temporal Action Detection with Structured Segment Networks

Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, Dahua Lin

Introduction

Temporal action detection has drawn increasing attention from the research community, owing to its numerous potential applications in surveillance, video analytics, and other areas . This task is to detect human action instances from untrimmed, and possibly very long videos. Compared to action recognition, it is substantially more challenging, as it is expected to output not only the action category, but also the precise starting and ending time points.

Over the past several years, the advances in convolutional neural networks have led to remarkable progress in video analysis. Notably, the accuracy of action recognition has been significantly improved . Yet, the performances of action detection methods remain unsatisfactory . For existing approaches, one major challenge in precise temporal localization is the large number of incomplete action fragments in the proposed temporal regions. Traditional snippet based classifiers rely on discriminative snippets of actions, which would also exist in these incomplete proposals. This makes them very hard to distinguish from valid detections (see Fig. 1). We argue that tackling this challenge requires the capability of temporal structure analysis, or in other words, the ability to identify different stages e.g. starting, course, and ending, which together decide the completeness of an actions instance.

Structural analysis is not new in computer vision. It has been well studied in various tasks, e.g. image segmentation , scene understanding , and human pose estimation . Take the most related object detection for example, in deformable part based models (DPM) , the modeling of the spatial configurations among parts is crucial. Even with the strong expressive power of convolutional networks , explicitly modeling spatial structures, in the form of spatial pyramids , remains an effective way to achieve improved performance, as demonstrated in a number of state-of-the-art object detection frameworks, e.g. Fast R-CNN and region-based FCN .

In the context of video understanding, although temporal structures have played an crucial role in action recognition , their modeling in temporal action detection was not as common and successful. Snippet based methods often process individual snippets independently without considering the temporal structures among them. Later works attempt to incorporate temporal structures, but are often limited to analyzing short clips. S-CNN models the temporal structures via the 3D convolution, but its capability is restricted by the underlying architecture , which is designed to accommodate only 1616 frames. The methods based on recurrent networks rely on dense snippet sampling and thus are confronted with serious computational challenges when modeling long-term structures. Overall, existing works are limited in two key aspects. First, the tremendous amount of visual data in videos restricts their capability of modeling long-term dependencies in an end-to-end manner. Also, they neither provide explicit modeling of different stages in an activity (e.g. starting and ending) nor offer a mechanism to assess the completeness, which, as mentioned, is crucial for accurate action detection.

In this work, we aim to move beyond these limitations and develop an effective technique for temporal action detection. Specifically, we adopt the proven paradigm of “proposal+classification”, but take a significant step forward by utilizing explicit structural modeling in the temporal dimension. In our model, each complete activity instance is considered as a composition of three major stages, namely starting, course, and ending. We introduce structured temporal pyramid pooling to produce a global representation of the entire proposal. Then we introduce a decomposed discriminative model to jointly classify action categories and determine completeness of the proposals, which work collectively to output only complete action instances. These components are integrated into a unified network, called structured segment network (SSN). We adopt the sparse snippet sampling strategy , which overcomes the computational issue for long-term modeling and enables efficient end-to-end training of SSN. Additionally, we propose to use multi-scale grouping upon the temporal actionness signal to generate action proposals, achieving higher temporal recall with less proposals to further boost the detection performance.

The proposed SSN framework excels in the following aspects: 1) It provides an effective mechanism to model the temporal structures of activities, and thus the capability of discriminating between complete and incomplete proposals. 2) It can be efficiently learned in an end-to-end fashion (55 to 1515 hours over a large video dataset, e.g. ActivityNet), and once trained, can perform fast inference of temporal structures. 3) The method achieves superior detection performance on standard benchmark datasets, establishing new state-of-the-art for temporal action detection.

Related Work

Action Recognition. Action recognition has been extensively studied in the past few years . Earlier methods are mostly based on hand-crafted visual features . In the past several years, the wide adoption of convolutional networks (CNNs) has resulted in remarkable performance gain. CNNs are first introduced to this task in . Later, two-stream architectures and 3D-CNN are proposed to incorporate both appearance and motion features. These methods are primarily frame-based and snippet-based, with simple schemes to aggregate results. There are also efforts that explore long-range temporal structures via temporal pooling or RNNs . However, most methods assume well-trimmed videos, where the action of interest lasts for nearly the entire duration. Hence, they don’t need to consider the issue of localizing the action instances.

Object Detection. Our action detection framework is closely related to object detection frameworks in spatial images, where detection is performed by classifying object proposals into foreground classes and a background class. Traditional object proposal methods rely on dense sliding windows and bottom-up methods that exploit low-level boundary cues . Recent proposal methods based on deep neural networks show better average recall while requiring less candidates . Deep models also introduce great modeling capacity for capturing object appearances. With strong visual features, spatial structural modeling remains a key component for detection. In particular, the RoI pooling is introduced to model the spatial configuration of object with minimal extra cost. The idea is further reflected in R-FCN where the spatial configuration is handled with the position sensitive pooling.

Temporal Action Detection. Previous works on activity detection mainly use sliding windows as candidates and focus on designing hand-crafted feature representations for classification . Recent works incorporate deep networks into the detection frameworks and obtain improved performance . S-CNN proposes a multi-stage CNN which boosts accuracy via a localization network. However, S-CNN relies on C3D as the feature extractor, which is initially designed for snippet-wise action classification. Extending it to detection with possibly long action proposals needs enforcing an undesired large temporal kernel stride. Another work uses Recurrent Neural Network (RNN) to learn a glimpse policy for predicting the starting and ending points of an action. Such sequential prediction is often time-consuming for processing long videos and it does not support joint training of the underlying feature extraction CNN. Our method differs from these approaches in that it explicitly models the action structure via structural temporal pyramid pooling. By using sparse sampling, we further enable efficient end-to-end training. Note there are also works on spatial-temporal detection and temporal video segmentation , which are beyond the scope of this paper.

Structured Segment Network

The proposed structured segment network framework, as shown in Figure 2, takes as input a video and a set of temporal action proposals. It outputs a set of predicted activity instances each associated with a category label and a temporal range (delimited by a starting point and an ending point). From the input to the output, it takes three key steps. First, the framework relies on a proposal method to produce a set of temporal proposals of varying durations, where each proposal comes with a starting and an ending time. The proposal methods will be discussed in detail in Section 5. Our framework considers each proposal as a composition of three consecutive stages, starting, course, and ending, which respectively capture how the action starts, proceeds, and ends. Thus upon each proposal, structured temporal pyramid pooling (STPP) are performed by 1) splitting the proposal into the three stages; 2) building temporal pyramidal representation for each stage; 3) building global representation for the whole proposal by concatenating stage-level representations. Finally, two classifiers respectively for recognizing the activity category and assessing the completeness will be applied on the representation obtained by STPP and their predictions will be combined, resulting in a subset of complete instances tagged with category labels. Other proposals, which are considered as either belonging to background or incomplete, will be filtered out. All the components outlined above are integrated into a unified network, which will be trained in an end-to-end way. For training, we adopt the sparse snippet sampling strategy to approximate the temporal pyramid on dense samples. By exploiting the redundancy among video snippets, this strategy can substantially reduce the computational cost, thus allowing the crucial modeling of long-term temporal structures.

At the input level, a video can be represented as a sequence of TT snippets, denoted as (St)t=1T(S_{t})_{t=1}^{T}. Here, one snippet contains several consecutive frames, which, as a whole, is characterized by a combination of RGB images and an optical flow stack . Consider a given set of NN proposals P={pi=[si,ei]}i=1NP=\{p_{i}=[s_{i},e_{i}]\}_{i=1}^{N}. Each proposal pip_{i} is composed of a starting time sis_{i} and an ending time eie_{i}. The duration of pip_{i} is thus di=ei−sid_{i}=e_{i}-s_{i}. To allow structural analysis and particularly to determine whether a proposal captures a complete instance, we need to put it in a context. Hence, we augment each proposal pip_{i} into pi′=[si′,ei′]p^{\prime}_{i}=[s^{\prime}_{i},e^{\prime}_{i}] with where si′=si−di/2s^{\prime}_{i}=s_{i}-d_{i}/2 and ei′=ei+di/2e^{\prime}_{i}=e_{i}+d_{i}/2. In other words, the augmented proposal pi′p^{\prime}_{i} doubles the span of pip_{i} by extending beyond the starting and ending points, respectively by di/2d_{i}/2. If a proposal accurately aligns well with a groundtruth instance, the augmented proposal will capture not only the inherent process of the activity, but also how it starts and ends. Following the three-stage notion, we divide the augmented proposal pi′p^{\prime}_{i} into three consecutive intervals: pis=[si′,si]p_{i}^{s}=[s^{\prime}_{i},s_{i}], pic=[si,ei]p_{i}^{c}=[s_{i},e_{i}], and pie=[ei,ei′]p_{i}^{e}=[e_{i},e^{\prime}_{i}], which are respectively corresponding to the starting, course, and ending stages.

2 Structured Temporal Pyramid Pooling

As mentioned, the structured segment network framework derives a global representation for each proposal via temporal pyramid pooling. This design is inspired by the success of spatial pyramid pooling in object recognition and scene classification. Specifically, given an augmented proposal pi′p^{\prime}_{i} divided into three stages pisp_{i}^{s}, picp_{i}^{c}, and piep_{i}^{e}, we first compute the stage-wise feature vectors fis\mathbf{f}_{i}^{s}, fic\mathbf{f}_{i}^{c}, and fie\mathbf{f}_{i}^{e} respectively via temporal pyramid pooling, and then concatenate them into a global representation.

Specifically, a stage with interval [s,e][s,e] would cover a series of snippets, denoted as {St∣s≤t≤e}\{S_{t}|s\leq t\leq e\}. For each snippet, we can obtain a feature vector vt\mathbf{v}_{t}. Note that we can use any feature extractor here. In this work, we adopt the effective two-stream feature representation first proposed in . Based on these features, we construct a LL-level temporal pyramid where each level evenly divides the interval into BlB_{l} parts. For the ii-th part of the ll-th level, whose interval is [sli,eli][s_{li},e_{li}], we can derive a pooled feature as

Then the overall representation of this stage can be obtained by concatenating the pooled features across all parts at all levels as fic=(ui(l)∣l=1,…,L, i=1, …,Bl)\mathbf{f}_{i}^{c}=(\mathbf{u}_{i}^{(l)}|l=1,\ldots,L,~{}i=1,~{}\ldots,B_{l}).

We treat the three stages differently. Generally, we observed that the course stage, which reflects the activity process itself, usually contains richer structure e.g. this process itself may contain sub-stages. Hence, we use a two-level pyramid, i.e. L=2,B1=1L=2,B_{1}=1, and B2=2B_{2}=2, for the course stage, while using simpler one-level pyramids (which essentially reduce to standard average pooling) for starting and ending pyramids. We found empirically that this setting strikes a good balance between expressive power and complexity. Finally, the stage-wise features are combined via concatenation. Overall, this construction explicitly leverages the structure of an activity instance and its surrounding context, and thus we call it structured temporal pyramid pooling (STPP).

3 Activity and Completeness Classifiers

On top of the structured features described above, we introduce two types of classifiers, an activity classifier and a set of completeness classifiers. Specifically, the activity classifier AA classifies input proposals into K+1K+1 classes, i.e. KK activity classes (with labels 1,…,K1,\ldots,K) and an additional “background” class (with label ). This classifier restricts its scope to the course stage, making predictions based on the corresponding feature fic\mathbf{f}_{i}^{c}. The completeness classifiers {Ck}k=1K\{C_{k}\}_{k=1}^{K} are a set of binary classifiers, each for one activity class. Particularly, CkC_{k} predicts whether a proposal captures a complete activity instance of class kk, based on the global representation {fis,fic,fie}\{\mathbf{f}_{i}^{s},\mathbf{f}_{i}^{c},\mathbf{f}_{i}^{e}\} induced by STPP. In this way, the completeness is determined not only on the proposal itself but also on its surrounding context.

Both types of classifiers are implemented as linear classifiers on top of high-level features. Given a proposal pip_{i}, the activity classifier will produce a vector of normalized responses via a softmax layer. From a probabilistic view, it can be considered as a conditional distribution P(ci∣pi)P(c_{i}|p_{i}), where cic_{i} is the class label. For each activity class kk, the corresponding completeness classifier CkC_{k} will yield a probability value, which can be understood as the conditional probability P(bi∣ci,pi)P(b_{i}|c_{i},p_{i}), where bib_{i} indicates whether pip_{i} is complete. Both outputs together form a joint distribution. When ci≥1c_{i}\geq 1, P(ci,bi∣pi)=P(ci∣pi)⋅P(bi∣ci,pi).P(c_{i},b_{i}|p_{i})=P(c_{i}|p_{i})\cdot P(b_{i}|c_{i},p_{i}). Hence, we can define a unified classification loss jointly on both types of classifiers. With a proposal pip_{i} and its label cic_{i}:

Here, the completeness term P(bi∣ci,pi)P(b_{i}|c_{i},p_{i}) is only used when ci≥1c_{i}\geq 1, i.e. the proposal pip_{i} is not considered as part of the background. Note that these classifiers together with STPP are integrated into a single network that is trained in an end-to-end way.

During training, we collect three types of proposal samples: (1) positive proposals, i.e. those overlap with the closest groundtruth instances with at least 0.70.7 IoU; (2) background proposals, i.e. those that do not overlap with any groundtruth instances; and (3) incomplete proposals, i.e. those that satisfy the following criteria: 80%80\% of its own span is contained in a groundtruth instance, while its IoU with that instance is below 0.30.3 (in other words, it just covers a small part of the instance). For these proposal types, we respectively have (ci>0,bi=1)(c_{i}>0,b_{i}=1), ci=0c_{i}=0, and (ci>0,bi=0)(c_{i}>0,b_{i}=0). Each mini-batch is ensured to contain all three types of proposals.

4 Location Regression and Multi-Task Loss

With the structured information encoded in the global features, we can not only make categorical predictions, but also refine the proposal’s temporal interval itself by location regression. We devise a set of location regressors {Rk}k=1K\{R_{k}\}_{k=1}^{K}, each for an activity class. We follow the design in RCNN , but adapting it for 1D temporal regions. Particularly, for a positive proposal pip_{i}, we regress the relative changes of both the interval center μi\mu_{i} and the span ϕi\phi_{i} (in log-scale), using the closest groundtruth instance as the target. With both the classifiers and location regressors, we define a multi-task loss over an training sample pip_{i}, as:

Here, Lreg\mathcal{L}_{reg} uses the smooth L1L_{1} loss function .

Efficient Training and Inference with SSN

The huge amount of frames poses a serious challenge in computational cost to video analysis. Our structured segment network also faces this challenge. This section presents two techniques which we use to reduce the cost and enable end-to-end training.

The structured temporal pyramid, in its original form, rely on densely sampled snippets. This would lead to excessive computational cost and memory demand in end-to-end training over long proposals – in practice, proposals that span over hundreds of frames are not uncommon. However, dense sampling is generally unnecessary in our framework. Particularly, the pooling operation is essentially to collect feature statistics over a certain region. Such statistics can be well approximated via a subset of snippets, due to the high redundancy among them.

Motivated by this, we devise a sparse snippet sampling scheme. Specifically, given a augmented proposal pi′p^{\prime}_{i}, we evenly divide it into L=9L=9 segments, randomly sampling only one snippet from each segment. Structured temporal pyramid pooling is performed for each pooling region on its corresponding segments. This scheme is inspired by the segmental architecture in , but differs in that it operates within STPP instead of a global average pooling. In this way, we fix the number of features needed to be computed regardless of how long the proposal is, thus effectively reducing the computational cost, especially for modeling long-term structures. More importantly, this enables end-to-end training of the entire framework over a large number of long proposals.

Inference with reordered computation.

In testing, we sample video snippets with a fixed interval of 66 frames, and construct the temporal pyramid thereon. The original formulation of temporal pyramid first computes pooled features and then applies the classifiers and regressors on top which is not efficient. Actually, for each video, hundreds of proposals will be generated, and these proposals can significantly overlap with each other – therefore, a considerable portion of the snippets and the features derived thereon are shared among proposals.

To exploit this redundancy in the computation, we adopt the idea introduced in position sensitive pooling to improve testing efficiency. Note that our classifiers and regressors are both linear. So the key step in classification or regression is to multiply a weight matrix W\mathbf{W} with the global feature vector f\mathbf{f}. Recall that f\mathbf{f} itself is a concatenation of multiple features, each pooled over a certain interval. Hence the computation can be written as Wf=∑jWjfj\mathbf{W}\mathbf{f}=\sum_{j}\mathbf{W}_{j}\mathbf{f}_{j}, where jj indexes different regions along the pyramid. Here, fj\mathbf{f}_{j} is obtained by average pooling over all snippet-wise features within the region rjr_{j}. Thus, we have

Temporal Region Proposals

In general, SSN accepts arbitrary proposals, e.g. sliding windows . Yet, an effective proposal method can produce more accurate proposals, and thus allowing a small number of proposals to reach a certain level of performance. In this work, we devise an effective proposal method called temporal actionness grouping (TAG).

This method uses an actionness classifier to evaluate the binary actionness probabilities for individual snippets. The use of binary actionness for proposals is first introduced in spatial action detection by . Here we utilize it for temporal action detection.

Our basic idea is to find those continuous temporal regions with mostly high actionness snippets to serve as proposals. To this end, we repurpose a classic watershed algorithm , applying it to the 1D signal formed by a sequence of complemented actionness values, as shown in Figure 3. Imagine the signal as 1D terrain with heights and basins. This algorithm floods water on this terrain with different “water level” (γ)(\gamma), resulting in a set of “basins” covered by water, denoted by G(γ)G(\gamma). Intuitively, each “basin” corresponds to a temporal region with high actionness. The ridges above water then form the blank areas between basins, as illustrated in Fig. 3.

Given a set of basins G(γ)G(\gamma), we devise a grouping scheme similar to , which tries to connect small basins into proposal regions. The scheme works as follows: it begins with a seed basin, and consecutively absorbs the basins that follow, until the fraction of the basin durations over the total duration (i.e. from the beginning of the first basin to the ending of the last) drops below a certain threshold τ\tau. The absorbed basins and the blank spaces between them are then grouped to form a single proposal. We treat each basin as seed and perform the grouping procedure to obtain a set of proposals denoted by G′(τ,γ)G^{\prime}(\tau,\gamma). Note that we do not choose a specific combination of τ\tau and γ\gamma. Instead we uniformly sample τ\tau and γ\gamma from ∈(0,1)\in(0,1) with an even step of 0.050.05. The combination of these two thresholds leads to multiple sets of regions. We then take the union of them. Finally, we apply non-maximal suppression to the union with IoU threshold 0.950.95, to filter out highly overlapped proposals. The retained proposals will be fed to the SSN framework.

Experimental Results

We conducted experiments to test the proposed framework on two large-scale action detection benchmark datasets: ActivityNet and THUMOS14 . In this section we first introduce these datasets and other experimental settings and then investigate the impact of different components via a set of ablation studies. Finally we compare the performance of SSN with other state-of-the-art approaches.

ActivityNet has two versions, v1.2 and v1.3. The former contains 96829682 videos in 100100 classes, while the latter, which is a superset of v1.2 and was used in the ActivityNet Challenge 2016, contains 1999419994 videos in 200200 classes. In each version, the dataset is divided into three disjoint subsets, training, validation, and testing, by 22:11:11. THUMOS14 has 10101010 videos for validation and 15741574 videos for testing. This dataset does not provide the training set by itself. Instead, the UCF101 , a trimmed video dataset is appointed as the official training set. Following the standard practice, we train out models on the validation set and evaluate them on the testing set. On these two sets, 220220 and 212212 videos have temporal annotations in 2020 classes, respectively. 22 falsely annotated videos (“270”,“1496”) in the test set are excluded in evaluation. In our experiments, we compare with our method with the states of the art on both THUMOS14 and ActivityNet v1.3, and perform ablation studies on ActivityNet v1.2.

Implementation Details.

We train the structured segment network in an end-to-end manner, with raw video frames and action proposals as the input. Two-stream CNNs are used for feature extraction. We also use the spatial and temporal streams to harness both the appearance and motion features. The binary actionness classifiers underlying the TAG proposals are trained with on the training subset of each dataset. We use SGD to learn CNN parameters in our framework, with batch size 128128 and momentum 0.90.9. We initialize the CNNs with pre-trained models from ImageNet . The initial learning rates are set to 0.0010.001 for RGB networks and 0.0050.005 for optical flow networks. In each minibatch, we keep the ratio of three types of proposals, namely positive, background, and incomplete, to be 11:11:66. For the completeness classifiers, only the samples with loss values ranked in the first 1/61/6 of a minibatch are used for calculating gradients, which resembles online hard negative mining . On both versions of ActivityNet, the RGB and optical flow branches of the two-stream CNN are respectively trained for 9.5K9.5K and 20K20K iterations, with learning rates scaled down by 0.10.1 after every 4K4K and 8K8K iterations, respectively. On THUMOS14, these two branches are respectively trained for 1K1K and 6K6K iterations, with learning rates scaled down by 0.10.1 per 400400 and 25002500 iterations.

Evaluation Metrics.

As both datasets originate from contests, each dataset has its own convention of reporting performance metrics. We follow their conventions, reporting mean average precision (mAP) at different IoU thresholds. On both versions of ActivityNet, the IoU thresholds are {0.5,0.75,0.95}\{0.5,0.75,0.95\}. The average of mAP values with IoU thresholds [0.5[0.5:0.050.05:0.95]0.95] is used to compare the performance between different methods. On THUMOS14, the IoU thresholds are {0.1,0.2,0.3,0.4,0.5}\{0.1,0.2,0.3,0.4,0.5\}. The mAP at 0.50.5 IoU is used for comparing results from different methods.

2 Ablation Studies

We compare the performance of different action proposal schemes in three aspects, i.e. recall, quality, and detection performance. Particularly, we compare our TAG scheme with common sliding windows as well as other state-of-the-art proposal methods, including SCNN-prop, a proposal networks presented in , TAP , DAP . For the sliding window scheme, we use 2020 exponential scales starting from 0.30.3 second long and step sizes of 0.40.4 times of window lengths.

We first evaluate the average recall rates, which are summarized in Table 1. We can see that TAG proposal have higher recall rates with the same number of proposals. Then we investigate the quality of its proposals. We plot the recall rates from different proposal methods at different IoU thresholds in Fig. 4. We can see TAG retains relatively high recall at high IoU thresholds, demonstrating that the proposals from TAG are generally more accurate. In experiments we also tried applying the actionness classifier trained on ActivityNet v1.2 directly on THUMOS14. We can still achieve a reasonable average recall of 39.6%39.6\%, while the one trained on THUMOS14 achieves 48.9%48.9\% in Table 1. Finally, we evaluate the proposal methods in the context of action detection. The detection mAP values using sliding window proposals and TAG proposals are shown in Table 3. The results confirm that, in most cases, the improved proposals can result in improved detection performance.

Structured Temporal Pyramid Pooling.

Here we study the influence of different pooling strategies in STPP. We denote one pooling configuration as (B1,…,BK)−A(B_{1},\ldots,B_{K})-A, where KK refers to the number of pyramid levels for the course stage and B1,…,BKB_{1},\ldots,B_{K} the number of regions in each level. A=1A=1 indicates we use augmented proposal and model the starting and ending stage, while A=0A=0 indicates we only use the original proposal (without augmentation). Additionally we compare two within-region pooling methods: average and max pooling. The results are summarized in Table 2. Note that these configurations are evaluated in the stage-wise training scenario. We observe that cases where A=0A=0 have inferior performance, showing that the introduction of the stage structure is very important for accurate detection. Also, increasing the depth of the pyramids for the course stage can give slight performance gain. Based on these results, we fix the configuration to (1,2)−1(1,2)-1 in later experiments.

Classifier Design.

In this work, we introduced the activity and completeness classifiers which work together to classify the proposal. We verify the importance of this decomposed design by studying another design that replaces it with a single set of classifiers, for which both background and incomplete samples are uniformly treated as negative. We perform similar negative sample mining for this setting. The results are summarized in Table 3. We observe that using only one classifier to distinguish positive samples from both background and incomplete would lead to worse result even with negative mining, where mAP decreased from 23.7%23.7\% to 17.9%17.9\%. We attribute this performance gain to the different natures of the two negative proposal types, which require different classifiers to handle.

Location Regression & Multi-Task Learning.

Because of the contextual information contained in the starting and ending stages of the global region features, we are able to perform location regression. We measure the contribution of this step to the detection performance in Table 3. From the results we can see that the location regression and multi-task learning, where we train the classifiers and the regressors together in an end-to-end manner, always improve the detection accuracy.

Training: Stage-wise v.s. End-to-end.

While the structured segment network is designed for end-to-end training, it is also possible to first densely extract features and train the classifiers and regressors with SVM and ridge regression, respectively. We refer to this training scheme as stage-wise training. We compare the performance of end-to-end training and stage-wise training in Table 3. We observe that models from end-to-end training can slightly outperform those learned with stage-wise training under the same settings. This is remarkable as we are only sparsely sampling snippets in end-to-end training, which also demonstrates the importance of jointly optimizing the classifiers and feature extractors and justifies our framework design. Besides, end-to-end training has another major advantage that it does not need to store the extracted features for the training set, which could become quite storage intensive as training data grows.

3 Comparison with the State of the Art

Finally, we compare our method with other state-of-the-art temporal action detection methods on THUMOS14 and ActivityNet v1.3 , and report the performances using the metrics described above. Note that the average action duration in THUMOS14 and ActivityNet are 44 and 5050 seconds. And the average video duration are 233233 and 114114 seconds, respectively. This reflects the distinct natures of these datasets in terms of the granularities and temporal structures of the action instances. Hence, strong adaptivity is required to perform consistently well on both datasets.

On THUMOS 14, We compare with the contest results and those from recent works, including the methods that use segment-based 3D CNN , score pyramids , and recurrent reinforcement learning . The results are shown in Table 4. In most cases, the proposed method outperforms previous state-of-the-art methods by over 10%10\% in absolute mAP values.

ActivityNet.

The results on the testing set of ActivityNet v1.3 are shown in Table 5. For references, we list the performances of highest ranked entries in the ActivityNet 2016 challenge. We submit our results to the test server of ActivityNet v1.3 and report the detection performance on the testing set. The proposed framework, using a single model instead of an ensemble, is able to achieve an average mAP of 28.2828.28 and perform well at high IOU thresholds, i.e., 0.750.75 and 0.950.95. This clearly demonstrates the superiority of our method. Visualization of the detection results can be found in the supplementary materials .

Conclusion

In this paper, we presented a generic framework for temporal action detection, which combines a structured temporal pyramid with two types of classifiers, respectively for predicting activity class and completeness. With this framework, we achieved significant performance gain over state-of-the-art methods on both ActivityNet and THUMOS14. Moreover, we demonstrated that our method is both accurate and generic, being able to localize temporal boundaries precisely and working well for activity classes with very different temporal structures.

This work is partially supported by the Big Data Collaboration Research grant from SenseTime Group (CUHK Agreement No. TS1610626), the General Research Fund (GRF) of Hong Kong (No. 14236516) and the Early Career Scheme (ECS) of Hong Kong (No. 24204215).

References