Graph Convolutional Networks for Temporal Action Localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, Chuang Gan
Introduction
Understanding human actions in videos has been becoming a prominent research topic in computer vision, owing to its various applications in security surveillance, human behavior analysis and many other areas . Despite the fruitful progress in this vein, there are still some challenging tasks demanding further exploration — temporal action localization is such an example. To deal with real videos that are untrimmed and usually contain the background of irrelevant activities, temporal action localization requires the machine to not only classify the actions of interest but also localize the start and end time of every action instance. Consider a sport video as illustrated in Figure 1, the detector should find out the frames where the action event is happening and identify the category of the event.
Temporal Action localization has attracted increasing attention in the last several years . Inspired by the success of object detection, most current action detection methods resort to the two-stage pipeline: they first generate a set of 1D temporal proposals and then perform classification and temporal boundary regression on each proposal individually. However, processing each proposal separately in the prediction stage will inevitably neglect the semantic relations between proposals.
We contend that exploiting the proposal-proposal relations in the video domain provides more cues to facilitate the recognition of each proposal instance. To illustrate this, we revisit the example in Figure 1, where we have generated four proposals. On the one hand, the proposals , and overlapping with each other describe different parts of the same action instance (i.e., the start period, main body and end period). Conventional methods perform prediction on by using its feature alone, which we think is insufficient to deliver complete knowledge for the detection. If we additionally take the features of and into account, we will obtain more contextual information around , which is advantageous especially for the temporal boundary regression of . On the other hand, describes the background (i.e., the sport field), and its content is also helpful in identifying the action label of , since what is happening on the sport field is likely to be sport action (e.g. “discus throwing”) but not the one happens elsewhere (e.g. “kissing”). In other words, the classification of can be partly guided by the content of even they are temporally disjointed.
To model the proposal-proposal interactions, one may employ the self-attention mechanism — as what has been conducted previously in language translation and object detection — to capture the pair-wise similarity between proposals. A self-attention module can affect an individual proposal by aggregating information from all other proposals with the automatically learned aggregation weights. However, this method is computationally expensive as querying all proposal pairs has a quadratic complexity of the proposal number (note that each video could contain more than thousands of proposals). On the contrary, Graph Convolutional Networks (GCNs) , which generalize convolutions from grid-like data (e.g. images) to non-grid structures (e.g. social networks), have received increasing interests in the machine learning domain . GCNs can affect each node by aggregating information from the adjacent nodes, and thus are very suitable for leveraging the relations between proposals. More importantly, unlike the self-attention strategy, applying GCNs enables us to aggregate information from only the local neighbourhoods for each proposal, and thus can help decrease the computational complexity remarkably.
In this paper, we regard the proposals as nodes of a specific graph and take advantage of GCNs for modeling the proposal relations. Motivated by the discussions above, we construct the graph by investigating two kinds of edges between proposals, including the contextual edges to incorporate the contextual information for each proposal instance (e.g., detecting by accessing and in Figure 1) and the surrounding edges to query knowledge from nearby but distinct proposals (e.g., querying for in Figure 1).
We then perform graph convolutions on the constructed graph. Although the information is aggregated from local neighbors in each layer, message passing between distant nodes is still possible if the depth of GCNs increases. Besides, we conduct two different GCNs to perform classification and regression separately, which is demonstrated to be effective by our experiments. Moreover, to avoid the overwhelming computation cost, we further devise a sampling strategy to train the GCNs efficiently while still preserving desired detection performance. We evaluate our proposed method on two popular benchmarks for temporal action detection, i.e., THUMOS14 and AcitivityNet1.3 .
To sum up, our contributions are as follow:
To the best of our knowledge, we are the first to exploit the proposal-proposal relations for temporal action localization in videos.
To model the interactions between proposals, we construct a graph of proposals by establishing the edges based on our valuable observations and then apply GCNs to do message aggregation among proposals.
We have verified the effectiveness of our proposed method on two benchmarks. On THUMOS14 especially, our method obtains the mAP of when , which significantly outperforms the state-of-the-art, i.e. by . Augmentation experiments on ActivityNet also verify the efficacy of modeling action proposal relationships.
Related work
Temporal action localization. Recently, great progress has been achieved in deep learning , which facilitates the development of temporal action localization. Approaches on this task can be grouped into three categories: (1) methods performing frame or segment-level classification where the smoothing and merging steps are required to obtain the temporal boundaries ; (2) approaches employing a two-stage framework involving proposal generation, classification and boundary refinement ; (3) methods developing end-to-end architectures integrating the proposal generation and classification .
Our work is built upon the second category where the action proposals are first generated and then used to perform classification and boundary regression. Following this paradigm, Shou et al. propose to generate proposals from sliding windows and classify them. Xu et al. exploit the 3D ConvNet and propose a framework inspired by Faster R-CNN . The above methods neglect the context information of proposals, and hence some attempts have been developed to incorporate the context to enhance the proposal feature . They show encouraging improvements by extracting features on the extended receptive field (i.e., boundary) of the proposal. Despite their success, they all process each proposal individually. In contrast, our method has considered the proposal-proposal interactions and leveraged the relations between proposals.
Graph Convolutional Networks. Kipf et al. propose the Graph Convolutional Networks (GCNs) to define convolutions on the non-grid structures . Thanks to its effectiveness, GCNs have been successfully applied to several research areas in computer vision, such as skeleton-based action recognition , person re-identification , and video classification . For real-world applications, the graph can be large and directly using GCNs is inefficient. Therefore, several attempts are posed for efficient training by virtue of the sampling strategy, such as the node-wise method SAGE , layer-wise model FastGCN and its layer-dependent variant AS-GCN . In this paper, considering the flexibility and implementability, we adopt SAGE method as the sampling strategy in our framework.
Our Approach
2 General Scheme of Our Approach
In this paper, we use a proposal graph to present the relations between proposals and then apply GCN on the graph to exploit the relations and learn powerful representations for proposals. The intuition behind applying GCN is that when performing graph convolution, each node aggregates information from its neighborhoods. In this way, the feature of each proposal is enhanced by other proposals, which helps boost the detection performance eventually.
Without loss of generality, we assume the action proposals have been obtained beforehand by some methods (e.g., the TAG method in ). In this paper, given an input video , we seek to predict the action category and temporal position for each proposal by exploiting proposal relations. Formally, we compute
where denotes any mapping functions to be learned. To exploit GCN for action localization, our paradigm takes both the proposal graph and the proposal features as input and perform graph convolution on the graph to leverage proposal relations. The enhanced proposal features (i.e., the outputs of GCN) are then used to jointly predict the category label and temporal bounding box. The schematic of our approach is shown in Figure 2. For simplicity, we denote our model as P-GCN henceforth.
In the following sections, we aim to answer two questions: (1) how to construct a graph to represent the relations between proposals; (2) how to use GCN to learn representations of proposals based on the graph and facilitate the action localization.
3 Proposal Graph Construction
For the graph of each video, the nodes are instantiated as the action proposals, while the edges between proposals are demanded to be characterized specifically to better model the proposal relations.
One way to construct edges is linking all proposals with each other, which yet will bring in overwhelming computations for going through all proposal pairs. It also incurs redundant or noisy information for action localization, as some unrelated proposals should not be connected. In this paper, we devise a smarter approach by exploiting the temporal relevance/distance between proposals instead. Specifically, we introduce two types of edges, the contextual edges and surrounding edges, respectively.
Contextual Edges. We establish an edge between proposal and if , where is a certain threshold. Here, represents the relevance between proposals and is defined by the tIoU metric, i.e.,
where and compute the temporal intersection and union of the two proposals, respectively. If we focus on the proposal , establishing the edges by computing will select its neighbourhoods as those have high overlaps with it. Obviously, the non-overlapping portions of the highly-overlapping neighbourhoods are able to provide rich contextual information for . As already demonstrated in , exploring the contextual information is of great help in refining the detection boundary and increasing the detection accuracy eventually. Here, by our contextual edges, all overlapping proposals automatically share the contextual information with each other, and these information are further processed by the graph convolution.
Surrounding Edges. The contextual edges connect the overlapping proposals that usually correspond to the same action instance. Actually, distinct but nearby actions (including the background items) could also be correlated, and the message passing among them will facilitate the detection of each other. For example in Figure 1, the background proposal will provide a guidance on identifying the action class of proposal (e.g., more likely to be sport action). To handle such kind of correlations, we first utilize to query the distinct proposals, and then compute the following distance
to add the edges between nearby proposals if , where is a certain threshold. In Eq. (3), (or ) represents the center coordinate of (or ). As a complement of the contextual edges, the surrounding edges enable the message to pass across distinct action instances and thereby provides more temporal cues for the detection.
4 Graph Convolution for Action Localization
Given the constructed graph, we apply the GCN to do action localization. We build -layer graph convolutions in our implementation. Specifically for the -th layer (), the graph convolution is implemented by
We apply an activation function (i.e., ReLU) after each convolution layer before the features are forwarded to the next layer. In addition, our experiments find it more effective by further concatenating the hidden features with the input features in the last layer, namely,
where denotes the concatenation operation.
Joining the previous work , we find that it is beneficial to predict the action label and temporal boundary separately by virtue of two GCNs—one conducted on the original proposal features and the other one on the extended proposal features . The first GCN is formulated as
where we apply a Fully-Connected (FC) layer with soft-max operation on top of to predict the action label . The second GCN can be formulated as
where the graph structure is the same as that in Eq. (6) but the input proposal feature is different. The extended feature is attained by first extending the temporal boundary of with of its length on both the left and right sides and then extracting the feature within the extended boundary. Here, we adopt two FC layers on top of , one for predicting the boundary and the other one for predicting the completeness label , which indicates whether the proposal is complete or not. It has been demonstrated by that, incomplete proposals that have low tIoU with the ground-truths could have high classification score, and thus it will make mistakes when using the classification score alone to rank the proposal for the mAP test; further applying the completeness score enables us to avoid this issue.
Adjacency Matrix. In Eq. (4), we need to compute the adjacency matrix . Here, we design the adjacency matrix by assigning specific weights to edges. For example, we can apply the cosine similarity to estimate the weights of edge by
In the above computation, we compute relying on the feature vector . We can also map the feature vectors into an embedding space using a learnable linear mapping function as in before the cosine computation. We leave the discussion in our experiments.
5 Efficient Training by Sampling
Typical proposal generation methods usually produce thousands of proposals for each video. Applying the aforementioned graph convolution (Eq. (4)) on all proposals demands hefty computation and memory footprints. To accelerate the training of GCNs, several approaches have been proposed based on neighbourhood sampling. Here, we adopt the SAGE method in our method for its flexibility.
The SAGE method uniformly samples the fixed-size neighbourhoods of each node layer-by-layer in a top-down passway. In other words, the nodes of the -th layer are formulated as the sampled neighbourhoods of the nodes in the -th layer. After all nodes of all layers are sampled, SAGE performs the information aggregation in a bottom-up fashion. Here we specify the aggregation function to be a sampling form of Eq. (4), namely,
where node is sampled from the neighbourhoods of node , i.e., ; is the sampling size and is much less than the total number . The summation in Eq. (10) is further normalized by , which empirically makes the training more stable. Besides, we also enforce the self addition of its feature for node in Eq. (10). We do not perform any sampling when testing. For better readability, Algorithm 1 depicts the algorithmic Flow of our method.
Experiments
THUMOS14 is a standard benchmark for action localization. Its training set known as the UCF-101 dataset consists of 13320 videos. The validation, testing and background set contain 1010, 1574 and 2500 untrimmed videos, respectively. Performing action localization on this dataset is challenging since each video has more than 15 action instances and its 71% frames are occupied by background items. Following the common setting in , we apply 200 videos in the validation set for training and conduct evaluation on the 213 annotated videos from the testing set.
ActivityNet is another popular benchmark for action localization on untrimmed videos. We evaluate our method on ActivityNet v1.3, which contains around 10K training videos and 5K validation videos corresponded to 200 different activities. Each video has an average of 1.65 action instances. Following the standard practice, we train our method on the training videos and test it on the validation videos. In our experiments, we contrast our method with the state-of-the-art methods on both THUMOS14 and ActivityNet v1.3, and perform ablation studies on THUMOS14.
2 Implementation details
Evaluation Metrics. We use mean Average Precision (mAP) as the evaluation metric. A proposal is considered to be correct if its temporal IoU with the ground-truth instance is larger than a certain threshold and the predicted category is the same as this ground-truth instance. On THUMOS14, the tIOU thresholds are chosen from ; on ActivityNet v1.3, the IoU thresholds are from , and we also report the average mAP of the IoU thresholds between 0.5 and 0.95 with the step of .
Features and Proposals. Our model is implemented under the two-stream strategy : RGB frames and optical flow fields. We first uniformly divide each input video into 64-frame segments. We then use a two-stream Inflated 3D ConvNet (I3D) model pre-trained on Kinetics to extract the segment features. In detail, the I3D model takes as input the RGB/optical-flow segment and outputs a 1024-dimensional feature vector for each segment. Upon the I3D features, we further apply max pooling across segments to obtain one 1024-dimensional feature vector for each proposal that is obtained by the BSN method . Note that we do not finetune the parameters of the I3D model in our training phase. Besides the I3D features and BSN proposals, our ablation studies in § 5 also explore other types of features (e.g. 2-D features ) and proposals (e.g. TAG proposals ).
Proposal Graph Construction. We construct the proposal graph by fixing the values of as 0.7 and as 1 for both streams. More discussions on choosing the values of and could be found in the supplementary material. We adopt 2-layer GCN since we observed no clear improvement with more than 2 layers but the model complexity is increased. For more efficiency, we choose in Eq. (10) for neighbourhood sampling unless otherwise specified.
Training. The initial learning rate is 0.001 for the RGB stream and 0.01 for the Flow stream. During training, the learning rates will be divided by 10 every 15 epochs. The dropout ratio is 0.8. The classification and completeness are trained with the cross-entropy loss and the hinge loss, respectively. The regression term is trained with the smooth loss. More training details can be found in the supplementary material.
Testing. We do not perform neighbourhood sampling (i.e. Eq. (10)) for testing. The predictions of the RGB and Flow steams are fused using a ratio of 2:3. We multiply the classification score with the completeness score as the final score for calculating mAP. We then use Non-Maximum Suppression (NMS) to obtain the final predicted temporal proposals for each action class separately. We use 600 and 100 proposals per video for computing mAPs on THUMOS14 and ActivityNet v1.3, respectively.
3 Comparison with state-of-the-art results
THUMOS14. Our P-GCN model is compared with the state-of-the-art methods in Table 1. The P-GCN model reaches the highest mAP over all thresholds, implying that our method can recognize and localize actions much more accurately than any other method. Particularly, our P-GCN model outperforms the previously best method (i.e. TAL-Net ) by 6.3% absolute improvement and the second-best result by more than 12.2%, when .
ActivityNet v1.3. Table 2 reports the action localization results of various methods. Regarding the average mAP, P-GCN outperforms SSN , CDC , and TAL-Net by 3.01%, 3.19%, and 6.77%, respectively. We observe that the method by Lin et al. (called LIN below) performs promisingly on this dataset. Note that LIN is originally designed for generating class-agnostic proposals, and thus relies on external video-level action labels (from UntrimmedNet ) for action localization. In contrast, our method is self-contained and is able to perform action localization without any external label. Actually, P-GCN can still be modified to take external labels into account. To achieve this, we assign the top-2 video-level classes predicted by UntrimmedNet to all the proposals in that video. We provide more details about how to involve external labels in P-GCN in the supplementary material. As summarized in Table 2, our enhanced version P-GCN* consistently outperforms LIN, hence demonstrating the effectiveness of our method under the same setting.
Ablation Studies
In this section, we will perform complete and in-depth ablation studies to evaluate the impact of each component of our model. More details about the structures of baseline methods (such as MLP and MP) can be found in the supplementary material.
As illustrated in § 3.4, we apply two GCNs for action classification and boundary regression separately. Here, we implement the baseline with a 2-layer MultiLayer-Perceptron (MLP). The MLP baseline shares the same structure as GCN except that we remove the adjacent matrix in Eq. (4). To be specific, for the -th layer, the propagation in Eq. (4) becomes , where are the trainable parameters. Without using , MLP processes each proposal feature independently. By comparing the performance of MLP with GCN, we can justify the importance of message passing along proposals. To do so, we replace each GCN with an MLP and have the following variants of our model including: (1) MLP1 + GCN2 where GCN1 is replaced; (2) GCN1 + MLP2 where GCN2 is replaced; and (3) MLP1 + MLP2 where both GCNs are replaced. Table 3 reads that all these variants decrease the performance of our model, thus verifying the effectiveness of GCNs for both action classification and boundary regression. Overall, our model P-GCN significantly outperforms the MLP protocol (i.e. MLP1 + MLP2), validating the importance of considering proposal-proposal relations in temporal action localization.
2 How does the graph convolution help?
Besides graph convolutions, performing mean pooling among proposal features is another way to enable information dissemination between proposals. We thus conduct another baseline by first adopting MLP on the proposal features and then conducting mean pooling on the output of MLP over adjacent proposals. The adjacent connections are formulated by using the same graph as GCN. We term this baseline as MP below. Similar to the setting in § 5.1, we have three variants of our model including: (1) MP1 + MP2; (2) MP1 + GCN2; and (3) GCN1 + MP2. We report the results in Table 4. Our P-GCN outperforms all MP variants, demonstrating the superiority of graph convolution over mean pooling on capturing between-proposal connections. The protocol MP1 + MP2 in Table 4 performs better than MLP1 + MLP2 in Table 3, which again reveals the benefit of modeling the proposal-proposal relations, even we pursue it using the naive mean pooling.
3 Influences of different backbones
Our framework is general and compatible with different backbones (i.e., proposals and features). Beside the backbones applied above, we further perform experiments on TAG proposals and 2D features . We try different combinations: (1) BSN+I3D; (2) BSN+2D; (3) TAG+I3D; (4) TAG+2D, and report the results of MLP and P-GCN in Figure 3. In comparison with MLP, our P-GCN leads to significant and consistent improvements in all types of features and proposals. These results conclude that, our method is generally effective and is not limited to the specific feature or proposal type.
4 The weights of edge and self-addition
We have defined the weights of edges in Eq. (9), where the cosine similarity (cos-sim) is applied. This similarity can be further extended by first embedding the features before the cosine computation. We call the embedded version as embed-cos-sim, and compare it with cos-sim in Table 5. No obvious improvement is attained by replacing cos-sim with embed-cos-sim (the mAP difference between them is less than ). Eq. (10) has considered the self-addition of the node feature. We also investigate the importance of this term in Table 5. It suggests that the self-addition leads to at least 1.7% absolute improvements on both RGB and Flow streams.
5 Is it necessary to consider two types of edges?
To evaluate the necessity of formulating two types of edges, we perform experiments on two variants of our P-GCN, each of which considers only one type of edge in the graph construction stage. As expected, the result in Table 6 drops remarkably when either kind of edge is removed. Another crucial point is that our P-GCN still boosts MLP when only the surrounding edges are remained. The rationale behind this could be that, actions in the same video are correlated and exploiting the surrounding relation will enable more accurate action classification.
6 The efficiency of our sampling strategy
We train P-GCN efficiently based on the neighbourhood sampling in Eq. (10). Here, we are interested in how the sampling size affects the final performance. Table 7 reports the testing mAPs corresponded to different varying from 1 to 5 (and also 10). The training time per iteration is also added in Table 7. We observe that when the model achieves higher mAP than the full model (i.e., ) while reducing 76% of training time for each iteration. This is interesting, as sampling fewer nodes even yields better results. We conjecture that, the neighbourhood sampling could bring in more stochasticity and guide our model to escape from the local minimal during training, thus delivering better results.
7 Qualitative Results
Given the significant improvements, we also attempt to find out in what cases our P-GCN model improves over MLP. We visualize the qualitative results on THUMOS14 in Figure 4. In the top example, both MLP and our P-GCN model are able to predict the action category correctly, while P-GCN predicts a more precise temporal boundary. In the bottom example, due to similar action characteristic and context, MLP predicts the action of “Shotput” as “Throw Discus”. Despite such challenge, P-GCN still correctly predicts the action category, demonstrating the effectiveness of our method. More qualitative results could be found in the supplementary material.
Conclusions
In this paper, we have exploited the proposal-proposal interaction to tackle the task of temporal action localization. By constructing a graph of proposals and applying GCNs to message passing, our P-GCN model outperforms the state-of-the-art methods by a large margin on two benchmarks, i.e., THUMOS14 and ActivithNet v1.3. It would be interesting to extend our P-GCN for object detection in image and we leave it for our future work.
. This work was partially supported by National Natural Science Foundation of China (NSFC) 61602185, 61836003 (key project), Program for Guangdong Introducing Innovative and Enterpreneurial Teams 2017ZT07X183, Guangdong Provincial Scientific and Technological Funds under Grants 2018B010107001, and Tencent AI Lab Rhino-Bird Focused Research Program (No. JR201902).
References
A Proposal Features
We have two types of proposal features and the process of feature extraction is shown in Figure A.
Proposal features. For the original proposal, we first obtain a set of segment-level features within the proposal and then apply max-pooling across segments to obtain one 1024-dimensional feature vector.
Extended proposal features. The boundary of the original proposal is extended with of its length on both the left and right sides, resulting in the extended proposal. Thus, the extended proposal has three portions: start, center and end. For each portion, we follow the same feature extraction process as the original proposal. Finally, the extended proposal feature is obtained by concatenating the feature of three portions.
B Network Architectures
P-GCN. The network architecture of our P-GCN model is shown in Figure B. Let and be the number of proposals in one video and the total number of action categories, respectively. On the top of GCN, we have three fully-connected (FC) layers for different purposes. The one with outputs is for boundary regression and the other two with outputs are designed for action classification and completeness classification, respectively.
MLP baseline. The network architecture of MLP baseline is shown in Figure C. We replace each of GCNs with a 2-layer multilayer perceptron (MLP). We set the number of parameters in MLP the same as GCN’s for a fair comparison. Note that MLP processes each proposal independently without exploiting the relations between proposals.
Mean-Pooling baseline. As shown in Figure D, the network architecture of Mean-Pooling baseline is the same as the MLP baseline’s except that we conduct mean-pooling on the output of MLP over the adjacent proposals.
C Training Details
We have three types of training samples chosen by two criteria, i.e., the best tIoU and best overlap. For each proposal, we calculate its tIoU with all the ground truth in that video and choose the largest tIoU as the best tIoU (similarly for best overlap). For simplicity, we denote the best tIoU and best overlap as tIoU and OL. Then, three types of training samples can be described as: (1) Foreground sample: ; (2) Incomplete sample: ; (3) Background sample: . These certain thresholds are slightly different on two datasets as shown in Table A. We consider all foreground proposals as the complete proposals.
Each mini-batch contains examples sampled from a single video. The ratio of three types of samples is fixed to (1):(2):(3)=1:6:1. We set the mini-batch size to 32 on THUMOS14 and 64 on ActivityNet v1.3.
For more efficiency, we fix the number of neighborhoods for each node to be 10 by selecting contextual edges with the largest relevance scores and surrounding edges with the smallest distances, where the ratio of contextual and surrounding edges is set to 4:1.
In addition, we empirically found that setting to 0 (when ) leads to better results.
D Loss function
Multi-task Loss. Our P-GCN model can not only predict action category but also refine the proposal’s temporal boundary by location regression. With the action classifier, completeness classifier and location regressors, we define a multi-task loss by:
where and is the predicted probability and ground truth action label of the -th proposal in a mini-batch, respectively. Here, 0 represents the background class. is the completeness label. and are the predicted and ground truth offset, which will be detailed below. In all experiments, we set .
Completeness Loss. Here, the completeness term is used only when , i.e., the proposal is not considered as part of the background.
Regression Loss. We devise a set of location regressors , each for an activity category. For a proposal, we regress the boundary using the closest ground truth instance as the target. Our P-GCN model does not predict the start time and end time of each proposal directly. Instead, it predicts the offset relative to the proposal, where and are the offset of center coordinate and length, respectively. The ground truth offset is denoted as and parameterized by:
where and denote the original center coordinate and length of the proposal, respectively. and account for the center coordinate and length of the closest ground truth, respectively. is the smooth L1 loss and used when and , i.e., the proposal is a foreground sample.
E Details of Augmentation Experiments on ActivityNet
Our P-GCN model can be further augmented by taking the external video-level labels into account. To achieve this, we replace the predicted action classes in Eq. (6) with the external action labels. Specifically, given an input video, we use UntrimmedNet to predict the top-2 video-level classes and assign these classes to all the proposals in this video. In this way, each proposal has two action classes.
To further compute mAP, the score of each proposal is required. In our implementation, we follow the settings in BSN by calculating
where and are the action score and completeness score associated with the action class. represents the confidence score produced by BSN and denotes for the action score predicted by UntrimmedNet.
The parameter is a threshold value for constructing contextual edges, i.e. . Since , can be chosen from . An ablation study is shown in Table B. Our method performs well when .
G Ablation study of boundary regression
We conducted an ablation study on boundary regression in Table C, whose results validate the necessity of using boundary regression.
H Additional runtime compared to [52]
The MLP baseline is indeed a particular implementation of , and it shares the same amount of parameters with our P-GCN. We compare the runtime between P-GCN and MLP in Table D. It reads that GCN only incurs a relatively small additional runtime compared to MLP but is able to boost the performance significantly.