ActionFormer: Localizing Moments of Actions with Transformers

Chenlin Zhang, Jianxin Wu, Yin Li

Introduction

Identifying action instances in time and recognizing their categories, known as temporal action localization (TAL), remains a challenging problem in video understanding. Significant progress has been made in developing deep models for TAL. Most previous works have considered using action proposals or anchor windows , and developed convolutional , recurrent , and graph neural networks for TAL. Despite a steady progress on major benchmarks, the accuracy of existing methods usually comes at a price of modeling complexity, with increasingly sophisticated proposal generation, anchor design, loss function, network architecture, and output decoding process.

In this paper, we adopt a minimalist design and develop a Transformer based model for TAL, inspired by the recent success of Transformers in NLP and vision . Originally developed for sequence data, Transformers use self-attention to model long-range dependencies, and thus are a natural fit for TAL in untrimmed videos. Our method, illustrated in Fig. 1, adapts local self-attention to model temporal context in an input untrimmed videos, classifies every moment, and regresses their corresponding action boundaries. The result is a deep model trained using standard classification and regression loss, and can localize moments of actions in a single shot, without using action proposals or pre-defined anchor windows.

Specifically, our model, dubbed ActionFormer, integrates local self-attention to extract a feature pyramid from an input video. Each location in the output pyramid represents a moment in the video, and is treated as an action candidate. A lightweight convolutional decoder is further employed on the feature pyramid to classify these candidates into foreground action categories, and to regress the distance between a foreground candidate and its action onset and offset. The results can be easily decoded into actions with their labels and temporal boundaries. Our method thus provides a single-stage anchor-free model for TAL.

We show that such a simple model, with proper design, can be surprisingly powerful for TAL. In particular, ActionFormer establishes a new state of the art across several major TAL benchmarks, surpassing previous works by a significant margin. For example, ActionFormer achieves 71.0% mAP at tIoU==0.5 on THUMOS14, outperforming the best prior model by 14.1 absolute percentage points. Further, ActionFormer reaches an average mAP of 36.6% on ActivityNet 1.3. More importantly, ActionFormer shows impressive results on EPIC-Kitchens 100 for egocentric action localization, with a boost of over 13.5 absolute percentage points in average mAP.

Our work is based on simple techniques, supported by favourable empirical results, and validated by extensive ablation experiments, at our best. Our main contributions are summarized as follows. First, we are among the first to propose a Transformer based model for single-stage anchor-free TAL. Second, we study key design choices of developing Transformer models for TAL, and demonstrate a simple model that works surprisingly well. Finally, our model achieves state-of-the-art results across major benchmarks and offers a solid baseline for TAL.

Related Works

Two-stage TAL. These approaches first generate candidate video segments as action proposals, and further classify the proposals into actions and refine their temporal boundaries. Previous works focused on action proposal generation, by either classifying anchor windows or detecting action boundaries , and more recently using a graph representation or Transformers . Others have integrated proposal generation and classification into a single model . More recent effort investigates the modeling of temporal context among proposals using graph neural networks or attention and self-attention mechanisms . Similar to previous approaches, our method considers the modeling of long-term temporal context, yet uses a self-attention within a Transformer model. Different from previous approaches, our model detects actions without using proposals.

Single-stage TAL. Several recent works focused on single-stage TAL, seeking to localize actions in a single shot without using action proposals. Many of them are anchor-based (e.g., using anchor windows sampled from sliding windows). Lin et al. presented the first single-stage TAL using convolutional networks, borrowing ideas from a single-stage object detector . Buch et al. presented a recurrent memory module for single-stage TAL. Long et al. proposed to use Gaussian kernels to dynamically optimize the scale of each anchor, based on a 1D convolutional network. Yang et al. explored the combination of anchor-based and anchor-free models for single-stage TAL, again using convolutional networks. More recently, Lin et al. proposed an anchor-free single-stage model by designing a saliency-based refinement module incorporated in convolutional network. Similar ideas were also explored in video grounding .

Our model falls into the category of single-stage TAL. Indeed, our formulation follows a minimalist design of sequence labeling by classifying every moment and regressing their action boundaries, previously discussed in . The key difference is that we design a Transformer network for action localization. The result is a single stage anchor-free model that outperforms all previous methods. A concurrent work from Liu et al. also used Transformer for TAL, yet considered a set prediction problem similar to DETR .

Spatial-temporal Action Localization. A related yet different task, known as spatial-temporal action localization, is to detect the actions both temporally and spatially, in the form of moving bounding boxes of an actor. It is possible that TAL might be used as a first step for spatial-temporal localization. Girdhar et al. proposed to use Transformer for spatial-temporal action localization. While both our work and use Transformer, the two models differ significantly. We consider a sequence of video frames as the inputs, while used a set of 2D object proposals. Moreover, our work addresses a different task of TAL.

Object Detection. TAL models have been heavily influenced by the developments of object detection models. Some of our model design, including the multiscale feature representation and convolutional decoder, is inspired by feature pyramid network and RetinaNet . Our training using center sampling also stems from recent single-stage object detectors .

Vision Transformer. Transformer models were originally developed for NLP tasks , and has demonstrated recent success for many vision tasks. ViT presented the first pure Transformer-based model that can achieve state-of-the-art performances on image classification. Subsequent works, including DeiT , T2T-ViT , Swin Transformer , Focal Transformer and PVT , have further pushed the envelope, resulting in vision Transformer backbones with impressive results on classification, segmentation, and detection tasks. Transformer have also been explored in object detection , semantic segmentation , and video representation learning . Our model builds on these developments and presents one of the first Transformer models for TAL.

ActionFormer: A Simple Transformer Model for Temporal Action Localization

Given an input video X\mathbf{X}, we assume that X\mathbf{X} can be represented using a set of feature vectors X={x1,x2,…,xT}\mathbf{X}=\{\mathbf{x}_{1},\mathbf{x}_{2},\ldots,\mathbf{x}_{T}\} defined on discretized time steps t={1,2,…,T}t=\{1,2,\ldots,T\}, where the total duration TT varies across videos. For example, xt\mathbf{x}_{t} can be the feature vector of a video clip at moment tt extracted from a 3D convolutional network. The goal of temporal action localization is to predict the action label Y={y1,y2,…,yN}\mathbf{Y}=\{\mathbf{y}_{1},\mathbf{y}_{2},\ldots,\mathbf{y}_{N}\} based on the input video sequence X\mathbf{X}. Y\mathbf{Y} consists of NN action instances yi\mathbf{y}_{i}, where NN also varies across videos. Each instance yi=(si,ei,ai)\mathbf{y}_{i}=(s_{i},e_{i},a_{i}) is defined by its starting time sis_{i} (onset), ending time eie_{i} (offset) and its action label aia_{i}, where si∈[1,T]s_{i}\in[1,T], ei∈[1,T]e_{i}\in[1,T], ai∈{1,..,C}a_{i}\in\{1,..,C\} (CC pre-defined categories) and si<eis_{i}<e_{i}. The task of TAL is thus a challenging problem of structured output prediction.

A Simple Representation for Action Localization. Our method builds on an anchor-free representation for action localization, inspired by . The key idea is to classify each moment as either one of the action categories or the background, and further regress the distance between this time step and the action’s onset and offset. In doing so, we convert the structured output prediction problem (X={x1,x2,...,xT}→Y={y1,y2,…,yN}\mathbf{X}=\{\mathbf{x}_{1},\mathbf{x}_{2},...,\mathbf{x}_{T}\}\rightarrow\mathbf{Y}=\{\mathbf{y}_{1},\mathbf{y}_{2},\ldots,\mathbf{y}_{N}\}) into a more approachable sequence labeling problem

The output y^t=(p(at),dts,dte)\mathbf{\hat{y}}_{t}=(p(a_{t}),d_{t}^{s},d_{t}^{e}) at time tt is defined as

p(at)p(a_{t}) consists of CC values, with each representing a binomial variable indicating the probability of action category ata_{t} (∈{1,2,…,C}\in\{1,2,\ldots,C\}) at time tt. This can be considered as the outputs of CC binary classification.

dts>0d_{t}^{s}>0 and dte>0d_{t}^{e}>0 are the distance between the current time tt to the action’s onset and offset, respectively. dtsd_{t}^{s} and dted_{t}^{e} are not defined if the time step tt lies on the background.

Intuitively, this formulation considers every moment tt in the video X\mathbf{X} as an action candidate, recognizes the action’s category ata_{t}, and estimates the distances between current step and the action boundaries (dtsd_{t}^{s} and dted_{t}^{e}) if an action presents. Action localization results can be readily decoded from y^t=(p(at),dts,dte)\mathbf{\hat{y}}_{t}=(p(a_{t}),d_{t}^{s},d_{t}^{e}) by

Method Overview. Our model — ActionFormer learns to label an input video sequence f(X)→Y^f(\mathbf{X})\rightarrow\mathbf{\hat{Y}}. Specifically, ff is realized using a deep model. ActionFormer follows an encoder-decoder architecture proven successful in many vision tasks, and decomposes ff as h∘gh\circ g. Here g ⁣:X→Zg\colon\mathbf{X}\rightarrow\mathbf{Z} encodes the input into a latent vector Z\mathbf{Z}, and h ⁣:Z→Y^h\colon\mathbf{Z}\rightarrow\mathbf{\hat{Y}} subsequently decodes Z\mathbf{Z} into the sequence label Y^\mathbf{\hat{Y}}.

Fig. 2 presents an overview of our model. Importantly, our encoder gg is parameterized by a Transformer network . Our decoder hh adopts a lightweight convolutional network. To capture actions at various temporal scales, we design a multi-scale feature representation Z={Z1,Z2,…,ZL}\mathbf{Z}=\{\mathbf{Z}^{1},\mathbf{Z}^{2},\ldots,\mathbf{Z}^{L}\} forming a feature pyramid with varying resolutions. Note that our model operates on a temporal axis defined by feature grids rather than the absolute time, allowing it to adapt to videos with different frame rates. We now describe the details of our model.

Our model first encodes an input video X={x1,x2,…,xT}\mathbf{X}=\{\mathbf{x}_{1},\mathbf{x}_{2},\ldots,\mathbf{x}_{T}\} into a multiscale feature representation Z={Z1,Z2,…,ZL}\mathbf{Z}=\{\mathbf{Z}^{1},\mathbf{Z}^{2},\ldots,\mathbf{Z}^{L}\} using an encoder gg. The encoder gg consists of (1) a projection function using a convolutional network that embeds each feature (xt\mathbf{x}_{t}) into a DD-dimensional space; and (2) a Transformer network that maps the embedded features to the output feature pyramid Z\mathbf{Z}.

Projection. Our projection E\mathbf{E} is a shallow convolutional network with ReLU as the activation function, defined as

A main advantage of MSA⁡\operatorname{MSA} is the ability to integrate temporal context across the full sequence, yet such a benefit comes at the cost of computation. A vanilla MSA⁡\operatorname{MSA} has a complexity of O(T2D+D2T)O(T^{2}D+D^{2}T) in both memory and time, and is thus highly inefficient for long videos. There has been several recent work on efficient self-attention . Here we adapt the local self-attention from by limiting the attention within a local window. Our intuition is the temporal context beyond a certain range is less helpful for action localization. Such a local self-attention significantly reduces the complexity to O(W2TD+D2T)O(W^{2}TD+D^{2}T) with WW the local window size (≪T\ll T). Importantly, local self-attention is used in tandem with the multiscale feature representation Z={Z1,Z2,…,ZL}\mathbf{Z}=\{\mathbf{Z}^{1},\mathbf{Z}^{2},\ldots,\mathbf{Z}^{L}\}, using the same window size on each pyramid level. With this design, a small window size (19) on a downsampled feature map (16x) will cover a large temporal range (304).

Multiscale Transformer. We now present the design of our Transformer encoder. Our Transformer has LL Transformer layers with each layer consisting of alternating layers of local multiheaded self-attention (MSA⁡\operatorname{MSA}) and MLP⁡\operatorname{MLP} blocks. Moreover, LayerNorm (LN⁡\operatorname{LN}) is applied before every MSA⁡\operatorname{MSA} or MLP⁡\operatorname{MLP} block, and residual connection is added after every block. GELU is used for the MLP⁡\operatorname{MLP}. To capture actions at different temporal scales, a downsampling operator ↓⁡(⋅)\operatorname{\downarrow}(\cdot) is optionally attached. This is given by

The downsampling operator ↓⁡\operatorname{\downarrow} is implemented using a strided depthwise 1D convolution due to its efficiency. We use 2x downsampling for our model. Our Transformer block is shown in Fig. 2 (right). Our model further combines several Transformer blocks with downsampling in between, resulting in a feature pyramid Z={Z1,Z2,…,ZL}\mathbf{Z}=\{\mathbf{Z}^{1},\mathbf{Z}^{2},\ldots,\mathbf{Z}^{L}\}.

2 Decoding Actions in Time

Next, our model decodes the feature pyramid Z\mathbf{Z} from the encoder gg into the sequence label Y^={y^1,y^2,…,y^T}\mathbf{\hat{Y}}=\{\mathbf{\hat{y}}_{1},\mathbf{\hat{y}}_{2},\ldots,\mathbf{\hat{y}}_{T}\} using the decoder hh. Our decoder is a lightweight convolutional network with a classification and a regression head.

Regression Head. Similar to our classification head, our regression head examines every moment tt across all LL levels on the pyramid. The difference is that the regression head predicts the distances to the onset and offset of an action (dtsd^{s}_{t}, dted^{e}_{t}), only if the current time step tt lies in an action. An output regression range is pre-specified for each pyramid level. The regression head, again, is implemented using a 1D convolutional network following the same design of the classification network, except that a ReLU is attached at the end for distance estimation.

3 ActionFormer: Model Design

Putting things together, ActionFormer is conceptually simple: each feature on the feature pyramid Z\mathbf{Z} outputs an action score p(a)p(a) and the corresponding temporal boundaries (s,e)(s,e), which are then used to decode an action candidate. Notwithstanding the simplicity, we find that several key architecture designs are important to ensure a strong performance. We discuss these design choices here.

Design of the Feature Pyramid. A critical component of our model is the design of the temporal feature pyramid Z={Z1,Z2,…,ZL}\mathbf{Z}=\{\mathbf{Z}^{1},\mathbf{Z}^{2},\ldots,\mathbf{Z}^{L}\}. The design choices include (1) the number of levels within the pyramid; (2) the downsampling ratio between successive feature maps; and (3) the output regression range of each pyramid level. Inspired by the design of feature pyramid in modern object detectors (FPN and FCOS ), we simplify our design choices by using a 2x downsampling of the feature maps, and roughly enlarging the output regression range by 2 accordingly. We explore different design choices in our ablation.

Loss Function. Our model outputs (p(at),dts,dte)(p(a_{t}),d_{t}^{s},d_{t}^{e}) for every moment tt, including the probability of action categories p(at)p(a_{t}) and the distances to action boundaries (dtsd^{s}_{t}, dted^{e}_{t}). Our loss function, again following minimalist design, only has two terms: (1) Lcls\mathcal{L}_{cls} a focal loss for CC way binary classification; and (2) Lreg\mathcal{L}_{reg} a DIoU loss for distance regression. The loss is defined for each video XX as

Importantly, Lcls\mathcal{L}_{cls} uses Focal loss to recognize CC action categories. Focal loss naturally handles imbalanced samples — there are much more negative samples than positive ones. Moreover, Lreg\mathcal{L}_{reg} adopts a differentiable IoU loss . Lreg\mathcal{L}_{reg} is only enabled when the current time step contains a positive sample.

4 Implementation Details

Training. Following , we use Adam with warm-up for training. The warm-up stage is critical for model convergence and good performance, as also pointed out by . When training with variable length input, we fix the maximum input sequence length, pad or cropped the input sequences accordingly, and add proper masking for operations in the model. This is equal to training with sliding windows as in . Varying the maximum input sequence length during training has little impact to the performance, as shown in our ablation.

Inference. At inference time, we feed the full sequences into the model, as no position embeddings are used in the model. Our model takes the input video X\mathbf{X}, and outputs {(p(at),dts,dte))}\{(p(a_{t}),d_{t}^{s},d_{t}^{e}))\} for every time step tt across all pyramid levels. Each time step tt further decodes an action instance (et=t−dts,st=t+dte,p(at))(e_{t}=t-d_{t}^{s},s_{t}=t+d_{t}^{e},p(a_{t})). ete_{t} and sts_{t} are the onset and offset of the action, and p(at)p(a_{t}) is an action confidence score. The result action candidates are further processed using Soft-NMS to remove highly overlapping instances, leading to the final outputs of actions.

Network Architecture. We used 2 convolutions for projection, 7 Transformer blocks for the encoder (all using local attention and with 2x downsampling for the last 5), and separate classification and regression heads as the decoder. The regression range on each pyramid level was normalized by the stride of the features. More details are presented in the appendix C.

Experiments and Results

We now present our experiments and results. Our main results include benchmarks on THUMOS14 , ActivityNet-1.3 and EPIC-Kitchens 100 . Moreover, we provide extensive ablation studies of our model.

Evaluation Metric. For all datasets, we report the standard mean average precision (mAP) at different temporal intersection over union (tIoU) thresholds, widely used to evaluate TAL methods. tIoU is defined as the intersection over union between two temporal windows, i.e., the 1D Jaccard index. Given a tIoU threshold, mAP computes the mean of average prevision across all action categories. An average mAP is also reported by averaging across several tIoUs.

Baseline and Comparison. For our main results on THUMOS14 and ActivityNet-1.3 . We compare to a strong set of baselines, including both two-stage (e.g., G-TAD , BC-GNN , TAL-MR ) and single-stage (e.g., A2Net , GTAN , AFSD , TadTR ) methods for TAL. Our close competitors are those single-stage methods. Despite our best attempt for a fair comparison, we recognize some of our baselines used different setups (e.g., video features). Our experiment setup follows previous works . And our intention here is to compare our results to the best results previously reported.

Dataset. THUMOS14 dataset contains 413 untrimmed videos with 20 categories of actions. The dataset is divided into two subsets: validation set and test set. The validation set contains 200 videos and the test set contains 213 videos. Following the common practice , we use the validation set for training and report results on the test set.

Experiment Setup. We used two-stream I3D pretrained on Kinetics to extract the video features on THUMOS14, following . mAP@[0.3[0.3:0.10.1:0.7]0.7] was used to evaluate our model. A window size of 19 was used for local self-attention based on our ablation. Further details are described in the appendix C. To show that our method can adapt to different video features, we also consider the pre-training method from using an R(2+1)D network .

Results. Table 1 (left) summarizes the results. Our method achieves an average mAP of 66.8% ([0.3:0.1:0.7][0.3:0.1:0.7]), with an mAP of 71.0% at tIoU==0.5 and an mAP of 43.9% at tIoU==0.7, outperforming all previous methods by a large margin (+14.1% mAP at tIoU==0.5 and +12.8% mAP at tIoU==0.7). Our results stay on top of all single-stage methods, and also beats all previous two-stage methods, including the latest ones from . Note that our method significantly outperforms the concurrent work of TadTR , which also designed a Transformer model for TAL. With the combination of a simple design and a strong Transformer model, our method establishes new state of the art on THUMOS14, crossing the 65% average mAP for the first time.

2 Results on ActivityNet-1.3

Dataset. ActivityNet-1.3 is a large-scale action dataset which contains 200 activity classes and around 20,000 videos with more than 600 hours. The dataset is divided into three subsets: 10,024 videos for training, 4,926 for validation, and 5,044 for testing. Following the common practice in , we train our model on the training set and report the performance on the validation set.

Experiment Setup. We used two-stream I3D for feature extraction.Following , the extracted features were downsampled into a fixed length of 160 using linear interpolation. For evaluation, we used mAP@[0.5[0.5:0.050.05:0.95]0.95] and also reported the average mAP. A window size of 11 was used for local self-attention. Further implementation details can be found in the appendix C. Moreover, we combined external classification results from following . Similarly, we consider the pre-training method from .

Results. Table 1 (right) shows the results. With I3D features, our method reaches an average mAP of 35.6% ([0.5:0.05:0.95][0.5:0.05:0.95]), outperforming all previous methods using the same features by at least 0.6%. This boost is significant as the result is averaged across many tIoU thresholds, including those tight ones e.g., 0.95. Using the pre-training method from TSP largely improves our results (36.6% average mAP). Our model thus outperforms the best method with the same features by a major margin (+0.7%). Again, our method outperforms TadTR . Our results are worse than TCANet —a latest two-stage method using stronger SlowFast features that are not publicly available. We conjecture our method will also benefit from better features. Nonetheless, our model clearly demonstrates state-of-the-art results on this challenging dataset.

3 Results on EPIC-Kitchens 100

Dataset. EPIC-Kitchens 100 is the largest egocentric action dataset. The dataset contains 100 hours of videos from 700 sessions capturing cooking activities in different kitchens. In comparison to ActivityNet-1.3, EPIC-Kitchens 100 has less number of videos, yet many more instances per video (average 128 vs. 1.5 on ActivityNet-1.3). In comparison to THUMOS14, EPIC-Kitchens is 3 times larger in terms of video hours and more than 10 times larger in terms of action instances. These egocentric videos also include significant camera motion. This dataset thus poses new challenges for TAL.

Experiment Setup. We used a SlowFast network pre-trained on EPIC-Kitchens for feature extraction. This model is provided by . Our model was trained on the training set and evaluated on the validation set. A window size of 9 was used for local self-attention. For evaluation, we used mAP@[0.1[0.1:0.10.1:0.5]0.5] and report the average mAP following . In this dataset, an action is defined as a combination of a verb (action) and a noun (object). As this dataset was recently released, we are only able to compare our methods to BMN and G-TAD , both using the same SlowFast features provided by . Again, implementation details are described in the appendix C.

Results. Table 2 presents the results. Our method achieves an average mAP ([0.1[0.1:0.10.1:0.5]0.5]) of 23.5% and 21.9% for verb and noun, respectively. Our results again largely outperform the strong baselines of BMN and G-TAD by over 13.5% in absolute percentage points. An interesting observation is that the gaps between our results and BMN / G-TAD are much larger on EPIC-Kitchens 100. A possible reason is that ActivityNet has a small number of actions per video (1.5), leading to imbalanced classification for our model; only a few moments (around the action center) are labeled positive while all rest are negative.

We adapt ActionFormer for EPIC-Kitchens 100 2022 Action Detection challenge. By combining features from SlowFast and ViViT , ActionFormer achieves 21.36% / 20.95% average mAP for actions on the validation / test set. Our results ranked 2nd with a gap of 0.32 average mAP to the top solution.

4 Ablation Experiments

We conduct extensive ablations on THUMOS14 to understand our model design. Results are reported using I3D features with a fixed random seed for training. Further ablations on loss weight, maximum input length during training, input temporal feature resolution and error analysis can be found in the appendix A.

Baseline: A Convolutional Network. Our ablation starts by re-implementing a baseline anchor-free method (AF Base) as described in (Table 3a row 1-2). This baseline shares the same action representation as our model, yet uses a 1D convolutional network as the encoder. We roughly match the number of layers and parameters of this baseline to our model. See the appendix C for more details. This baseline achieves an average mAP of 46.6% on THUMOS14 (Table 3a row 3), outperforms the numbers reported in by 6.2%. We attribute the difference to variations in architectures and training schemes. This baseline, when using score fusion, reaches 52.9% average mAP (Table 3a row 4).

Transformer Network. Our next step is to simply replace the 1D convolutional network with our Transformer model using vanilla self-attention. This model achieves an average mAP of 62.7% (Table 3a row 5) — a major boost of 16.1%. We note that this model already outperforms the best reported results (56.9% mAP at tIoU==0.5 from ). This result shows that our Transformer model is very powerful for TAL, and serves as the main course of performance gain.

Layer Norm, Center Sampling, Position Encoding, & Score Fusion. We further add layer norm in the classification and regression heads, apply center sampling during training, and explore position encoding as well as score fusion (Table 3a row 6-9). Adding layer norm boosts the average mAP by 2.7%, and using center sampling further improves the performance by 1.4%. The commonly used position encoding, however, does not bring performance gain. We postulate that our projection using convolutions as well as the depthwise convolutions in our Transformer blocks already leak the location information, as also pointed out in . Further fusing the classification scores will decrease the largely performance. As a reference, when replacing the vanilla self-attention with a local version (window size=19), the average mAP remains the same.

Window Size for Local Self-Attention. Next, we study the effects of window size for local self-attention in our model. We vary the window size, re-train the model, and present both model accuracy, complexity (in GMACs), and run time in Table 3b. All results are reported without score fusion. Due to our design of a multiscale feature pyramid, even using the global self-attention only leads to 26% increase in MACs when compared to the baseline convolutional network. Reducing the window size cuts down the MACs yet maintains a similar accuracy. In addition to MACs, we also evaluate the normalized run time of these models on GPUs, where the base convolutional model is set to 1.0x. In spite of similar MACs, Transformer-based models are roughly 2x slower in run time compared to a convolution-based model (AF Base). It is known that self-attention is not easily parallelizable on GPUs. Also, our current implementation of Transformer used PyTorch primitives without leveraging customized CUDA kernels.

Feature Pyramid. Further, we study the design of the feature pyramid. As discussed in Sec. 3.3, our design space is specified by (1) the number of pyramid levels and (2) an initial regression range for the first pyramid level. We vary these parameters and report results in Table 3c. We follow our best design with a local window size=19 and layer norm, and center sampling.

First, we disable the feature pyramid and attach the heads to the feature map with the highest resolution. This is done by setting the number of pyramid to 1 with an initial regression range of [0, +∞+\infty), Removing the feature pyramid results in a major performance drop (-19.3% in average mAP), suggesting that using feature pyramid is critical for our model. Next, we set the number of pyramid levels to 3 and experiment with different initial regression ranges. The best results are achieved with the range of [0, 4). Further increase of the range decreases the mAP scores. Finally, we fix the initial regression range to [0, 4) and increase the number of pyramid levels. The performance of our method generally increases with more pyramid levels, yet is saturated when using 6 levels.

Result Visualization. Finally, we visualize the outputs of our model (before Soft-NMS) in Fig. 3, including the action scores, and the regression outputs weighted by the action scores (as a weighted histogram). Our model outputs a strong peak near the center of an action, potentially due to the employment of center sampling during training. The regression of action boundaries seems less accurate. We conjecture that our regression heads can be further improved.

Conclusion and Discussion

In this paper, we presented ActionFormer—a Transformer-based method for temporal action localization. ActionFormer has a simple design, falls into the category of single-stage anchor-free method, yet achieves impressive results across several major TAL benchmarks including THUMOS14, ActivityNet-1.3, and the more recent EPIC-Kitchens 100 (egocentric videos). Through our experiments, we showed that the power of ActionFormer lies in our orchestrated design, in particular the combination of local self-attention and a multiscale feature representation to model longer range temporal context in videos. We hope that our model, notwithstanding its simplicity, can shed light on the task of temporal action localization, as well as the broader field of video understanding.

Appendix

In the appendix, we describe (1) additional ablation experiments (Appendix A); (2) further error analysis of our results (Appendix B); (3) implementation details and how to reproduce our results (Appendix C); (4) additional visualizations of our results (Appendix D); and (5) limitation of our approach and furture directions (Appendix E). For sections, figures, tables, and equations, we use numbers (e.g.,, Sec. 1) to refer to the main paper and capital letters (e.g.,, Sec. A) to refer to this appendix.

Appendix A Additional Ablation Experiments

Here we present additional ablation experiments, as mentioned in Sec. 4.4 of the main paper. These are omitted from the main paper due to lack of space. All experiments are reported on THUMOS14, consistent with our ablation experiments in the main paper. We follow our best design and use a local window size=19 with layer norm, center sampling, and score fusion enabled.

Loss Weight. We provide additional ablation on the loss weight λreg\lambda_{reg} in Eq. 7. Specifically, we varied the loss weight λreg∈[0.2,0.5,1,2,5]\lambda_{reg}\in[0.2,0.5,1,2,5], retrained the model, and reported the mAP scores. The results are presented in Table A. For a large range of λreg\lambda_{reg}, our model has quite stable results with a maximum gap of 1.4% in average mAP. λreg=1\lambda_{reg}=1 yields the best results, as we used in all our experiments.

Maximum Input Sequence Length during Training. A possible explanation of our superior results is that our model might benefit from training using a long sequence (2304 time steps as in our previous experiments). Here we examine the effects of maximum input sequence length during training. Table B reports mAP scores for different training sequence lengths. The results of our model remain fairly consistent even with much shorter input sequence length. Note that when truncating an input sequence, our training scheme is equal to training with sliding windows as in . The differences are (1) the windows are dynamically sampled rather than pre-generated; (2) windows without foreground actions are removed. When using a input sequence length of 512, similar to what was considered in (512), our method only has a minor drop in average mAP (-1.1%) and significantly outperforms .

Temporal Feature Resolution. Some of the previous works considered video features with lower temporal resolution. For example, a feature stride of 8 was used by PGCN and ContextLoc . To understand the effects of temporal feature resolution, we downsample our input I3D features and study the performance variation when using different feature strides. Table C report the results. When using a lower resolution (stride=8), the results of our model only drop slightly (-0.5% in average mAP). Further reducing the resolution (e.g.,, stride=16) leads to larger performance degradation, yet our results remains favourable.

Appendix B Further Error Analyses

We present further analyses of our results on THUMOS14 using the tool provided by . We refer the readers to for more details.

Metrics. In , several characteristic metrics were defined given a dataset (e.g., THUMOS14), including coverage, length, and the number of instances. Specifically, coverage presents the relative length of the actions (compared to the whole video), categorized into five bins: Extra Small (XS: (0, 0.02]), Small (S: (0.02, 0.04]), Medium (M: (0.04, 0.06]), Large (L:(0.06, 0.08]), and Extra Large (XL: (0.08, 1.0]). Length denotes the absolute length (in seconds) of actions, organized into five length groups: Extra Small (XS: (0, 3]), Small (S: (3, 6]), Medium (M: (6, 12]), Long (L: (12, 18]), and Extra Long (XL: >> 18). Moreover, number of instances refers to the total count of instances (from the same class) in a video. This number is further divided into four parts, including Extra Small (XS: 1); Small (S: ); Medium (M: ); Large (L: >> 80).

Results and Analyses. Fig. A presents the false negative profiling. In Fig. A, we breakdown the false negative rates under the different coverage, length, and the number of instances. Our results have similar false negative rates across different coverage categories, yet have much higher false negative rates on action instances that are either very shot or very long (length), and on videos that contains many action instances (#instances). These action instances and videos are naturally more challenging.

Fig. B presents the sensitivity analysis of our results, i.e.,, normalized mAP at tIoU==0.5 under different characteristic metrics (left) and the variance of mAP across categories (right). Our model performs better on simple context scenarios, including XS/S/M/L coverage, S/M length and XS #instances, and worse on more complicated scenarios. The trend is similar to the false negative profiling in Fig. A. Moreover, our model is robust across different categories in coverage, length and #instances with small variances.

Appendix C Implementation Details

We now present implementation details including the network architecture, training and inference. Further details can be found in our code.

Network Architecture. We present our network architecture in Table D, as described in Sec. 3. In the ablation study (Sec. 4), we also considered a baseline that replaces the Transformer Units in Table D with convolution blocks, following the design of a bottleneck block in ResNet . Specifically, a stack of three 1D convolutional layers were used. The kernel size of three convolutional layers were 1, 3 and 1, respectively. The expansion factor of the bottleneck block was 2. We added an extra strided convolutional layer with kernel size=1 and stride=2 to perform downsampling when necessary.

Training Details. For training, we considered both fixed length inputs (ActivityNet) and variable length inputs (THUMOS14, ActivityNet, and EPIC-Kitchens 100). For variable length inputs, we capped the input length to 2304 (around 5 minutes on THUMOS14 and around 20 minutes on EPIC-Kitchens 100), and randomly selected a subset of consecutive clips from an input video. Position embedding was disabled by default except for ActivityNet. Model EMA and gradient clipping were also implemented to further stabilize the training. Hyperparameters were slightly different across datasets and discussed later in our experiment details.

Inference Details. For fixed length inputs (ActivityNet-1.3), we fed the full sequence into our model. For variable length inputs (THUMOS14 and EPIC-Kitchens 100), we sent the full sequence into the model. When using position embeddings in our ablation study, we adopted the technique from . Specifically, for input sequences shorter than the training sequence length (2304), we fed the full sequence into our model and clipped the position embedding using the actual length of the video. For input sequences longer than the training sequence length, we again fed the full sequence into our model, yet used linear interpolation to upsample the position embeddings.

Score Fusion. For our experiments on THUMOS14 and ActivityNet-1.3, we sometimes consider score fusion using external classification scores. Specifically, given an input video, the top-2 video-level classes given by external classification scores were assigned to all detected action instances in this video, where the action scores from our model were multiplied with the external classification scores. Each detected action instance from our model thus creates two action instances. We refer the readers to (Appendix E) for a more detailed description of the score fusion strategy.

Experiment Details. Our experiment details vary across datasets, as each dataset includes videos of different resolution and frame rate, and considers different types of features. We now describe our experiment details for THUMOS14, ActivityNet-1.3, and EPIC-Kitchens 100.

THUMOS14: We used two-stream I3D pretrained on Kinetics to extract the video features on THUMOS14, following . We fed 16 consecutive frames as the input to I3D, used a sliding window with stride 4 and extracted 1024-D features before the last fully connected layer. The two-stream features were further concatenated (2048-D) as the input to our model. mAP@[0.3[0.3:0.10.1:0.7]0.7] was used to evaluate our model. Our model was trained for 50 epochs with a linear warmup of 5 epochs. The initial learning rate was 1e-4 and a cosine learning rate decay is used. The mini-batch size was 2, and a weight decay of 1e-4 was used.

ActivityNet-1.3: We used two-stream I3D and TSP for feature extraction, and increased the stride of the sliding window to 16. Following , the extracted features were downsampled into a fixed length of 160 and 192 using linear interpolation for I3D and TSP features, respectively. For evaluation, we used mAP@[0.5[0.5:0.050.05:0.95]0.95] and also reported the average mAP. Our model was trained for 15 epochs with a linear warmup of 5 epochs. The learning rate was 1e-3, the mini-batch size was 16, and the weight decay was 1e-4. For ActivityNet, we find it is helpful to train our model to generate proposals by considering all actions from a single category, and then use external classification scores for the recognition. This strategy was also used in previous single-stage TAL methods .

EPIC-Kitchens 100: We used a SlowFast network pre-trained on EPIC-Kitchens for feature extraction. This model is provided by . We fed 32 frame window with a stride of 16 to extract 2304-D features. Our model was trained on the training set and evaluated on the validation set. A window size of 9 was used for local self-attention. For evaluation, we used mAP@[0.1[0.1:0.10.1:0.5]0.5] and report the average mAP following . Our model was trained for 30 epochs with learning rate 1e-4, mini-batch size 2, and weight decay of 1e-4.

Reproducibility of Our Results. All results reported in the paper were obtained with the same random seed using PyTorch 1.10, CUDA 10.2 and CUDNN 7.6.5 on an NVIDIA Titan Xp GPU, using deterministic GPU computing routines. On the same machine, our code will always produce the same results when using the same random seed. Across machines/GPUs and computing environments, we have observed minor variation of average mAP scores (up to 0.5% average mAP on THUMOS, less than 0.2% average mAP on ActivityNet, and under 0.8% average mAP on EPIC Kitchens), yet those minor variations do not erode the clear performance gains of our method. Our code is made publicly available.

Appendix D Additional Visualizations

Further, we present more visualizations of our results in Fig. D, extending Fig. 3 of the main paper. Our model is able to detect the occurrence of actions and estimate their temporal boundaries for the most of the cases (see the first column of Fig. D). The major failure modes of our model, as demonstrated in the second column of Fig. D, include (1) incorrect classification of action centers, i.e., background confusion (classification errors); (2) inaccurate regression of the action’s onset and offset (localization errors). We plan to address these issues in our future work.

Appendix E Limitations and Future Work

A main limitation of our method is the use of pre-extracted video features, also faced by many previous approaches. Another limitation is the need for many human labeled videos for training and the constraint of a pre-defined vocabulary of actions. Interesting future directions include pre-training for action localization , and learning from videos and text corpus without human labels.

References