Unsupervised Pre-training for Temporal Action Localization Tasks

Can Zhang, Tianyu Yang, Junwu Weng, Meng Cao, Jue Wang, Yuexian Zou

Introduction

Model pre-training is an effective technique for training deep networks in many computer vision tasks. The core idea is to learn general representations on large-scale labeled or unlabeled data, and utilize the learned representations to improve the performance of downstream tasks with limited data. This is especially beneficial for tasks that require enormous human effort to annotate data, such as temporal action localization (TAL).

Despite the prevailing use of ready-made feature extractors pre-trained on temporal action classification (TAC) in TAL, this pre-training strategy is sub-optimal as the inherent discrepancy between TAC and TAL exists. Without a doubt, this discrepancy impedes further performance improvement of TAL. Though some recent works attempt to tackle this issue, they still rely on large-scale annotated video data. Recently, unsupervised pre-training has attracted great attention due to its potentials in exploiting large amounts of unlabeled data. Contrastive learning is one of the most popular directions that focus on instance discrimination, which pulls instance-level positive pairs closer while repelling negative ones apart in the embedding space. To fill the gap between the upstream pre-training and the downstream tasks, recent contrastive learning methods focus on specifically designing pretext tasks for various downstream image tasks, e.g., object detection , semantic segmentation , etc. In contrast, the progress of unsupervised pre-training in video domain is relatively lagging behind and most existing methods are still designed and evaluated for classification tasks.

In this paper, we make the first attempt on unsupervised pre-training for TAL tasks. One possible way to achieve this is to directly extend the image contrastive learning idea to the video domain, where a video is treated as an instance and the clips are regarded as views of instances. Those clip embeddings from the same video are pulled closer while those from different videos are pushed apart. Clearly, this way only focuses on instance (video-level) discrimination, i.e., learning time-invariant features for specific video instances, which is required by TAC task in essence. In contrast, TAL expects the representations to be equivariant to temporal translation and scale. For example, if we change the start time and duration of an action instance in the input video, the output classification responses of TAC should be unchanged, while the output localization predictions of TAL need to be altered accordingly. The inherent discrepancy between these two tasks attracts our attention to question the suitability of the existing instance discrimination paradigm for TAL. Indeed, as shown in Fig. 1, such video-level discrimination is beneficial for TAC tasks, but not well-aligned with TAL tasks. So, it is desirable and challenging to design a new learning scheme that can be transferred well on TAL tasks.

Motivated by the inherent discrepancy between TAC and TAL, we introduce temporal equivariant contrastive learning paradigm by designing a new unsupervised pretext task called Pseudo Action Localization (PAL). Specifically, to mimic the TAL-tailored data with temporal boundaries, we first construct our training set by transforming the existing large-scale TAC datasets in a cheap manner. We randomly crop two temporal regions with random temporal lengths and scales from one video as pseudo actions. Each of these regions includes multiple consecutive clips. Then we paste them onto different temporal positions of other randomly selected background videos. With the preset temporal transformation (paste location, clip length, sampling scale), the model is able to align the pseudo action features of two synthesized videos. Such transformation and alignment process are named as input-level transformation and feature-level equi-transformation in our paper. Moreover, to better align the upstream pre-training pipeline to the downstream TAL architecture, we follow the way of estimating temporal locations in TAL tasks by applying several layers of temporal convolutions to process the sequential clip-level features. Thereby, the information of surrounding background clips is highly involved in the final output features of pseudo action regions. With the random paste operation, the diversity of background-involvement is increased. Further, we propose to maximize the agreement between two aligned pseudo action region features such that the learned features are forced to focus on the most discriminative and background-irrelevant parts, thus enhancing their robustness and achieving the equivariance requirement in TAL.

We summarize our main contributions as follows: (1) To our best knowledge, this is the FIRST work focusing on unsupervised pre-training for temporal action localization tasks (UP-TAL). (2) We design an intuitive and effective self-supervised pretext task customized for TAL, called PAL. A time-equivariant contrastive learning paradigm is also introduced to perform transformed foreground discrimination, customized for TAL representation learning. (3) Extensive experiments on ActivityNet v1.3 , Charades-STA and THUMOS’14 datasets show that PAL transfers well on various downstream TAL-related tasks: Temporal Action Detection (TAD), Action Proposal Generation (APG) and Video Grounding (VG). Notably, our PAL even surpasses the supervised pre-training when using the same amount of video data.

Related work

Contrastive Video Representation Learning. Recently, contrastive learning has gained increasing attention due to its outstanding performance. Essentially, these contrast-based methods focus on instance discrimination , i.e., distinguishing each instance from the rest. Following this direction, recent researches extend the contrastive learning idea to the video domain, where clips from the same video are considered as positives and clips from the different videos as negatives. Besides, other directions, such as: dense future prediction , cross-modal supervision , etc, have also been studied in the literature. Notably, most of these methods are designed for TAC tasks that learn time-invariant features. In contrast, we propose a novel pretext task tailored for TAL, which follows a temporal equivariant learning scheme. A concurrent work also focuses on time-equivariant representation learning. Two clips from different videos but with the same relative transformation (overlap/order) are considered as positive pairs, which promotes detailed learning of motion patterns and thus is beneficial for TAC tasks. Our method differs essentially from it in the fact that positive pairs are constructed from two transformed regions (multiple clips) of the same foreground video but with different backgrounds. This facilitates the learning of TAL-friendly features such that they are robust to background interference but sensitive to temporal transformation (scale and location).

Temporal Action Localization (TAL) Tasks. Unlike TAC , the target of TAL is to temporally localize the action of interest in untrimmed videos. In general, TAL covers a range of tasks, such as: Action Proposal Generation (APG), Temporal Action Detection (TAD) and Video Grounding (VG), etc. APG aims at generating temporal proposals which are likely to contain human actions. Previous methods design temporal anchor instances for feature sequences or directly predict boundary probabilities . TAD aims at predicting the temporal extent as well as the class labels of action instances. Most existing fully-supervised TAD methods integrate the proposal generation and classification procedures in a unified network. Some recent works have also designed TAD algorithms under weaker supervision . VG, a.k.a, text-to-video temporal grounding, aims to localize the time interval corresponding to a given text query. The current literature can be roughly divided into two categories, namely proposal-based and proposal-free architectures. For these TAL tasks, we choose three representative works (BMN , G-TAD and LGI ) with officially released code to validate the efficacy of our PAL.

Supervised Pre-training for TAL. Due to the GPU memory constraint, the common practice in TAL is to first pre-train a feature encoder on large-scale trimmed TAC datasets, and then use it to extract frame-level or segment-level features in untrimmed TAL videos. Inevitably, this will result in a task discrepancy problem, since feature encoders are trained on TAC while used for TAL. This domain gap has not been fully studied though it is common in TAL. Recent advances try to bridge this gap through boundary type classification , foreground region classification and end-to-end training . Unfortunately, they all belong to the supervised pre-training paradigm and therefore rely on large-scale labeled videos. In contrast, we propose a novel method, for the first time (to our best knowledge), focusing on Unsupervised Pre-training for TAL (UP-TAL).

Cut-Paste for Data Synthesis. Cut-Paste, which cuts a part of one data sample and pastes it onto another sample, is found to be a useful data augmentation strategy when facing the data shortage issue. It has been widely adopted in supervised learning for object detection , instance segmentation , and self-supervised learning for image/video classification , object detection and anomaly detection , etc. The most recent work relevant to ours is BSP , which also synthesizes videos through temporal Cut-Paste. The essential difference lies in the fact that BSP supervisedly generates different types of temporal boundaries and learn to predict them to facilitate the learning of video features while our PAL synthesizes videos without using any label information and train the backbone by aligning pseudo action region features from two synthetic videos and maximizing their agreement with temporal equivariant contrastive learning.

Method

As mentioned in Sec. 1, the most essential difference between TAC and TAL is that the former requires temporal invariance while the latter desires temporal equivariance representations. This motivates us to question the suitability of the existing “TAC features for TAL” paradigm. Thus, in this section, we delve into the design of unsupervised pre-training customized for TAL, to reach the task alignment goal, i.e., “TAL features for TAL”.

For the TAC task, given a video vi\bm{v}_{i} from a dataset V={vi}i=1N\bm{V}=\{\bm{v}_{i}\}_{i=1}^{N}, the goal is to learn a feature encoding function F(v)\mathcal{F}(\bm{v}) with which the extracted representation is ensured to be insensitive to the temporal transformation T\mathcal{T}, i.e. ∀v∈V:F(T(v))=F(v)\forall\bm{v}\in\bm{V}:\mathcal{F}(\mathcal{T}(\bm{v}))=\mathcal{F}(\bm{v}), as illustrated in Fig. 2(a) top part. To achieve this objective, the learning strategy can be basically designed as pushing F(T(v))\mathcal{F}(\mathcal{T}(\bm{v})) and F(v)\mathcal{F}(\bm{v}) close to each other in the feature space. To be more general, two random transformations T\mathcal{T} and T′\mathcal{T}^{\prime} are applied to v\bm{v} to implement the strategy, and contrastive learning is involved to enforce the consistency:

in which the identity mapping T0(v)=v\mathcal{T}_{0}(\bm{v})=\bm{v} is also considered.

Under the scenario of TAL, we require F\mathcal{F} to be sensitive to the transformation T\mathcal{T}, i.e. ∀v∈V:F(T(v))=T(F(v))\forall\bm{v}\in\bm{V}:\mathcal{F}(\mathcal{T}(\bm{v}))=\mathcal{T}(\mathcal{F}(\bm{v})), which can be re-written as F(v)=G(F(T(v)))\mathcal{F}(\bm{v})=\mathcal{G}(\mathcal{F}(\mathcal{T}(\bm{v}))) and G≜T\mathcal{G}\triangleq\mathcal{T}-1\scriptscriptstyle\text{-}\,\!1 (See Fig. 2(a) bottom part). Similar to Eqn. 1, we apply two random transformations to v\bm{v}, and therefore have

Intuitively, contrastive learning can be introduced here to model the temporal equivariance by forcing the features processed by two transformation pairs (T,G)(\mathcal{T},\mathcal{G}) and (T′,G′)(\mathcal{T}^{\prime},\mathcal{G}^{\prime}) respectively to be analogous to one another:

In the following sections, a parameterized temporal transformation T\mathcal{T} tailored for TAL tasks is introduced. We delicately design a new self-supervised task called Pseudo Action Localization (PAL) with self-generated transformation signals T\mathcal{T}, and apply the contrastive strategy to learn temporal translation and scale equivariance encoding F\mathcal{F}.

2 Pseudo Action Localization

As illustrated in Fig. 2(b), given a large-scale trimmed video dataset (e.g., Kinetics ), we randomly select two temporal regions from one video (viewed as pseudo action regions), and then paste them onto another two videos (viewed as pseudo background) at various scales and locations. With the self-generated temporal locations and scales treated as prior during pre-training, the model is expected to localize the pseudo action regions from the synthesized new videos. Instead of directly predicting the paste locations and scales, we introduce the contrastive strategy to enforce the consistency between the features of two random regions defined by the priors for temporal equivariance representation learning, as illustrated in Eqn. 3.

In this pipeline, we first perform transformation T\mathcal{T} in the input space for TAL-tailored video generation (Sec. 3.2.1). Then we use backbone F\mathcal{F} and multiple heads to map the transformed videos into the feature space (Sec. 3.2.2). Next, equi-transformation G\mathcal{G} is applied to inverse the transformation T\mathcal{T} in the feature space (Sec. 3.2.3). We finally conduct region contrastive learning for TAL-customized pre-training (Sec. 3.2.4).

To learn the temporal equivariance encoding function F\mathcal{F}, we define the transformation T\mathcal{T} as a video region sampling and paste operation. Specifically, given a video vi\bm{v}_{i} as pseudo action video as well as a randomly selected video vn\bm{v}_{n} as pseudo background, we first sample a random region from action video vi\bm{v}_{i}, and paste it onto the background video vn\bm{v}_{n} to generate a synthesized video vi→n\bm{v}_{i\rightarrow n}. The input-level transformation T\mathcal{T} is then defined as follows:

where ss and ee represent the start and end clipHere we perform the temporal transformation in a clip-wise manner to align with the clip-level video encoder. indices of the pseudo action region in the new video vi→n\bm{v}_{i\rightarrow n}.

To improve the robustness of the learned representation, we soften the paste operation in the implementation by changing it to a blending one with blending ratio β\beta, which takes β\beta of action region and mix it with (1−β)(1-\beta) of background region to generate the blended region. β\beta is randomized from the range [0.6,1][0.6,1]. Besides, spatial data augmentation is involved to increase the diversity of training data. Following the convention , we apply random cropping, horizontal flipping, Gaussian blurring and color jittering, and all are temporally consistent. In particular, instead of sampling action regions with fixed stride, we propose a scale-aware sampling strategy to add some randomness to the action timescale. Here, we refer to the timescale as how fast an action goes. It stems from the observation that an action video played at different speeds contains almost identical semantics. We simply model the timescale variation by sampling action region frames with different strides.

Overall, by this sample-and-paste way, our input-level transformation mimics the temporal location and scale variance in real untrimmed action videos, which also provides a strong supervision signal for TAL-tailored temporal equivariant contrastive learning.

2.2 Feature Encoding

Our feature encoder F\mathcal{F} contains a backbone Fb\mathcal{F}^{b} with non-linear projection head Fn\mathcal{F}^{n} and a temporal embedding head Ft\mathcal{F}^{t}, namely F=Fb∘Fn∘Ft\mathcal{F}=\mathcal{F}^{b}\circ\mathcal{F}^{n}\circ\mathcal{F}^{t}. Formally, given the synthesized video vi→n\bm{v}_{i\rightarrow n}, the corresponding clip feature sequence {ci→n(j)}j=1J\{\bm{c}_{i\rightarrow n}^{(j)}\}_{j=1}^{J} is obtained by:

in which the backbone Fb\mathcal{F}^{b} is a clip-level encoder, and Ft\mathcal{F}^{t} is a video-level head for temporal modeling among clips. JJ is the number of sampled clips. It is noted that applying temporal convolutions (Ft\mathcal{F}^{t}) on chronological clip-level features is crucial in our setting. This enables information aggregation among neighboring clips and therefore the features of pseudo action regions near the boundary can be highly affected by the nearby background. In this way, our PAL can learn background-insensitive boundary features by maximizing the agreement (Sec. 3.2.4) between region features of the same video but influenced by different pseudo backgrounds.

2.3 Feature-Level Equi-Transformation

Recall that we aim at designing a TAL-tailored pre-training paradigm by extending the contrastive strategy to learn temporal equivariance representations. To this end, we propose to utilize additional free region-wise supervision in the form of inverse temporal transformations. In our case, the changes of pseudo action location in the input composited videos (vi→n\bm{v}_{i\rightarrow n}) will be reflected in the corresponding ones in their feature sequences ({ci→n(j)}j=1J\{\bm{c}_{i\rightarrow n}^{(j)}\}_{j=1}^{J}) obtained by Eqn. 5.

To echo the input-level transformation T\mathcal{T} introduced in Sec. 3.2.1, we here define the feature-level equi-transformation G\mathcal{G} as an alignment operation. Formally, this feature alignment process is defined as:

then the region representation can be obtained by temporally averaging pooling the corresponding sequential clip-level features, i.e., ri→n(s,e)=TempAvgPool({ci→n(j)}j=se)\bm{r}_{i\rightarrow n}^{(s,e)}={\rm TempAvgPool}(\{\bm{c}_{i\rightarrow n}^{(j)}\}_{j=s}^{e}).

2.4 Contrastive Training Objective

Following the transformation T\mathcal{T} and alignment G\mathcal{G} operations introduced above, two pseudo action regions [s,e][s,e] and [s′,e′][s^{\prime},e^{\prime}] from video vi\bm{v}_{i} are extracted, and pasted onto two pseudo background videos vn\bm{v}_{n} and vm\bm{v}_{m} to obtain region representations ri→n(s,e)\bm{r}_{i\rightarrow n}^{(s,e)} and ri→m(s′,e′)\bm{r}_{i\rightarrow m}^{(s^{\prime},e^{\prime})}. These two representations are set as the query and positive key pair (rq,rk+)(\bm{r}_{q},\bm{r}_{k^{+}}) in contrastive learning, namely rq=ri→n(s,e)\bm{r}_{q}=\bm{r}_{i\rightarrow n}^{(s,e)} and rk+=ri→m(s′,e′)\bm{r}_{k^{+}}=\bm{r}_{i\rightarrow m}^{(s^{\prime},e^{\prime})}. The region features from other composited videos are viewed as negatives. Given the encoded query rq\bm{r}_{q}, positive key rk+\bm{r}_{k^{+}} and negatives {rki}i=1K\{\bm{r}_{k_{i}}\}_{i=1}^{K}, the contrastive learning essentially encourages the query to be similar to the positive sample and dissimilar to the negative ones. Our PAL is a pretext task and independent of the detailed loss function, so we simply extend the InfoNCE contrastive loss to ensure region consistency in this work:

where τ\tau is a temperature hyper-parameter and KK is the number of negative samples. Equipped with our proposed PAL, by minimizing the region contrastive loss, the encoding backbone Fb\mathcal{F}^{b} is encouraged to learn temporal equivariant features, which we believe is beneficial to the TAL tasks.

Experiments

To evaluate our proposed PAL, we follow the pre-training and transferring procedures: first pre-train the feature network on a large-scale trimmed dataset without category labels, then transfer the features pre-computed by the frozen backbone to the downstream TAL tasks.

Datasets. To have an apple-to-apple comparison with other self-supervised video representation learning methods, we use Kinetics as the initial pre-training dataset, without using any labels. Kinetics is a large-scale trimmed action recognition benchmark. Each video has a single action class and lasts around 10 seconds. The typical version Kinetics-400 (K400) includes ∼\sim300k videos with 400 human action classes, and the latest version Kinetics-700 (K700) contains ∼\sim650k videos with 700 action classes.

Implementation Details. We choose I3D , the commonly used feature encoder in TAL, as the default backbone (Fb\mathcal{F}^{b}) in our experiments. For temporal embedding head (Ft\mathcal{F}^{t}), we employ two-layer temporal convolutions with a kernel size of 3 followed by ReLU activation function. We uniformly sample 8 clips (8 frames per clip) with a resolution of 112×112112\times 112 for each video, and the max clip length of the pseudo action region is limited to 6. The range of the blending ratio β\beta is set as [0.6,1.0][0.6,1.0]. For our scale-aware sampling strategy, the sampling stride of frames within a clip is chosen from $.Following,wealsomaintainamemoryqueueof16,384negativesamplesandusesynchronizedBNacrossalllayers.WeapplyL2normtotheoutputfeaturesfrom. Following , we also maintain a memory queue of 16,384 negative samples and use synchronized BN across all layers. We apply L2 norm to the output features from\mathcal{F}^{t}.Thetemperature. The temperature\tauissetto0.07forallexperiments.Foroptimization,wetrainourPALusingtheAdamalgorithmwithaweightdecayofis set to 0.07 for all experiments. For optimization, we train our PAL using the Adam algorithm with a weight decay of10^{-5}.Theinitiallearningrateissetas. The initial learning rate is set as10^{-4}$ and decreases by a factor of 10 when the validation loss saturates. The training takes 200 epochs in total with a batch size of 512 on 64 NVIDIA Tesla V100 GPUs.

1.2 Transferring to TAL tasks

Target TAL Tasks. We choose three popular temporal localization tasks to evaluate our PAL features: Temporal Action Detection (TAD), Action Proposal Generation (APG) and Video Grounding (VG).

Datasets. (1) ActivityNet v1.3 is a popular large-scale benchmark for TAD and APG tasks, including 10,024 training videos, 4,926 validation videos corresponded to 200 action classes. Each video contains 1.65 action instances on average; (2) Charades-STA is commonly used for VG task, containing 12,408 and 3,720 text query pairs in training and test set, respectively. The average duration of videos is 30 seconds and the maximum length of a text query is 10; (3) THUMOS’14 is a standard benchmark for TAD and APG tasks, containing 200 validation videos and 213 test videos of 20 action categories. The video length varies greatly, from less than a second to about 26 minutes. On average, each video contains ∼\sim16 action instances.

Evaluation Metrics. We follow the standard evaluation protocol. For the TAD task, we report mean Average Precision (mAP) values under different temporal Intersection over Union (tIoU) thresholds. For the APG task, we report the Area Under the Curve (AUC) of the average recall vs. average number (AR-AN) of proposals per video. For the VG task, the top-1 recall at three tIoU thresholds and their mean value (mIoU) are reported.

Implementation Details. To validate the efficacy of our pre-training strategy, we retrain several state-of-the-art TAL methods by only replacing the original features with our PAL features. We choose those representative works with publicly-available codes. Specifically, we choose G-TAD for TAD task, BMN for APG task, and LGI for VG task.

2 Main Results

In this section, we compare the performance of our PAL with other state-of-the-art pre-training approaches on three challenging TAL tasks. For those self-supervised methods designed for TAC tasks, we directly use their released pre-trained models to extract the video features for the downstream TAL task evaluations.

Temporal Action Detection (TAD) & Action Proposal Generation (APG). In Table 1, we report our TAD and APG results on ActivityNet v1.3 and compare them with state-of-the-art pre-training methods. When pre-trained on K400, our PAL consistently outperforms other self-supervised methods, which strongly demonstrates the effectiveness of our method. Although these self-supervised pre-training competitors have achieved promising results on TAC tasks, the task discrepancy issue still harms their transferability on TAL tasks, which verifies the necessity of our work. Compared to our baseline MoCo-v2 , which focuses on learning temporal invariant features, our proposed temporal equivariant learning scheme is more suitable for TAL, so it yields an improvement of +3.1% mAP@AVG and +2.8% AUC gains under the same settings. Notably, when using the same backbone (I3D) and pre-training dataset (K400), our unsupervised PAL even surpasses the supervised counterpart TAC by gains of +0.9% on mAP@AVG and +1.2% on AUC. It suggests that the proper use of data may benefit more than action label annotation information in TAL, which is consistent with our motivation. When pre-trained on a larger dataset K700, our PAL further improves the performance, showing its potential benefit of leveraging large-scale web videos. Compared to recent fully-supervised pre-training methods including BSP , LoFi-E2E and TSP , our first attempt on unsupervised TAL pre-training achieves competitive results. Note that both LoFi-E2E and TSP use the downstream dataset ActivityNet (ANet) for feature pre-training which can lead to unfair comparison.

Video Grounding (VG). The VG results on Charades-STA are reported in Table 2. Note that the original LGI exploits I3D features fine-tuned on the downstream Charades-STA dataset. For fair comparison, we retrain the LGI model using K400 pre-trained I3D features, without changing any hyper-parameters in the original codebase. Clearly, our PAL achieves the best VG performance under the unsupervised pre-training setting and even surpasses the supervised TAC trained features. Note that BSP feature, which is pre-trained in a supervised manner and has much more per-clip FLOPs (33G vs. 3.6G), performs better than ours as expected.

Overall, to our best knowledge, as the first unsupervised pre-training work customized for TAL, PAL consistently surpasses other unsupervised pre-training methods on three typical TAL tasks, demonstrating the efficacy of our idea.

3 Ablation Study

In this section, we conduct ablation experiments to fully understand the concept of our PAL. For ease of experimentation, all the ablation studies are conducted with 100 training epochs on K400 and evaluated on the TAD task.

Effectiveness of the key PAL components. In Table 3, we examine how each design in PAL affects the overall performance. We consider three key components in PAL: (1) dense sampling strategy to select multiple clips as region sample; (2) scale-aware sampling strategy to sample pseudo action regions with different strides; (3) paste operation to paste the selected regions onto the background videos. We start with the basic setting that does not involve any of the above designs, where only one clip is randomly sampled from each video and clip-level contrastive learning is performed. Then we introduce the dense sampling strategy to contrast region-level embeddings, which brings +0.5% improvement due to more temporal clues being included. Next, a scale-aware sampling strategy is applied along with dense sampling, leading to an overall +1.1% gain. This verifies that adding some randomness to the temporal scale facilitates representation learning. The biggest improvement is achieved after introducing paste operation. We infer that this is because attracting the region features of the same video but influenced by different backgrounds yields more background-insensitive features, which benefits localization tasks. Finally, pasting pseudo action regions onto different background videos further contributes a reasonable gain of +0.7% and the final improvement reaches +3.2% compared with baseline. In summary, by adding these key components step by step, the performance consistently boosts, verifying the effectiveness of our PAL.

Number of the temporal embedding head layers. In Table 4.3, we experiment with the different numbers of the temporal embedding head layers. When the layer number is 0, the input clips are processed independently without temporal fusion, and therefore the surrounding background clips have no substantial effect on the action regions. It’s obvious that attaching a single temporal convolution layer atop the backbone significantly boosts the performance (+1.3%). This verifies our hypothesis that introducing background semantics can help promote the localization power required by TAL. Since using two temporal embedding layers yields the best performance, we choose this setting as our default.

Hard paste vs. Soft paste. In PAL, β\beta controls the paste ratio of pseudo action regions onto the background videos. We evaluate different β\beta from 0 to 1. In particular, β=1\beta=1 means the “hard” paste and β<1\beta<1 is the “soft” paste. For the soft way, we test several representative intervals: [0,0.4][0,0.4], [0.4,0.6][0.4,0.6], [0.6,1.0][0.6,1.0] and [0,1.0][0,1.0], which indicates four cases, i.e., background-dominated, half-and-half, action-dominated and purely random, respectively. As shown in Table 4.3, the β∈[0.6,1.0]\beta\in[0.6,1.0] setting outperforms hard paste and achieves the best result, partially because the action-dominated soft paste serves as an effective data augmentation strategy. So we use this setting by default.

Evaluation on THUMOS’14. THUMOS’14 is a relatively small-scale dataset compared with ActivityNet v1.3 (cf. Sec. 4.1.2). We list the experimental results in Table 6. Compared with the relative performance gain on ActivityNet v1.3, our improvement on THUMOS’14 is more prominent, which confirms the generalization ability of PAL under the small-scale data condition.

Evaluation on TAC task. We investigate the transferring ability of our PAL on TAC downstream task. Following the common practice, all layers are fine-tuned end-to-end. The results are evaluated on UCF101 and HMDB51 datasets. Although our PAL feature is designed for TAL tasks, we observe in Table 7 that it still achieves competitive performance on TAC tasks. In detail, our PAL outperforms baseline MoCo-v2 by +2.7% & +3.1% at top-1 accuracy on UCF101 and HMDB51 respectively. Notably, it even exceeds the recently proposed VTHCL with the same backbone.

4 Feature Visualization

Recall that PAL is proposed to guide the network in learning temporal translation and scale equivariance ability. To confirm this, we apply temporal transformations on the action instances in real-world videos and investigate whether these changes will be reflected in feature space accordingly. Specifically, given a video from ActivityNet v1.3, we first crop the action instance based on the temporal annotations, then re-sample the action instance with different temporal strides and insert them back into random temporal locations. Here, we consider two temporal transformations: (1) 2×\times down-sampling the action instance and moving backward along the time axis; and (2) 2×\times up-sampling the action instance and moving forward along the time axis. Next, MoCo-v2 (baseline), VideoMoCo and our PAL encoders are applied to extract features for the original video and the two transformed videos.

We visualize the cosine similarity between each clip features pair within the same video in Fig. 3. We also plot the ground-truth annotations (green bars) to indicate the action clips. As can be seen, MoCo-v2 and VideoMoCo learn time-invariant features that are insensitive to the temporal transformations. There is a high similarity between the pseudo action region and the background, while the salient area of PAL features changes accordingly. This confirms that our method successfully learns the time-equivariant characteristic, which is naturally more beneficial for TAL tasks. Besides, our introduced time-equivariant learning scheme can not only better separate the action and background clips, but also enable sharper contrast between action and its surrounding background clips. In this way, the clip features become more informative and boundary-aware, which facilitates localization. More visualization results can be found in our supplementary materials.

Discussion and Conclusion

This paper presents a new pretext task called Pseudo Action Localization (PAL), which is delicately designed to pre-train representations in an unsupervised manner for TAL tasks (UP-TAL). Motivated by the essential discrepancy between TAC and TAL, we also introduce a temporal scale and location equivariance learning scheme to facilitate better task alignment for the downstream transferring process. On a variety of downstream TAL tasks including temporal action detection, action proposal generation and video grounding, we demonstrate the effectiveness of our proposed method, which consistently surpasses its TAC counterpart and other unsupervised pre-training methods.

Acknowledgements. This paper was partially supported by the National Natural Science Foundation of China (NSFC) 62176008. Special acknowledgements are given to Aoto-PKUSZ Joint Lab for its support.

References