Learning by Aligning Videos in Time
Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed, Andrey Konin, Muhammad Zeeshan Zia, Quoc-Huy Tran
Introduction
There are just three problems in computer vision: registration, registration, and registration.
Lukas-Kanade and Iterative Closest Point have been amongst the most ubiquitous building blocks in artificial perception literature. Yet spatio-temporal registration has received little attention in the present deep learning renaissance. Correspondingly, we add to a small number of recent approaches that have revived temporal alignment as a means of improving video representation learning. In order to learn perfect alignment of two videos, a learning algorithm must be able to disentangle phases of the activity in time while simultaneously associating visually similar frames in the two different videos. We demonstrate that learning in this manner generates representations that are effective for downstream tasks that rely on fine-grained temporal features.
In the context of using temporal alignment for learning video representations, some recent works use cycle-consistency losses to perform local alignment between individual frames. At the same time, some works have explored global alignment for video classification and segmentation . We adapt such global alignment ideas for video representation learning in this work.
A few of approaches have been proposed for supervised action recognition and action segmentation . Unfortunately, these approaches require fine-grained annotations which can be prohibitively expensive . We note the seemingly infinite supply of public video data, and contrast it with the high cost of fine-grained annotation. This discrepancy emphasizes the importance of exploring self-supervised methods. We are further motivated by datasets and downstream tasks that specifically benefit from temporal alignment, such as video streams of semi-repetitive activities from manufacturing assembly lines to surgery rooms. It is desirable to measure the variability and anomalies across such datasets, where representations that optimize for temporal alignment may be highly performant.
Our approach, Learning by Aligning Videos (LAV), utilizes the task of temporally aligning videos for learning self-supervised video representations. Specifically, we use a differentiable version of an alignment metric which has been widely used in the time series literature, namely Dynamic Time Warping (DTW) . DTW is a global alignment metric, taking into account entire sequences while aligning. Unfortunately, in a self-supervised representation learning context, optimizing solely for DTW may converge to trivial solutions wherein the learned representations are not meaningful. To address this issue, we combine the above alignment metric with a regularization, as shown in Fig. 1. In particular, we propose a regularization term that optimizes for temporally disentangled representations, i.e., frames that are close in time are mapped to spatially nearby points in the embedding space and vice versa.
We introduce a novel self-supervised method for learning video representations by temporally aligning videos as a whole, leveraging both frame-level and video-level cues.
We adopt the classical DTW as our temporal alignment loss, while proposing a new temporal regularization. The two components have mutual benefits, i.e., the latter prevents trivial solutions, whereas the former leads to better performance.
Our approach performs on par with or better than the state-of-the-art on various temporal understanding tasks on Pouring, Penn Action, and IKEA ASM datasets. The best performance is sometimes achieved by combining our method with a recent work . Further, our approach offers significant accuracy gain when lacking labeled data.
We manually annotate dense per-frame labels for 2123 videos of Penn Action.
Related Work
In this section, we review recent literature in self-supervised learning with a focus on image and video data.
Image-Based Self-Supervised Representation Learning. Early self-supervised representation learning methods explore image content as supervision signals. They propose pretext tasks based on artificial image cues as labels and train deep networks for solving those tasks . These pretext tasks include objectives such as image colorization , object counting , solving jigsaw puzzles , and predicting image rotations . Even earlier approaches learn representations simply by reconstructing the input image or recovering it from noise . In this work, we focus on self-supervised representation learning from videos, which leverages both spatial and temporal information in videos.
Video-Based Self-Supervised Representation Learning. With the advent of deep architectures for video understanding , various pretext tasks have been introduced as supervision signals for self-supervised representation learning from videos. One popular class of methods learn representations by predicting future frames or forecasting their encoding features . Another group of methods leverage temporal information, for example, temporal order and temporal coherence are used as labels in and respectively. Recently, Donglai et al. train a deep model for classifying temporal direction, while Sermanet et al. learn representations via consistency across different viewpoints and neighboring frames. The above methods usually optimize over a single video at a time, whereas our approach jointly optimizes over a pair of videos at once, potentially extracting more information from both videos.
Temporal Video Alignment. There exists a lot of literature on time series alignment, yet only a few ideas have been carried over to aligning videos. Unfortunately, traditional methods for time series alignment, for example, DTW , are not differentiable and hence can not be directly used for training neural networks. To address this weakness, a smooth approximation of DTW, namely Soft-DTW, is introduced in . More recently, Soft-DTW formulations have been used in a weakly supervised setting for aligning a video to a transcript or in a few-shot supervised setting for aligning videos . In the present paper, we adapt Soft-DTW for learning self-supervised representations from videos, using temporal video alignment as the pretext task. The closest work to ours is Temporal Cycle Consistency (TCC) , which learns self-supervised representations by finding frame correspondences across videos. While TCC aligns each frame separately, our approach aligns the video as a whole, leveraging both frame-level and video-level cues.
Our Approach
In this section, we discuss our main contribution which is a self-supervised method to learn video representations via temporal video alignment. Specifically, we learn an embedding space where two videos with similar contents can be conveniently aligned in time. We first aim to optimize the embedding space solely for the global alignment cost between the two videos, which can lead to trivial solutions. To overcome this problem, we regularize the embedding space such that for each input video, temporally close frames are mapped to nearby points in the embedding space, whereas temporally distant frames are correspondingly mapped far away in the embedding space. Fig. 2 shows an overview of our loss and regularization (right) and our encoder (left). Below we first define some notations and then provide the details of our temporal alignment loss, temporal regularization, final loss, and encoder network in Secs. 3.1, 3.2, 3.3, and 3.4 respectively.
Notations. We denote the embedding function as , namely a neural network with parameters . Our method takes as input two videos and , where and are the numbers of frames in and respectively. For a frame in and in , the embedding frames of and are written as and respectively. In addition, we denote and as the embedding videos of and respectively.
We adopt the classical DTW discrepancy as our temporal alignment loss. DTW has been widely used with non-visual data, such as time series, and has just recently been applied to video data, but in a weakly supervised setup for video-to-transcript alignment or in a few-shot supervised setup for video alignment . Unlike , we explore the use of DTW for self-supervised video representation learning by leveraging temporal video alignment as the pretext task.
Here, is the set of all possible (binary) alignment matrices, which correspond to paths from the top-left corner of to the bottom-right corner of using only moves. is a typical alignment matrix, with if in is aligned with in . DTW can be computed using dynamic programming, particularly solving the below cumulative distance function:
Due to the non-differentiable operator, DTW is not differentiable and unstable when used in an optimization framework. We therefore employ a continuous relaxation version of DTW, namely Soft-DTW, proposed by . In particular, Soft-DTW replaces the discrete operator in DTW by the smoothed one, defined as:
where is a smoothing parameter. Soft-DTW returns the alignment cost between and by finding the soft-minimum cost path in , which can be written as:
Note that since the smoothed operator converges to the discrete one when approaches 0, Soft-DTW produces similar results as DTW when is near 0. In addition, although using does not make the objective convex, it does help the optimization by enabling smooth gradients and providing better optimization landscapes.
2 Temporal Regularization
Since (Soft-)DTW measures the (soft-)minimum cost path in , optimizing for (Soft-)DTW alone can result in trivial solutions, wherein all the entries in are close to 0, as we will show later in Sec. 5.1. In other words, all the frames in and are mapped to a small cluster in the embedding space. To avoid that, we opt to add a temporal regularization, which is applied separately on and . Below we discuss our regularization for only, while the same one can be applied for .
Motivated by , we adapt Inverse Difference Moment (IDM) as our regularization, which can be written as:
However, we notice one problem with the above IDM regularization, in particular, it treats temporally close and far way frames in similar ways. In Eq. 5, it maximizes similarities between temporally far away frames, though with smaller weights. Similarly, for Eq. 6, it still maximizes distances between temporally close frames, though with smaller weights. To address that, we propose separate terms for temporally close and far away frames. Specifically, we introduce a contrastive version of Eq. 6, which we call Contrastive-IDM, as our regularization:
Here, is a window size for separating temporally far away frames ( = 1 or negative pairs) and temporally close frames ( = 0 or positive pairs) and is a margin parameter. Contrastive-IDM encourages temporally close frames (positive pairs) to be nearby in the embedding space, while penalizing temporally far away frames (negative pairs) when the distance between them is smaller than margin in the embedding space. Note that, if we drop the weights and in Eq. 7, it becomes equivalent to Slow Feature Analysis (SFA), also referred to as temporal coherence , which treats all pairs equally. We would emphasize that, leveraging temporal information by adding weights to different pairs based on their temporal gaps leads to performance gain, as we will show in Sec. 5.1.
3 Final Loss
Our final loss is a combination of Soft-DTW alignment loss in Eq. 4 and Contrastive-IDM regularization in Eq. 7:
Here, is the weight for the regularization. The final loss encourages embedding videos to have minimum alignment costs while encouraging discrepancies among embedding frames. Both the alignment loss and the regularization are differentiable and can be optimized using backpropagation.
4 Encoder Network
We use ResNet-50 as our backbone network and extract features from the output of the layer. The extracted features have dimensions of . We then stack context frame features along the temporal dimension for each frame. Next, the combined features are passed through two 3D convolutional layers for aggregating temporal information. It is then followed by a 3D global max pooling layer, two fully-connected layers, and a linear projection layer to output embedding frames, with each having 128 dimensions. We resize input video frames to before feeding to our encoder network.
Datasets, Annotations, and Metrics
Datasets and Annotations. We use three datasets, namely Pouring , Penn Action , and IKEA ASM . While Pouring videos capture human hands interacting with objects, Penn Action and IKEA ASM videos show humans playing sports and assembling furniture respectively. We manually annotate dense frame-wise labels (i.e., key events and phases) for Penn Action using the same protocol of , since the authors of do not release them. See Fig. 3 for an example. For Pouring and IKEA ASM, we obtain the labels from the authors of and respectively. Actions/videos in IKEA ASM (17 phases) are more complicated/longer than those in Pouring (5 phases) and Penn Action (2-6 phases). We use the training/validation splits from the original datasets. For Pouring, we use all videos (70 for training, 14 for validation). Following , we use 13 actions of Penn Action (for each action, 40-134 videos for training, 42-116 videos for validation). For IKEA ASM, we use all Kallax_Drawer_Shelf videos (61 for training, 29 for validation).
Evaluation Metrics. We use four evaluation metrics computed on the validation set. The network is first trained on the training set and then frozen. Next, an SVM classifier or linear regressor is trained on top of the frozen network features (without any fine-tuning of the network). For all metrics, a high score means a better model. We summarize the metrics below:
Phase Classification: is the average per-frame phase classification accuracy, implemented by training an SVM classifier on top of the frozen network features to predict the phase labels.
Phase Progression : measures the prowess of representations learnt to predict action progress temporally, implemented by training a linear regressor on top of the frozen network features to predict the phase progression values (defined using the key event labels).
Kendall’s Tau : measures how well videos are aligned temporally if we use nearest neighbor matching. It does not require any labels for evaluation.
Average Precision: is the fine-grained frame retrieval accuracy, computed as the ratio of the retrieved frames with the same phase labels as the query frame.
We follow to use the first three metrics above, while we add the last metric for our fine-grained frame retrieval experiments in Sec. 5.4. Phase Progression and Kendall’s Tau assume no repetitive frames/labels in a video.
Experiments
In this section, we benchmark our approach (namely LAV, short for Learning by Aligning Videos) against state-of-the-art methods for video-based self-supervised representation learning on various temporal understanding tasks on Pouring, Penn Action, and IKEA ASM datasets.
Implementation Details. We use the same encoder in Sec. 3.4 for all methods for Pouring and Penn Action experiments. For IKEA ASM experiments, since the actions are more complex, we opt to extract features from the output of the layer (instead of ) for all methods. We initialize ResNet-50 layers with pre-trained weights for ImageNet classification, while remaining layers are initialized randomly. We L2-normalize the frame-embeddings before feeding them to our loss (LAV). We use ADAM optimization with a learning rate of and a weight decay of . We minimize our final loss in Eq. 8 computed over all video pairs in the training set. We randomly pair videos of the same action, regardless of their viewpoints. For datasets with a single action (e.g., Pouring and IKEA ASM), videos are randomly paired. For datasets with many actions (e.g., Penn Action), videos of the same action are randomly paired. For each video pair, we calculate the final loss using sampled frames from each video, i.e., we divide a video into uniform chunks and randomly sample one frame per chunk. We implement our network and loss in PyTorch . For more details, please refer to supplementary materials.
Competing Methods. Below are the competing methods:
Self-Supervised Learning: We compare LAV with recent self-supervised video representation learning methods, namely SAL , TCN , and TCC.
Fully-Supervised Learning: We test LAV against a fully-supervised method with explicit supervision. Specifically, following , we train a network on the downstream task by attaching a 1-layer classifier to the encoder in Sec. 3.4.
Random/ImageNet Features: For completeness, we include the results obtained by using random features or pre-trained features for ImageNet classification.
Here, we perform ablation studies on Pouring dataset to show the effectiveness of our design choices in Sec. 3.
Performance of Individual Losses. We first study the performance of individual components of our approach, i.e., Soft-DTW and Contrastive-IDM, as separate baselines. Also, we include other methods such as IDM in Eq. 6 and an SFA approach proposed in . Tab. 1 (top) presents the quantitative results. We observe that the model trained with Soft-DTW alone achieves the lowest accuracy across all metrics. In fact, it has similar classification accuracy to that of random features in Tab. 2 (i.e., vs. ), which shows that training solely with Soft-DTW yields trivial solutions and the network is unable to learn any useful representations. This is also confirmed by plotting the distance matrix between the embedding frames learned solely with Soft-DTW in Fig. 4(a), where all entries are near zero. In other words, the frames are mapped to a small cluster in the embedding space. Moreover, it can be seen from Tab. 1 (top) that Contrastive-IDM outperforms IDM by significant margins on all metrics (e.g., for Kendall’s Tau, vs. ), showing the advantage of using separate terms for temporally close and far away frame pairs. Lastly, although SFA and Contrastive-IDM have competitive performances on classification and progression, Contrastive-IDM outperforms SFA significantly on Kendall’s Tau (i.e., vs. ), supporting our idea of adding weights to different frame pairs based on their temporal gaps.
Performance of Combined Losses. We now study the impact of adding IDM, SFA, or Contrastive-IDM as regularization to Soft-DTW. Tab. 1 (bottom) presents the quantitative results. From the results, the addition of regularization boosts the performance of Soft-DTW significantly across all metrics (e.g., for progression, for Soft-DTW vs. for Soft-DTW+Contrastive-IDM). More importantly, utilizing our proposed Contrastive-IDM as regularization leads to the best performance across all metrics, outperforming using IDM or SFA as regularization by significant margins, especially on progression and Kendall’s Tau (e.g., for progression, for Soft-DTW+Contrastive-IDM vs. and for Soft-DTW+IDM and Soft-DTW+SFA respectively). This validates our ideas of separating temporally close and far away frame pairs, as well as leveraging temporal gaps to weight the frame pairs accordingly. We also visualize the distance matrix between the embedding frames learned with Soft-DTW+Contrastive-IDM in Fig. 4(b), where entries have diverse values. Below, we use Soft-DTW+Contrastive-IDM as our method (LAV).
2 Phase Classification Results
In this section, we evaluate the utility of our representations for action phase classification. Tab. 2 presents the quantitative results of all methods on Pouring, Penn Action, and IKEA ASM datasets. For Penn Action experiments, we follow to train 13 different models (i.e., 1 encoder + 1 SVM classifier, for each action) and report the average results across all actions. It can be seen from Tab. 2 that our method (LAV) outperforms other self-supervised video representation learning methods, namely SAL , TCN , and TCC , on all datasets. This shows that LAV is more capable of learning useful features that allow good classification performance when combined with a relatively simple classifier. Moreover, the best accuracy on Pouring and IKEA ASM is achieved by the combined LAV+TCC, which is similar to the observation in , where combining multiple losses leads to better classification performance. Next, the relative gaps between LAV and other self-supervised methods are the largest on IKEA ASM, which has more complex actions than Pouring and Penn Action. This implies that LAV is more capable of handling complex actions. Finally, compared to the fully-supervised baseline, self-supervision with LAV provides a significant performance boost in the low labeled data regimes. Specifically, with just 10% labeled data, LAV achieves very similar performance to the fully-supervised baseline trained with 100% labeled data (e.g., on Penn Action, vs. ).
Few-Shot Phase Classification Results. Following the above observation, we consider the application of our representations in a few-shot learning setting, i.e., there are many training videos, but only a few of them have frame-wise labels. We use the same setup as the above experiment, and compare our approach with other self-supervised methods and the fully-supervised baseline. For learning self-supervised features, all training videos are used, whereas the fully-supervised baseline is trained with a few labeled videos. Specifically, we study the classification performance with increasing the number of labeled videos. The results for two actions of Penn Action are reported in Fig. 5. Although all self-supervised methods offer a significant performance boost in the low labeled data settings, LAV provides the largest gain. Moreover, self-supervision using LAV with only 1 labeled video performs similarly to the fully-supervised baseline trained with the whole dataset. For instance, on Bowling, with just 1 labeled video, LAV achieves 71%, whereas the fully-supervised baseline trained with the entire dataset (134 labeled videos) obtains 77%.
3 Phase Progression and Kendall’s Tau Results
We now evaluate the performance of our approach on action phase progression and Kendall’s Tau. Tab. 3 presents the quantitative results of different self-supervised methods on Pouring and Penn Action. We do not evaluate on IKEA ASM, since its labels are repeated (i.e., the actions of picking up left side panel and picking up right side panel are both labeled as Pick Up Side Panel, thus Pick Up Side Panel is repeated). From the results, we achieve competitive numbers for both progression and Kendall’s Tau on both Pouring and Penn Action. On Pouring, LAV marginally beats TCN on both metrics (e.g., for progression, vs. ), while on Penn Action, LAV significantly outperforms TCC on Kendall’s Tau (i.e., vs. ). Moreover, on Penn Action, the combination of LAV+TCC yields a significant performance gain over TCC on both metrics (e.g., for Kendall’s Tau, vs. ).
4 Fine-Grained Frame Retrieval Results
Here, we utilize our representations for the task of fine-grained frame retrieval. We perform evaluations using the validation set of Pouring and Penn Action. In particular, we alternatively consider each video of the validation set as a query video and all the remaining videos of the validation set as a support set. For each query frame in the query video, we retrieve its most similar frames in the support set by finding its nearest neighbors in the embedding space. We report Average Precision at , which is the average percentage of the retrieved frames with the same action phase labels as the query frame. Tab. 4 presents the quantitative results of various self-supervised methods on Pouring and Penn Action. It is evident from Tab. 4 that LAV consistently achieves the best performance across different values of on both datasets (e.g., on Pouring, for AP@5, for LAV vs. , , and for TCC, TCN, and SAL respectively). This shows that our method is better at learning fine-grained features, which are important to this task. Also, the combined LAV+TCC leads to a significant performance gain over TCC (e.g., vs. ).
Moreover, we present some qualitative results with in Fig. 6, showing that LAV is more capable of capturing fine-grained features than TCC. In Fig. 6(a), the person in the query image has one leg elevated above the ground, which is also seen in 4 out of 5 images retrieved by LAV, whereas TCC fails to capture that in all of its retrieved images (see cyan circles). In Fig. 6(b), the actor in the query image is at the start of Golf Swing with the ball on the ground, which is also seen in all of LAV’s retrieved images, whereas TCC retrieves images with wrong phases (i.e., the person has finished Golf Swing with the ball not visible on the ground, see magenta circles).
5 Joint All-Action Model Results
So far, we have followed to train a separate model for each action of Penn Action and report the average results across all actions. This is not convenient both in terms of training time and memory requirement. In this section, we explore another experimental setup, where we jointly train a single model for all actions of Penn Action. In particular, we train 13 SVM classifiers (1 for each action) but share a single encoder. It is more challenging, since the network needs to jointly learn useful features for all actions. Tab. 5 shows the quantitative results of different self-supervised methods in the above setup. We observe that the performance of all methods is reduced as compared to Tabs. 2 and 3. Moreover, we notice LAV achieves the best performance across all metrics, outperforming TCC, TCN, and SAL in Tab. 5. This can be attributed to the fact that LAV leverages information from across videos in addition to cues from each individual video.
Additional Results. Note that due to space limits, we provide several additional experimental results, including training-from-scratch results and ablation results of hyperparameter settings, in supplementary materials.
Conclusion
In this work, we propose a novel fusion of temporal alignment loss and temporal regularization for learning self-supervised video representations via temporal video alignment, utilizing both frame-level and video-level cues. The two components are complementary to each other, i.e., temporal regularization prevents degenerate solutions while temporal alignment loss leads to higher performance. We show superior performance over prior methods for video-based self-supervised representation learning on various temporal understanding tasks on Pouring, Penn Action, and IKEA ASM datasets. Also, our method offers significant accuracy gain when lacking labeled data. Our future work will explore other temporal alignment losses, e.g., , to allow local temporal permutations and arbitrary video starting/ending points.
Acknowledgements. We would like to thank D. Dwibedi for releasing the code and answering questions about TCC.
Appendix A Supplementary Material
In this supplementary material, we first present results on a subset of 11 actions of Penn Action in Sec. A.1 and fine-grained frame retrieval results on IKEA ASM in Sec. A.2. We then show training-from-scratch results in Sec. A.3 and ablation results of , , and in Sec. A.4. Next, in Sec. A.5 we show results of combining LAV with TCC and TCN while we show results on a recent frame-shuffling method in Sec. A.6. Moreover, we visualize our embeddings and provide our labels for Penn Action in Secs. A.7 and A.8 respectively. Finally, we describe our additional implementation details in Sec. A.9.
Among the 3 datasets that we use in Sec. 5 of the main paper (i.e., Pouring, Penn Action, and IKEA ASM), we notice that for Penn Action while TCC performs well on most actions, it struggles on 2 actions i.e., Baseball Swing and Tennis Forehand. As we can see in Fig. 7, the Kendall’s Tau results by TCC on the above 2 actions (red and orange curves) do not go higher than , which hurts its overall performance on Penn Action in Tabs. 2-4 of the main paper. This might be due to the fact that the beginning and ending frames of the above 2 actions are visually similar, and TCC does not have an explicit mechanism to avoid aligning the beginning frames with the ending ones and vice versa (see Fig. 8 for examples). Other self-supervised methods, i.e., SAL, TCN, and LAV, do not suffer from the above problem as TCC, since they leverage temporal order information, i.e., SAL performs temporal order verification, TCN uses temporal coherence, while LAV exploits both temporal coherence and dynamic time warping prior. Also, the above problem for TCC might be alleviated by tuning the number of context frames and context stride, however, that requires further exploration.
For completeness, we filter out the results of the above 2 actions from Tabs. 2-4 of the main paper, and present the results of the remaining 11 actions of Penn Action in Tabs. 6 and 7. From the results, the performance of TCC is improved significantly when excluding the above 2 actions. In Tab. 6, TCC has competitive numbers with TCN and LAV (e.g., TCC performs the best on progression, while TCN and LAV perform the best on Kendall’s Tau and classification respectively). In Tab. 7, TCC and LAV have very competitive numbers (e.g., TCC slightly outperforms LAV for AP@5, while LAV marginally outperforms TCC for AP@10 and AP@15), outperforming SAL and TCN.
A.2 Fine-Grained Frame Retrieval Results on IKEA ASM
We now conduct fine-grained frame retrieval experiments on IKEA ASM and report the quantitative results of different self-supervised methods in Tab. 8. It is evident from the results that LAV consistently achieves the best performance across different values of , outperforming other methods by significant margins. For example, for AP@5, LAV obtains , while TCC, TCN, and SAL get %, %, and % respectively. Furthermore, the combined LAV+TCC leads to significant performance increase over TCC. For instance, for AP@5, LAV+TCC achieves , while TCC obtains . The above observations on IKEA ASM are similar to those on Penn Action and Pouring reported in Sec. 5.4 of the main paper, confirming the utility of our self-supervised representation for fine-grained frame retrieval.
A.3 Training-from-Scratch Results
All of the experiments in Sec. 5 of the main paper utilize an encoder network initialized with pre-trained weights from ImageNet classification. For completeness, we now experiment with learning from scratch. We use a smaller backbone network, i.e., VGG-M , (instead of ResNet-50) for this experiment. Tab. 9 shows the quantitative results of different self-supervised methods when learning from scratch on Pouring. It can be seen from Tab. 9 that the performance of all methods drops as compared to Tabs. 2 and 3 of the main text. Moreover, SAL and TCN are inferior to TCC and LAV across all metrics. Lastly, although LAV has slightly lower classification accuracy than TCC, LAV outperforms TCC on both progression and Kendall’s Tau.
Next, Tab. 10 shows training-from-scratch results on Penn Action, using a single joint model for all actions (similar as Sec. 5.5 of the main paper). For all methods, the performance in Tab. 10 is lower than Tab. 5 of the main text. Also, SAL and TCN are inferior to TCC and LAV. TCC performs the best on progression, while LAV performs the best on the other two metrics.
Finally, we obtain training-from-scratch results on IKEA ASM, which show LAV achieves the best performance (i.e., for classification, 23.84 for LAV vs. 22.04, 20.45, and 20.42 for TCC, TCN, and SAL respectively).
A.4 Ablation Results of α𝛼\alpha, σ𝜎\sigma, and p𝑝p
We first present ablation results of on Pouring in Fig. 9(a). We observe that the performance is generally stable across values of , and yields the best results. Next, Figs. 9(b) and 9(c) illustrate ablation results of and respectively on Pouring. From the results, the performance is generally stable across values of and . Particularly, performs the best, and large is preferred.
A.5 Performance of LAV+TCC and LAV+TCN
We note that LAV+TCC does not consistently perform better than LAV in Tabs. 2 and 3 of the main paper. This might be attributed to the fact that LAV works on L2-normalized embeddings while TCC does not. Since the two components operate on different embedding spaces, combining the two might not always lead to better results.
In addition, we evaluate LAV+TCN on Pouring. We notice LAV+TCN suffers from the same problem as LAV+TCC (i.e., normalized/unnormalized embeddings). LAV+TCN obtains 91.22, 0.7866, and 0.7925 for classification, progression, and Kendall’s Tau respectively, which are comparable to TCN but lower than LAV.
A.6 Performance of a Recent Frame-Shuffling Method
We evaluate the clip order prediction (COP) method of Xu et. al. on Pouring. As it is a clip-based method, we use sliding windows to generate embeddings for frames at window centers. As mentioned in Sec. 4 of the main paper, the network is first trained for the pretext task and then frozen while we train SVM classifier/linear regressor for the main tasks. It achieves 79.44, 0.5309, and 0.6656 for classification, progression, and Kendall’s Tau respectively, which are lower than SAL in Tabs. 2 and 3 of the main paper. This is likely because the pretext task (i.e., COP) is clip-based, whereas the main tasks are frame-based and require capturing fine-grained frame-based details. Further, since we freeze the network while training SVM classifier/linear regressor, it could not disregard irrelevant clip-based details to focus on the one frame that matters.
A.7 Visualization of Embeddings
We present the t-SNE visualization of the embeddings learned by LAV on 3 example actions of Penn Action in Fig. 10. For each action, we show 4 videos with each plotted using a unique color. In addition, we use different shades of the same color to distinguish different frames of the same video, i.e., beginning frames have light shades, while later frames have progressively darker shades. The visualization in Fig. 10 shows that LAV encodes each video as an overall smooth trajectory in the embedding space, where temporally close frames are mapped to nearby points in the embedding space and vice versa. Moreover, corresponding frames from different videos are generally aligned in the embedding space, e.g., points of different colors but similar shades are nearby in the embedding space and vice versa. We also sample one random time-step (highlighted by a black circle), and plot corresponding frames from different videos (each bordered by a distinct color), which are shown to belong to the same action phase. The above observations show the potential application of our self-supervised representation for temporal video alignment.
A.8 Labels for Penn Action
We have made our dense per-frame labels for 2123 videos of Penn Action publicly available at https://bit.ly/3f73e2W. Please refer to Tab. 2 of TCC for more details on actions, numbers of phases, lists of key events, and numbers of videos for training and validation.
A.9 Implementation Details
For fair evaluations, we use the same data augmentation techniques and encoder networks for all the competing methods. More specifically, we follow the same data augmentation procedures and borrow the encoder networks from TCC . Please refer to the supplementary material of TCC for more details on data augmentation techniques and encoder networks. In addition, we list the hyperparameter settings for our method in Tab. 11. For other methods, we use the same hyperparameter settings suggested by TCC.