Representation Learning via Global Temporal Alignment and Cycle-Consistency
Isma Hadji, Konstantinos G. Derpanis, Allan D. Jepson
Introduction
Temporal sequences (\eg, videos) are an appealing data source as they provide a rich source of information and additional constraints to leverage in learning. By far the main focus on temporal sequence analysis has been on learning representations targeting distinctions at the global signal level, \eg, action classification, where abundant labeled data is available for training. In this paper, we target a weakly-supervised training regime for representation learning, capable of making fine-grained temporal distinctions.
Most previous approaches to temporally fine-grained understanding of sequential signals have considered fully-supervised training methods (\eg, ), where labels are provided at the sub-sequence level, \eg, frames. The major drawback of these methods is the expense in acquiring dense labels, and their subjective nature. In contrast, a key consideration in our work is the selection of training signals capable of scaling up to large amounts of data yet supporting finer-grained video understanding.
As outlined in Fig. 1, given a set of paired sequences capturing the same process (\eg, tennis forehand) but highly varied (\ie, different participants, action executions, scenes, and camera viewpoints), our method trains an embedding network to support the recovery of their latent temporal alignment. We refer to our method as weakly supervised, as only readily available sequence-level labels (\eg, tennis forehand) are required to construct training pairings containing the same process. Given such a pair of sequences, we use their latent temporal alignment as a supervisory signal to learn fine-grained temporal distinctions.
Key to our proposed method is a novel dynamic time warping (DTW) formulation to score global alignments between paired sequences. DTW enforces a stronger constraint over simply considering local (soft) nearest neighbour correspondences , since the temporal ordering of the matches in the sequences are taken into account. We depart from previous differentiable DTW methods by taking a probabilistic path finding view of DTW that encompasses the following three key features. First, we introduce a differentiable smoothMin operator that effectively selects each successive path extension. Moreover, we show that this operator has a contrastive effect across paths which is missing in previous differentiable DTW formulations . Second, the pairwise embedding similarities that form our cost function are defined as probabilities, using the softmax operator. Optimizing our loss is shown to correspond to finding the maximum probability of any feasible alignment between the paired sequences. The softmax operator over element pairs also provides a contrastive component which we show is crucial to prevent the model from learning trivial embeddings. This forgoes the need for a downstream discriminative loss and the corresponding non-trivial task of defining negative alignments, \eg, . Third, as an additional supervisory signal, our probabilistic framework admits a straightforward global cycle-consistency loss that matches the alignments recovered through a cycle of sequence pairings. Collectively, our method takes into account long-term temporal information that allows us to learn embeddings sensitive to fine-grained temporal distinctions (\eg, human pose), while being invariant to nuisance variables, \eg, camera viewpoint, background, and appearance.
Contributions. We make the following key contributions:
A novel weakly supervised method for representation learning tasked with discovering the alignment between sequence pairings for the purpose of fine-grained temporal understanding.
A differentiable DTW formulation with two novel features: (i) a smoothMin operation that admits a probabilistic path interpretation and is contrastive across alternative paths, and (ii) a probabilistic data term that is contrastive across alternative data pairs.
A global cycle consistency loss to further enforce the temporal alignment.
An extensive set of evaluations, ablations, and comparisons with previous methods. We report significant performance increases on several tasks requiring fine-grained temporal distinctions.
Two downstream applications, namely 3D pose reconstruction and audio-visual retrieval.
Our code and trained models will be available at: https://github.com/hadjisma/VideoAlignment.
Related work
Representation learning. Most focus in representation learning with videos has been cast in a fully supervised setting, \eg, . Self-supervised learning with images or videos has emerged as a viable alternative to supervised learning, where the supervisory signal is obtained from the data. For video, a variety of proxy tasks have been defined in lieu of training with annotations, such as classifying whether video frames are in the correct temporal order (\eg, ), predicting whether a video is played at a normal or modified rate , solving a spatiotemporal jigsaw puzzle task , predicting figure-ground segmentation , predicting pixel or region correspondences across neighbouring video frames, and predicting some aspect of future frames conditioned on past frames . Others have considered multimodal settings, such as predicting video-audio misalignment . Similar to , our method is best characterized as weakly supervised, where sequence-level labels are used to determine sequence pairings for training.
Sequence alignment. Several methods assume paired, temporally synchronized videos of the same physical event for the purpose of representation learning. In contrast, and more closely related to our work, are methods that seek the alignment between sequences capturing the same process. One approach is to cast the learning objective as maximizing the number of elements between sequences that can be brought into one-to-one correspondence via (soft) nearest neighbours. This method does not leverage the long-term temporal structure of the sequences as done in dynamic time warping (DTW) . Given a cost function, DTW finds the optimal alignment between two sequences defined between elements comprising the sequences. Recent efforts have explored differentiable approximations of the discrete operations underlying DTW to allow gradient-based training. Similar to recent work , we also incorporate a relaxed DTW as our loss for sequence alignment. Our formulation is probabilistic and includes a contrastive definition of the element-wise similarities. A key distinction with these prior works and our own, beyond differences in the target application domain (\eg, video-transcript alignment ), is that rather than incorporate contrastive modelling after the DTW step (\eg, through the use of margin-based loss), our method includes contrastive signals in both the differentiable min approximation and the pairwise matching cost function used in our DTW framework.
Contrastive learning. One can also draw parallels with contrastive learning using the cross-entropy loss (\ie, negative log softmax) , where the goal is to learn a representation that brings different views of the same data together (\ie, positives) in the embedding space, while pushing views of different data (\ie, negatives) apart. This amounts to encoding information shared across the views, while eschewing unique factors to each view. To construct different views, previous work has explored a variety of augmentation and sampling schemes and correspondences across different modalities (\eg, video-audio, video-text, and luminance-depth ). In these works, the positive and negative pairings are known by construction, \eg, via image augmentation. We also make use of contrastive losses, but note that the correspondences (\ie, positives) between the sequences are latent rather than known.
Cycle-consistency. In addition to alignment, our method also incorporates cycle-consistency as a supervisory signal, where the objective is to verify matches across sets. Similar to recent work , we apply cycle-consistency across two temporal sequences. A key difference is that previous work applies cycle-consistency independently to local matches across sequences; whereas, we consider both local matches and their global temporal ordering, which we demonstrate empirically leads to improved alignments.
Technical approach
In this section, we describe our weakly-supervised approach to representation learning based on the alignment of sets of sequence pairs that capture the same process. Our learning objective is the training of a shared embedding function applied to each sequence element. In the case of multimodal sequences (\eg, audio-video), we have separate encoders for each modality that map their inputs to a common embedding space. Figure 2 provides an illustrative overview of our alignment approach to representation learning, which we fully unpack in the following subsections.
The number of feasible alignments in DTW grows exponentially with the sequence lengths. Fortunately, the structure of DTW with an appropriate cost function, , is amenable to dynamic programming which shares both quadratic time and space complexity. The optimal alignment (\ie, path through the cost matrix) is found by evaluating the following recurrence :
where , , and stores the partial accumulated cost along the optimal and feasible path ending with the alignment between and . The minimum operation amounts to a first-order Markov assumption, where the local path routing is deterministic.
Due to the discrete nature of the min operator in (1), responsible for local correspondence decisions, several works have considered smooth variants suitable for gradient-based training. In Sec. 3.2, we introduce a smooth relaxation of the min operator with favourable properties for our representation learning setting. Then in Sec. 3.3, we define our cost function which introduces a contrastive learning signal throughout the alignment process.
2 Local differentiable decisions
to denote the incoming optimal accumulated costs from the feasible paths leading into . We first modify (1) as
where is a smooth approximation of the minimum operator. The term can be seen as a (non-differentiable) additional penalty on any path that reaches . With this added penalty term, (3) reduces to simply
For an appropriate choice of the right hand side is now differentiable. Note that any path that is optimal according to (4) will correspond to a feasible path for the original DTW problem (although perhaps not optimal for that problem). Moreover, the cost of such a path according to (4) will be the original cost plus the sum of the penalties over all points on that path.
We are left with choosing a smooth approximation for the minimum operator. Here, we use a standard relaxation of the min operator (specifically, the expected value for ):
where denotes a temperature hyper-parameter. We refer to solving the recurrence relation (4), with the function taken to be smoothMin, as the smoothDTW problem.
Previous alignment methods have instead used the following formulation as a continuous approximation of the min operator:
where again denotes a temperature hyper-parameter.
While both the and smoothMin operators are differentiable approximations of the min operator (with the min operator subsumed as a special case), their different behaviours have profound effects on learning in our setting. These differences are illustrated in the plot in Fig. 3, where without loss of generality we assume the costs are sorted in increasing order, with for some and . As can be seen, the function is strictly monotonically increasing. As a result, with all other things being equal, minimizing this function encourages ties, \ie, the penalty is minimized only when . This is an undesirable behaviour as we seek the resulting embeddings to yield a well-defined path (\ie, optimal alignment) in our cumulative cost matrix, . In contrast, our smoothMin operator defines a contrastive watershed (at approximately ), where values to the left of the watershed encourage ties, while to the right the values are encouraged to be well separated. Moreover, since for smoothMin, our smoothDTW approach always provides an upper bound on the cost of the optimal path. The supplemental presents an expanded discussion and comparison.
3 Contrastive cost function
To complete the definition of our smoothDTW recurrence in (4), we now specify the cost function . Specifically, for each element in , we wish to express the cost of matching to any single item in , given that at least one of the elements in must match. Moreover, in keeping with our probabilistic path finding formulation, should be the negative log probability of matching the given to a selected in . This leads to the (non-symmetric) contrastive formulation
With this cost function the resulting values are the optimum value of the negative log probability for any feasible path starting at and ending at . Here, this negative log probability is the sum of the matching cost, (cf. (8)), and the smoothness penalty (cf. (14)), at each vertex, , along the optimal path ending at . Correspondingly, we define the alignment loss for matching to as:
We use the sum of the alignment losses for matching to and vice versa as the overall alignment loss, as shown in the left panel in Fig. 2.
The combination of the softmax over elements in , in (8), and the use of smooth DTW to formulate the alignment cost (9), rewards embeddings that are both: a) contrastive across the elements in ; and, moreover, b) have their best matching pairs arranged along a feasible path from to . We show in our ablation study that the ability to leverage these two properties during training are key to our utilization of the temporal alignment proxy. In contrast, dropping the softmax and simply computing the inner-product between elements could lead to the collapse of the embeddings around a single point during training, thus allowing the network to trivially minimize the alignment cost. To avoid such collapse, previous DTW-based methods have resorted to adding a discriminative loss downstream .
4 Global cycle-consistency
An additional loss is based on the notion that the match from sequence to , composed with the match from to , should ideally be the identity. We formulate this directly in terms of the cumulative cost matrix for matching sequence to (as defined by (4), (5), and (8)), along with the cost matrix for matching to , namely . Given the interpretation that is the optimal negative log probability of a path from to for matching to , consider the implied conditional distribution for matching the prefix sequences to for different ’s, namely
Note that this distribution does not use any information for elements from sequence , and is only obtained from the forward pass of matching to . We use the notation
to denote the matrix with elements .
Ideally, the contrastive matching in (8) is sharp and forms a feasible path from to , thereby providing a strongly peaked conditional distribution for each . However, without knowing the ground truth matching , we cannot use an explicit log-likelihood loss. This issue can be avoided by considering the composed conditional distribution
which is formed by treating and as conditionally independent distributions. It is easy to verify that this is indeed a distribution over elements of . Moreover, following the above matrix notation, it is represented by the matrix .
In the ideal case, this transport from elements in one sequence to another and back again should return to the same starting element. From (12) this corresponds to equaling the identity matrix, . Thus, our global cycle-consistency loss is the sum of cross-entropy losses:
The right panel in Fig. 2 provides a summary of our global cycle-consistency loss.
5 Training and implementation details
Our final loss function is obtained by combining the contrastive alignment loss, (9), and the global cycle consistency loss, (13), according to
where and are weights used to balance the two losses and are empirically set to and , respectively. The temperature hyper-parameters, and used in (5) and (8) are both set to , while in (10) is .
This overall loss is used to train a convolutional architecture composed of a backbone encoder applied framewise followed by an embedding network. Specifically, ResNet50-v2 is used as our backbone encoder where we extract features from Conv4c layer. We adopt the same embedding network used in previous related work comprised of two 3D convolutional layers, a global 3D max pooling layer, two fully connected layers, and a linear projection layer. The final embedding is L2 normalized. To learn over sequence pairs , we randomly extract frames from each sequence. Sampling of video frames is random to avoid learning potential trivial solutions that may arise from strided sampling. For a fair comparison to previous approaches using the same architecture (\ie, ), we use the same batch size of four sequences. Finally, our learning rate is fixed to for all our experiments.
Empirical evaluation
We evaluate the efficacy of our learned embeddings on challenging temporal fine-grained tasks, thereby going beyond traditional clip-level recognition tasks. In particular, our proposed loss is evaluated on fine-grained action recognition (\ie, action phase classification), few-shot fine-grained classification, and video synchronization. In addition, we also show that learning to align temporal sequences supports different downstream applications such as synchronous playback, 3D pose reconstruction and fine-grained audio/visual retrieval.
To evaluate our method, we use the PennAction and FineGym video datasets. Both datasets contain a diverse set of videos of human-related sports or fitness activities. These datasets are selected as they allow for learning framewise alignments and evaluating on tasks where fine-grained temporal distinctions are critical.
PennAction contains 2326 videos of humans performing 15 different sports or fitness actions. The videos are tightly cropped temporally around the start and end of the action. Notably, 13 categories contain non-repetitive actions. The dataset also includes ground truth 2D keypoint labels which we later use to demonstrate a 3D pose reconstruction application grounded on our alignment method.
FineGym is a recent large-scale fine-grained action recognition dataset that was specifically designed to evaluate the ability of an algorithm to parse and recognize the different phases of an action. Each video in FineGym is annotated according to a three-level hierarchy denoting the event being performed in the video, the different sets involved in performing the event, and the framewise elements (\ie, action phases) involved in each set. To perform any event-level action, a gymnast may perform the different sets in any order. To train embeddings using our alignment-based method, we re-organize the FineGym dataset such that all sets belonging to the same event appear in the same order in any given video. An example of this re-organization is provided in the supplemental. Also, it should be noted that while the videos of the vault event (VT) in FineGym depict the gymnast performing the action in three phases, the first two phases of the action are not explicitly labeled in the original dataset. For the sake of completeness we use the provided start and end times for these phases, thereby adding two new fine-grained action labels in FineGym. The remaining events in FineGym are otherwise unchanged. To account for these minor additions to the annotations, we refer to the extensions of FineGym99 and 288 as FineGym101 and FineGym290, respectively. Notably, we also report results on the original FineGym99 and 288 in the supplemental. Importantly, we do not use the element-level (\ie, action phase) labels during training.
2 Baselines
We compare our approach to other weakly supervised and self-supervised methods that entail temporal understanding in their definition. A detailed description of the baselines is provided in the supplemental.
3 Ablation study
We first present an ablation study that validates the contribution of each component of our loss. For this purpose, we evaluate fine-grained action recognition performance on FineGym101. Following previous work , we use a Support Vector Machine (SVM) classifier on top of the learned embeddings to report framewise fine-grained classification accuracy. Notably, the classifier is trained on the extracted embeddings with no additional fine-tuning of the network. The results in Table 1 show the pivotal role of our contrastive cost. In fact, turning off the contrastive component of our cost, (8), and simply relying on the cosine distance always leads to no improvement in learning from the onset of training. Also, these results show the advantage of adding our global cycle consistency, which further validates the correspondences. Finally, we also compare the performance of our smoothMin definition vs. the more widely used . The superiority of the adopted smooth definition supports the laid out arguments in Section 3.2.
4 Fine-grained action recognition
We now compare our fine-grained action recognition performance to our baselines using the FineGym dataset and consider two training settings for the backbone framewise encoder. (i) scratch: the backbone ResNet50 is trained from scratch with our proposed loss, (ii) only-bn: we fine-tune batch norm layers of ResNet50 from a model pre-trained on ImageNet . The embedder is otherwise trained from scratch in both cases. A third setting where all layers are fine-tuned is presented in the supplemental.
The results summarized in Table 2 (and the supplemental) speak decisively in favour of our method, where we outperform all other weakly and self-supervised methods with sizable margins. The gap is especially striking in the case of SpeedNet. This poor performance can largely be attributed to the fact that the task optimized in SpeedNet does not require detailed framewise understanding. On the other hand, the closest approach to ours is TCC as it also relies on pairwise local matchings between videos of the same class; however, the matchings are realized independently and thus ignores informative long-term sequence structure. This is in contrast to SaL which uses frames from the same videos to solve tasks requiring temporal understanding. Importantly, the superiority of our results compared to TCC demonstrates that the global nature of our loss makes the learned embeddings more robust to the presence of repeated sub-actions as is the case for most of the FineGym videos. Interestingly, the results obtained with D3TW* highlight the limits of the downstream discriminative loss, which requires an explicit construction of positive and negative examples for training. Notably, while our method outperforms all alternatives under both training settings, the best overall results are obtained under the only-bn setting and it is therefore used for all other experiments reported in this paper.
We also considered classification results of each event separately where we also outperformed all alternatives. Importantly, visualizations of the learned features suggests that the proposed loss learns to adapt and identify the most reliable cues to learn the alignments. Please see supplemental for detailed results, discussions, and visualizations.
5 Few-shot fine-grained action recognition
An advantage of our proposed weakly supervised method is that it does not rely on framewise labels for training. To evaluate this advantage, we also report few-shot classification results. In this case, the entire training set is used to learn the embeddings, but only a few videos per class are used to train the classifier. In particular, we use FineGym101 for this experiment with an increasing number of videos per class to train the classifier, starting from the 1-shot setting all the way to using the entire dataset.
The plot in Fig. 4 further confirms the superiority of our proposed method. As can be seen, our method outperforms all others across the range of number of labeled training videos used, with an especially strong performance even under the challenging 1-shot setting.
6 Video synchronization
To evaluate the quality of the synchronization (\ie, alignment) between two videos we use Kendall’s Tau metric, which does not require framewise labels. However, this metric assumes little to no repetitions in the aligned videos. We therefore follow previous work and report results only on the 13 classes without repetition in PennAction (\ie, strumming guitar and jumping rope are not included). Importantly, while previously reported results were obtained by training a different network for each class of PennAction, we consider the more challenging setting of learning a joint representation over all classes (\ie, as done in all previous experiments with FineGym). For a fair comparison, all baselines are also trained over all classes of PennAction rather than one network per class.
Table 3 summarizes our video alignment results, where we once again achieve state-of-the-art performance compared to the baselines. Importantly, our performance is also superior to previously reported results under the multi-network training setting . We report results under this setting in the supplemental.
In addition to the quality of synchronization, we also report in Table 3 state-of-the-art results on action phase classification. Notably, the action phase labels used here are provided by the original authors of the PennAction dataset as the labels used previously were not made publically available.
Downstream applications
3D pose reconstruction. As demonstrated by our evaluations, our proposed weakly-supervised method is capable of temporally aligning videos of similar actions even while they are captured in different environments, at different execution rates, and from different viewpoints. As a result, we can align poses of the same action from different viewpoints by forcing different videos to play synchronously. Examples of this video synchronization are provided in the supplemental. Importantly, the synchronization of videos taken at various viewpoints can serve as the basis for 3D pose reconstruction. To demonstrate this ability, we use videos from PennAction and their corresponding ground truth 2D keypoint labels. In particular, given a random query video from a given class in PennAction, we first start by aligning the remaining videos (from the same class) to it. Given these aligned frames and their corresponding 2D keypoints we use the Tomasi-Kanade factorization algorithm , followed by bundle adjustment , to compute a temporally aligned 3D model of the action performed. Fig. 5 presents an example of our 3D pose reconstructions; additional results are available in the supplemental material.
Audio-visual alignment. Finally, we demonstrate that our alignment based representation learning method can be applied to other types of sequences by applying it on separate audio and visual inputs. Given that any audio-visual pair is by default aligned and to avoid learning trivial solutions, we sample audio segments and video frames differently to make the task of learning the alignment harder on the networks and consequently learn strong audio and visual embeddings. In particular, while the audio signal is uniformly sampled into consecutive one second long segments, video frames are on the other hand randomly sampled along the temporal dimension. This sampling strategy encourages our model to learn different alignment paths due to the randomness in the video frame selection. For visual features, the same backbone and embedding network described in Sec. 3.5 is used to encode video frames, while we use VGGish to encode audio signals. In particular, each one second long audio segment is first converted into a log mel spectogram and used as an input for the VGGish network. Training is otherwise performed as described in Sec. 3.5.
For this sample application, we use the firing cannon class from the VGGSound dataset for training and testing. This class is selected as it is strongly visually indicated with an easily identifiable salient auditory signal (\ie, the explosion sound emitted upon firing of a cannon).
To demonstrate the quality of the learned audio and visual embeddings, we evaluate them on the task of fine-grained audio/visual retrieval. In particular, given a one second long audio signal corresponding to the moment of firing a cannon, we extract its corresponding audio embedding from the VGGish network trained using our approach. This embedding is then used to query corresponding visual embeddings from the vision network. The top-5 nearest frame embeddings from all videos in the test set are extracted. Sample correspondences, shown in Fig. 6, clearly depict that the corresponding visual embeddings also capture the moment of the firing visually. More details and sample results are presented in the supplemental.
Conclusion
In summary, this work introduced a novel weakly supervised method for representation learning relying on sequence alignment as a supervisory signal and taking a probabilistic view in tackling this problem. Because the latent supervisory signal entails detailed temporal understanding, we judge the effectiveness of our learned representation on tasks requiring fine-grained temporal distinctions and show that we establish a new state of the art. In addition, we present two applications of our temporal alignment framework, thereby opening up new avenues for future investigations grounded on the proposed approach.
References
Appendix A Supplementary material
Our supplementary material is organized as follows: Sec. B derives several properties of the two smooth minimum approximations introduced in the main manuscript. Sec. C provides details of our evaluation baselines. Sec. D further explains our re-organization of FineGym to satisfy our training requirements. For completeness, Secs. E and F provide additional fine-grained action recognition results. Similarly, Sec. G provides additional alignment results on PennAction. Finally, Sec. H probes our learned representations via visualizations to ascertain what information is captured. A supplemental video highlighting the different applications of the proposed loss is also provided.
Appendix B Smooth Minimum Properties
smoothMin properties. The smoothMin function is defined by
where denotes a temperature hyper-parameter. Define to be the minimum coefficient in , that is, . Then for it follows that
Therefore, the smoothness penalty satisfies
Note that from (15) it follows that , so we have , \ie, it provides an upper bound on the true minimum.
The maximum possible value of the penalty is also of interest. Given (17) it is sufficient to maximize subject to and for . By setting and simplifying, we find
for . This implies all the ’s for are equal at the maximum. Using (15) in (18), setting and for , and simplifying, we find must be the solution of
For , (19) has a unique solution . This implies that the penalty function (17) only has maxima, , for which there is exactly one distinct minimum value , and all the other values are in an -way tie for second largest at for , and where is as in (19).
To compute the value of this maximum, we substitute the resulting vector into and, by simplifying, we find
where we used (19) in the second line above.
For concrete examples, we find , for and , respectively. Also, for it follows from (19) that (since the left hand side less minus the right is negative at , positive for , and the derivative of this difference with respect to is positive). Therefore, (20) implies that the maximum smoothness penalty when is roughly , and for , which are in agreement with the plot in Fig. 3 of the main manuscript. Moreover, for large the maximum penalty is .
properties. Similarly we examine , which is defined by
For and we have
Recall that so for at least one . Given that all the terms in the sum are non-negative, and one of them is 1, we conclude that . Therefore, we have , \ie, it provides an lower bound on the true minimum.
Moreover, the maximum sum (22) occurs when for all (\ie, there is an -way tie for the minimum). In this case . Therefore, we have
Finally, we see that the smoothness penalty when is used is given by
Therefore, we have shown that is always negative and achieves its most negative value of if and only if the input vector represents an -way tie, \ie, .
Appendix C Baselines
We compare our approach to other weakly supervised and self-supervised methods that entail temporal understanding in their definition.
SpeedNet learns a video representation by learning to distinguish between videos played at different speeds.
Time Contrastive Networks (TCN) makes fine-grained temporal distinctions by optimizing a contrastive loss that encourages embeddings of an anchor image and an image taken simultaneously from a different camera viewpoint to be similar, while the embedding of an image taken from the same sequence of the anchor but at a different time instant to be distant from that of the anchor.
Shuffle and Learn (SaL) learns to predict whether a triplet of frames are in the correct order or shuffled.
Discriminative Differentiable Time Warping (D3TW) was originally introduced to learn alignments between video frames and action labels. Its definition includes a discriminative component requiring explicit definitions of positive and negative pairs. A pair is considered positive if all action labels are used in the alignment table, whereas negative pairs are constructed by dropping some of the action labels. We adapt this loss to our video pair alignment setting and define D3TW* as follows: a positive pair is constructed by sampling frames from the entire duration of both video-1 and video-2 in any pair, whereas we randomly drop portions of video-2 to constitute a negative pair, thereby mimicking the dropping of action labels as proposed in the original paper.
Temporal Cycle Consistency (TCC) learns fine-grained temporal correspondences between individual video frames by imposing a soft version of cycle consistency on the individual matches.
Appendix D FineGym re-organization
As mentioned in the main mansucript, each video in FineGym is annotated according to a three-level hierarchy denoting the event being performed in the video, the different sets involved in performing the event, and the framewise elements (i.e., action phases) involved in each set. To perform any event-level action, a gymnast may perform the different sets in any order. To train our embedding network using our alignment-based method, we re-organize the FineGym dataset such that all sets belonging to the same event appear in the same order in any given video. For example, given floor exercise events, gymnasts can perform four different sets of exercises in any order, we re-organize the clips in each video according to a selected prototype order, as shown Figure 7. These organized event-level videos are used during training and testing.
Appendix E Fine-grained action recognition
In Sec. 4.4 of the main manuscript, we compared our fine-grained action recognition performance to our baselines using FineGym with two training settings for the backbone framewise encoder. Here, we provide our complete comparison. In addition, to training from scratch (train-all) and fine-tuning the batch norm layers (only-bn), we include a third experiment consisting of fine-tuning all layers of a ResNet50 model pre-trained on ImageNet (train-all).
The full results of all experiments are summarized in Table 4. Consistent with our conclusions in Sec. 4.4, we outperform all the weakly and self-supervised baseline methods by significant margins under all training settings with the best results obtained under the only-bn setting.
For completeness, we also provide results of training the SVM classifier for framewise fine-grained action recognition on the original FineGym99 and 288 short clips. The results of this experiment, summarized in Table 5, once again demonstrate the superiority of the proposed approach even under this more challenging setting. Notably, comparison between our method and those reported in is not direct for two main reasons. First, we are targeting framewise accuracy, whereas focuses on clip-level accuracy. Second, we are the first to report weakly supervised results on FineGym.
Appendix F Details of fine-grained action recognition
To further investigate the utility of the learned embeddings, we also consider classification results of each event separately. In particular, using the embeddings learned on the entire FineGym101 dataset, we train a separate SVM classifier for each event. The results summarized in Table 6, further confirm the superiority of our approach. These results also show that classifying the sub-actions in the floor exercise (FX) event is the most challenging for all methods. Careful examination of videos in this class revealed wide variations in the way gymnasts perform each sub-action in the floor exercise event, which makes learning a proper alignment especially challenging.
Appendix G Video synchronization
In the main manuscript, we provide results of video synchronization under the challenging setting of training a single network for all classes in PennAction, whereas trained a different network for each class. For completeness, we provide additional results here where we also trained a network per class using our loss on PennAction to directly compare with . The results summarized in Table 7 speak decisively in favor of our approach where we outperform all approaches with a sizeable margin.
Appendix H Visualizing learned features
To investigate what our learned representation captures, we adapt the Class Activation Map (CAM) method to visualize the learned features. In particular, we extract feature maps from the last convolutional layer of our embedding network and simply average them along the channel dimension. The resulting activation maps are then normalized between 0 and 1 framewise, upsampled to match the input dimensions, and superimposed on the input video frames. For PennAction, the heatmaps in Figure 8 (a) shows that our embeddings are selective to body parts most involved in performing an action. This can explained by the fact that videos in PennAction are carefully curated with relatively clean, similar actions with no repetitions. On the other hand, we can see from Figure 8 (b), that our embeddings are tuned to human contact with surfaces to learn alignments in FineGym. This is an especially desired behaviour as while there can be significant variations in the way gymnasts perform different phases of an action, they generally share some commonality in the manner that they make contact with surfaces. More generally, these visualizations suggest that the proposed loss learns to adapt and identify the most reliable cues to learn the alignments. Additional activation visualizations are provided in the supplemental video.
Appendix I Downstream applications
Please see supplemental video see video at: https://github.com/hadjisma/VideoAlignment for various downstream application results.