Memory-augmented Dense Predictive Coding for Video Representation Learning

Tengda Han, Weidi Xie, Andrew Zisserman

Introduction

Recent advances in self-supervised representation learning for images have yielded impressive results, e.g\xperiod , with performance matching or exceeding that of supervised representation learning on downstream tasks. However, in the case of videos, although there have been similar gains for multi-modal self-supervised representation learning, e.g\xperiod , progress on learning only from the video stream (without additional audio or text streams) is lagging behind. The objective of this paper is to improve the performance of video only self-supervised learning.

Compared to still images, videos should be a more suitable source for self-supervised representation learning as they naturally provide various augmentation, such as object out of plane rotations and deformations. In addition, videos contain additional temporal information that can be used to disambiguate actions e.g\xperiod open vs. close. The temporal information can also act as a free supervisory signal to train a model to predict the future states from the past either passively by watching videos or actively in an interactive environment , and thereby learn a video representation.

Unfortunately, the exact future is indeterministic (a problem long discussed in the history of science, and known as “Laplace’s Demon”). As shown in Figure 1, this problem is directly apparent in the stochastic variability of scenes, e.g\xperiod trying to predict the exact motion of each leaf on a tree when the wind blows, or the changing reflections on the water. More concretely, consider the action of ‘playing golf’ – once the action starts, a future frame could have the hands and golf club in many possible positions, depending on the person who is playing. Learning visual representation by predicting the future therefore requires designing specific training schemes that simultaneously circumvents the unpredictable details in exact frames, and also handles multiple hypotheses and incomplete data – in particular only one possible future is exposed by the frames of one video.

Various approaches have been developed to deal with the multiple possible futures for an action. Vondrick et al\xperiod explicitly generates multiple hypotheses, and only the hypothesis that is closest to the true observation is chosen during optimization, however, this approach limits the number of possible future states. Another line of work circumvents this difficulty by using contrastive learning – the model is only asked to predict one future state that assigns higher similarity to the true observation than to any distractor observation (from different videos or from elsewhere in the same video). Recalling the ‘playing golf’ example, the embedding must capture the hand movement for this action, but not necessarily the precise position and velocity, only sufficiently to disambiguate future frames.

In this paper, we continue the idea of contrastive learning, but improve it by the addition of a Compressive Memory, which maps “lifelong” experience to a set of compressed memories and helps to better anticipate the future. We make the following four contributions: First, we propose a novel Memory-augmented Dense Predictive Coding (MemDPC) architecture. It is trained with a predictive attention mechanism over the set of compressed memories, such that any future states can always be constructed by a convex combination of the condensed representations, allowing it to make multiple hypotheses efficiently. Second, we investigate visual only self-supervised video representation learning from RGB frames, or from unsupervised optical flow, or both. Third, we argue that, in addition to the standard linear probes and fine-tuning , that have been used for evaluating representation from self-supervised learning, a non-linear probe should also be used, and demonstrate the difference that this probe makes. Finally, we evaluate the quality of learnt feature representation on four different downstream tasks: action recognition, learning under low-data regime (scarce annotations), video retrieval, and unintentional action classification; and demonstrate state-of-the-art performance over other approaches with similar settings on all tasks.

Related Work

Self-supervised learning for images has undergone rapid progress in visual representation learning recently . Generally speaking, the success can be attributed to one specific training paradigm, namely contrastive learning , i.e\xperiod contrast the positive and negative sample pairs.

Self-supervised learning for videos has explored various ideas to learn representations by exploiting spatio-temporal information . Of more relevance here is the line of research using contrastive learning, e.g\xperiod learn from visual-audio correspondence, learns from video and narrations, and our previous work learns video representations by predicting future states.

Memory models have been considered as one of the fundamental building blocks towards intelligence. In the deep learning literature, two different lines of research have received extensive attention, one is to build networks that involve an internal memory which can be implicitly updated in a recurrent manner, e.g\xperiod LSTM and GRU . The other line of research focuses on augmenting feed-forward models with an explicit memory that can be read or written to with an attention-based procedure . In this work, our compressive memory falls in the latter line, i.e\xperiod an external memory module.

Methodology

The proposed Memory-augmented Dense Predictive Coding (MemDPC), is a conceptually simple model for learning a video representation with contrastive predictive coding. The key novelty is to augment the previous DPC model with a Compressive Memory. This provides a mechanism for handling the multiple future hypotheses required in learning due to the problem that only one possible future is exposed by a particular video.

The architecture is shown in Figure 2. As in the case of DPC, the video is partitioned into 8 blocks with 5 frames each, and an encoder network ff generates an embedding zz for each block. For inference, these embeddings are aggregated over time by a function gg into a video level embedding cc. During training, the future block embeddings z^\hat{z} are predicted and used to select the true embedding in the dense predictive coding manner. In MemDPC, the prediction of z^\hat{z} is from a convex combination of memory elements (rather from cc directly as in DPC), and it is this restriction that also enables the network to handle multiple hypotheses, as will be explained below.

Video Block Encoder. As shown in Figure 2, we partition the input video sequence into multiple blocks x1,...,xt,xt+1,...x_{1},...,x_{t},x_{t+1},..., where each block is composed of multiple video frames. Then a shared feature extractor f(.)f(.) (architecture details are given in the appendix) extracts the video features ziz_{i} from each video block xix_{i}:

Temporal Aggregation. After acquiring block representations, multiple block embeddings are aggregated to obtain a context feature ctc_{t}, summarizing the information over a longer temporal window. Specifically,

in our case, we simply adopt Recurrent Neural Networks (RNNs) for g(.)g(.), but other auto-regressive model should also be feasible for temporal aggregation.

Compressive Memory. In order to enable efficient multi-hypotheses estimation, we augment the predictive models with an external common compressive memory. This external memory bank is shared for the entire dataset during training, and is accessed by a predictive addressing mechanism that infers a probability distribution over the memory entries, where each memory entry acts as a potential hypothesis.

Multiple Hypotheses. The dot product of the predicted and desired future pairs can be rewritten as:

where mi⊤zm_{i}^{\top}z refers to the dot product (i.e\xperiod similarity) between a single memory slot and the feature states from the observation. The objective of ϕ(.)\phi(.) is to predict a probability distribution over kk hypotheses in the memory bank, such that the expectation of mi⊤zm_{i}^{\top}z for a positive pair is larger than that of negative pairs. Since the future is uncertain, the desired future feature zz might be similar to one of the multiple hypotheses in the memory bank, say either mpm_{p} or mqm_{q}, for instance. To handle this uncertainty, the future prediction function ϕ(.)\phi(.) just needs to put a higher probability on both the pp and qq slots, such that Equation 5 is always large no matter which state the future is. In this way, the burden of modelling the future uncertainly is allocated to the memory bank M\mathbf{M} and future prediction function ϕ(.)\phi(.), thus the backbone encoder f(.)f(.) and g(.)g(.) can save capacity and capture the high-level action trajectory.

Memory Mechanism Discussion. Note, in contrast to the memory mechanism in Wu et al\xperiod and MoCo , which has the goal of storing more data samples to increase the number of negative samples during contrastive learning, our Compressive Memory has the goal of aiding learning by compressing all the potential hypotheses within the entire dataset, and allowing access through the predictive addressing mechanism. The memory mechanism shares similarity with NetVLAD , which represents a feature distribution with “trainable centroids”. However, in NetVLAD the goal is for compact and discriminative feature aggregation, and it encodes a weighted sum of residuals between feature vectors and the centroids. In contrast, our goal with ϕ(.)\phi(.) is to predict attention weights for the entries in the memory bank M\mathbf{M}, in order to construct the future state as a convex combination these entries. The model can also sequentially predict further into the future with the same memory bank.

2 Contrastive Learning

Contrastive Learning generally refers to the training paradigm that forces the similarity scores of positive pairs to be higher than those of negatives.

Specifically, in MemDPC, we predict the future states recursively, ending up with a sequence of predicted features z^t+1,z^t+2,…,z^end\hat{z}_{t+1},\hat{z}_{t+2},\dots,\hat{z}_{end} and the video feature from the true observations zt+1,zt+2,…,zendz_{t+1},z_{t+2},\dots,z_{end}. As shown in Figure 3, each predicted z^\hat{z} in practise is a dense feature map. To simplify the notation, we denote temporal index with ii and denote other indexes including spatial index and batch-wise index as kk, where batch-wise index means the index in the current mini-batch, k∈{(1,1,1),(1,1,2),...,(B,H,W)}k\in\{(1,1,1),(1,1,2),...,(B,H,W)\}. The objective function to minimize becomes:

where ψ(⋅)\psi(\cdot) is acting as a critic function, in our case, we simply use dot product between the two vectors (we also experiment with L2-normalization, and find it gives similar performance on downstream tasks). Essentially, the objective function acts as a multi-way classifier, and the goal of optimization is to learn the video block encoder that assigns the highest values for (z^i,k,zi,k)(\hat{z}_{i,k},z_{i,k}) i.e\xperiod higher similarity between the predicted future states and that from true observations originating from the same video and spatial-temporal aligned position.

3 Improving Performance by Extensions

As MemDPC is a general self-supervised learning framework, it can be combined with other ‘modules’ like two-stream networks and bi-directional RNN to improve the quality of the visual representations.

Two-stream Architecture. We represent dense optical flow as images by stacking the xx and yy displacement fields and another zero-valued array to make them 3-channel images. There is no need to modify the MemDPC framework, and it can be directly applied to optical flow inputs by simply replacing the input xtx_{t} from RGB frames to optical flow frames. We use late fusion like to combine both streams.

Bi-directional MemDPC. From the perspective of human perception, where only the future is actively predicted, MemDPC is initially designed to be single-directional. However, when passively taking the videos as input, predicting backwards becomes feasible. Bi-directional MemDPC has a shared feature extractor f(.)f(.) to extract the features z1,z2,...,ztz_{1},z_{2},...,z_{t}, but has two identical aggregators gf(.)g^{f}(.) and gb(.)g^{b}(.) denoting forward and backward aggregation. They aggregate the bi-directional context features ctfc^{f}_{t} and ctbc^{b}_{t}. Then MemDPC predicts the past and the future features with the shared ϕ(.)\phi(.) and shared memory bank M\mathbf{M}, and constructs contrastive losses for both directions, namely Lf\mathcal{L}^{f} and Lb\mathcal{L}^{b}. The final loss is the average of the losses from both directions.

How to Evaluate Self-Supervised Learning?

The standard way to evaluate the quality of the learned representation is to assess the performance on downstream tasks using two protocols: (i) a linear probe – freezing the network and only train a linear head for the downstream task; or (ii) fine-tuning the entire network for the downstream task. For example, in (i) if the downstream task is classification, e.g\xperiod of UCF101, then a linear classifier is trained on top of the frozen base network. In (ii) the self-supervised training of the base network only provides the initialization. However, there is no particular reason why self-supervision should lead to features that are linearly separable, even if the representation has encoded semantic information. Consequently, in addition to the two protocols mentioned above, we also evaluate the frozen features with non-linear probing, e.g\xperiod in the case of a classification downstream task, a non-linear MLP head is trained as the final classifier. In the experiments we evaluate the representation on four different downstream tasks.

Action Classification is a common evaluation task for self-supervised learning on videos and it allows us to compare against other methods. After self-supervised training, our MemDPC can be evaluated on this task under two settings: (i) linear and non-linear probing with a fixed network (here the entire backbone network, namely f(.)f(.), g(.)g(.)); and (ii) fine-tuning the entire network end-to-end with supervised learning. For the embedding, as shown in Figure 4, we take the input video blocks x1,x2,...,xtx_{1},x_{2},...,x_{t} in the same way as MemDPC and extract the context feature ctc_{t} using the feature extractor f(.)f(.) for each block and temporal aggregator g(.)g(.); then we spatially pool the context feature ctc_{t} to obtain the embedding. We describe the training details in Section 5.3. The detailed experiment can be found in Section 5.4.

Data Efficiency and Generalizability are reflected by the effectiveness of the representation under a scarce-annotation regime. For this task, we take the MemDPC representation and finetune it for action classification task, but limit the model to only use 10%, 20% and 50% of the labelled training samples, then we report the accuracy on the same testing set. The classifier has the identical training pipeline as shown in Figure 4, and training details are explained in Section 5.3. The detailed experiment can be found in Section 5.5.

Video Action Retrieval directly evaluates the quality of the representation without any further training, aiming to provide a straightforward understanding on the quality of the learnt representation. Here, we use the simplest non-parametric classifier, i.e\xperiod k-nearest neighbours, to determine whether semantically similar actions are close in the high-dimensional space. Referring to Figure 4, for each video, we truncate it into blocks x1,x2,...,xtx_{1},x_{2},...,x_{t} and extract the context feature ctc_{t} with the f(.)f(.) and g(.)g(.) trained with MemDPC. We spatially pool ctc_{t} to get a context feature vector, which is directly used as a query vector for measuring the similarity with other videos in the dataset. The detailed experiment can be found in Section 5.6.

Unintentional Actions is a straightforward application for a predictive framework like MemDPC. We evaluate our representation on the task of unintentional event classification that is proposed in the recent Oops dataset . The core of unintentional events detection in video is a problem of anomaly detection. Usually, one of the predicted hypotheses tends to match true future relatively well for most of the videos. The discrepancy between them yields a measurement of future predictability, or ‘surprise’ level. A big surprise or a mismatched prediction can be used to locate the failing moment. In detail, we design the model as follows: first, we compute both the predicted feature z^i\hat{z}_{i} and the corresponding true feature ziz_{i}, and let a function ξ(.)\xi(.) to measure their discrepancy. We train the model with two settings: (i) freezing the representation and only train the classifier ξ(.)\xi(.); (ii) finetuning the entire network. The detailed structure for the classification task can be found in Section 5.7.

Experiments

For the self-supervised training, two video action recognition datasets are used, but labels are dropped during training: UCF101 , containing 13k videos spanning over 101 human actions; and Kinetics400 (K400) with 306k 10-second video clips covering 400 human actions. For the downstream tasks we also use UCF101, and additionally we use: HMDB51 containing 7k videos spanning over 51 human actions; and Oops containing 20k videos of daily human activities with unexpected failed moments, among them 14k videos have the time stamps of the failed moments manually labelled.

2 Self-Supervised Training

In our experiment, we use a (2+3D)-ResNet, following , as the encoder f(.)f(.), where the first two residual blocks res2 and res3 have 2D convolutional kernels, and only res4 and res5 have 3D kernels. Specifically, (2+3D)-ResNet18 and (2+3D)-ResNet34 are used in our experiments, denoted as R18 and R34 below. For the temporal aggregation, g(⋅)g(\cdot), we use an one-layer GRU with kernel size 1×11\times 1, with the weights shared among all spatial positions on the feature map. The future prediction function, ϕ(.)\phi(.), is a two-layer convolutional network. We choose the size of the memory bank M\mathbf{M} to be 1024 based on experiments in Table 1. Network architecture are given in the appendix.

For the data, raw videos are decoded at a frame rate 24-30 fps, and each data sample consists of 40 consecutive frames, sampled with a temporal stride of 3 from the raw video. As input to MemDPC, they are divided into 8 video clips – so that each encoder f(.)f(.) inputs 5 frames, covering around 0.5 seconds, and the 40 frames around 4 seconds. For optical flow, in order to eliminate extra supervisory signals in the self-supervised training stage, we use the un-supervised TV-L1 algorithm , and follow the same pre-processing procedures as , i.e\xperiod truncating large motions with more than ±20\pm 20 in both channels, appending a third s channel, and transforming the values from toto. For data augmentation, we apply clip-wise random crop and horizontal flip, and frame-wise color jittering and random greyscale, for both the RGB and optical flow streams. We experiment with both 128×128128\times 128 and 224×224224\times 224 input resolution. The original video resolution is 256×256256\times 256 and it is firstly cropped to 224×224224\times 224 then rescaled if needed. Self-supervised training uses the Adam optimizer with initial learning rate 10−310^{-3}. The learning rate is decayed once to 10−410^{-4} when the validation loss plateaus. We use a batch size of 16 samples per GPU.

3 Supervised Classification

For all action classification downstream tasks, the input follows the same frame sampling procedure as when the model is trained with self-supervised learning, and then we train the classifier with cross-entropy loss as shown in Figure 4. A dropout of 0.9 is applied on the final layer. For data augmentation, we use clip-wise random crop, random horizontal flip, and random color jittering. The classifier is trained with Adam with a 10−310^{-3} initial learning rate, and decayed once to 10−410^{-4} when the validation loss plateaus. During testing, we follow the standard pipeline, i.e\xperiod ten-crop (center and four corner crops, w/o horizontal flip), take the same sequence length as training from the video, and average the prediction from the sampling temporal moving window.

4 Evaluation: Action Classification

We conduct two sets of experiments: (i) ablation studies on the effectiveness of the different modules in the MemDPC, by self-supervised learning on UCF101, (ii) to compare with other state-of-the-art approaches, we run MemDPC on K400 with self-supervised learning. For both settings, the representation quality is evaluated on UCF101 and HMDB51 with linear probing, non-linear probing, and end-to-end finetuning.

In this section, we conduct extensive experiments to validate the effectiveness of compressive memory, bidirectional aggregation, and self-supervised learning on optical flow. Note that, in each experiment, we keep the settings identical, and only vary one variable at a time.

As shown in Table 1, the following phenomena can be observed: First, comparing experiment C2\mathcal{C}2 against B1\mathcal{B}1 (68.268.2 vs. 61.861.8), networks initialized with self-supervised MemDPC clearly present better generalization than a randomly initialized network; Second, comparing with a strong baseline (A\mathcal{A}), the proposed compressive memory boost the learnt representation by around 5%5\% (68.268.2 vs. 63.663.6), and the optimal memory size for UCF101 is 10241024; Third, MemDPC acts as a general learning framework that can also help to boost the generalizability of motion representations, a 7.3%7.3\% boost can be seen from D1\mathcal{D}1 vs. B2\mathcal{B}2 (81.981.9 vs. 74.674.6); Fourth, the bidirectional aggregation provides a small boost to the accuracy by about 1% (E1\mathcal{E}1 vs. C2\mathcal{C}2, E2\mathcal{E}2 vs. D1\mathcal{D}1, E3\mathcal{E}3 vs. D2\mathcal{D}2). Lastly, after fusing both streams, D2\mathcal{D}2 achieves 84% classification accuracy, confirming our claim that self-supervised learning with only the video stream (without additional audio or text streams) can also end up with strong action recognition models.

Comparison with others.

In this section, we train MemDPC on K400 and evaluate the action classification performance on UCF101 and HMDB51. Specifically, we evaluate three settings: (1) finetuning the entire network (denoted as Freeze=✗); (2) freeze the backbone and only train a linear classifier, i.e\xperiod linear probe (denoted as Freeze=✓); (3) freeze the backbone and only train a non-linear classifier, i.e\xperiod non-linear probe (denoted as ‘n.l.’).

As shown in Table 2, for the same amount of data (K400) and visual-only input, MemDPC surpasses all previous state-of-the-art self-supervised methods on both UCF101 and HMDB51 (although there exist small differences in architecture, e.g\xperiod for 3DRotNet, ST-Puzzle, DPC, SpeedNet). When freezing the representation, it can be seen that a non-linear probe gives better results than a linear probe, and in practice a non-linear classifier is still very cheap to train.

Other self-supervised training methods on the same benchmarks are not directly comparable, even ignoring the architecture differences, due to the duration of videos used or to the number of modalities used. For example, CBT uses a longer version of K600 (referred to as K600+ in the table), the size is about 9 times that of the standard K400 that we use, and CBT requires RotNet initialization while MemDPC can be trained from scratch. Nevertheless, our performance exceeds that of CBT. Other works use additional modalities for pre-text tasks like audio , or narrations , and train on larger datasets. Despite these disadvantages, we demonstrate that MemDPC trained with only visual inputs, can achieve competitive results on the finetuning protocol.

5 Evaluation: Data Efficiency

In Figure 5, we show the data efficiency of MemDPC on both RGB input and optical flow with action recognition on the UCF101 dataset. As we reduce the labelled training samples, action classifier trained on MemDPC representation generalize significantly better than the classifier trained from scratch. Also, to match the performance of a random initialized classifier trained on 100% labelled data, a classifier trained on MemDPC initialization only requires less than 50% labelled data for both RGB and optical flow input.

6 Evaluation: Video Retrieval

In this protocol, we evaluate our representation with nearest-neighbour video retrieval, features are extracted from the model, which is only trained with self-supervised learning, no further finetuning is allowed.

Experiments are shown on two datasets: UCF101 and HMDB51. For both datasets, within the training set or within the testing set, multiple clips could be from the same source video, hence they are visually similar and make the retrieval task trivial. We follow the practice of , and use each clip in the test set to query the kk nearest clips in the training set.

For each clip, we sample multiple 88 video blocks with a sliding window, and extract the context representation ctc_{t} for each window. We spatial-pool each ctc_{t} and take the average over all the windows. For distance measurement, we use cosine distance. We report Recall at kk (R@kk) as the evaluation metric. That is, as long as one clip of the same class is retrieved in the top kk nearest neighbours, a correct retrieval is counted.

In Table 3, we show the retrieval performance on UCF101 and HMDB51. Note that the MemDPC benchmarked here is only trained on UCF101, the same as . For fair comparison, MemDPC in this experiment uses a R18 backbone, which has the same depth but less parameters than the 3D-ResNet used in . With RGB inputs, our MemDPC gets state-of-the-art performance on all the metrics except R@1 in UCF101, where the method from Buchler et al\xperiod specializes well on R@1. While for Flow inputs, MemDPC significantly outperforms all previous methods by a large margin. We also qualitatively show video retrieval results in the Appendix.

7 Evaluation: Unintentional Actions

We evaluate MemDPC on the Oops dataset on unintentional action classification. In Oops, there is one failure moment in the middle of each video. When cutting the video into short clips, the clip overlapping the failure moment is defined as a ‘transitioning’ action, the clips before are ‘intentional’ actions, and the clips afterwards are ‘unintentional’ actions. The core task is therefore to classify each short video clip into one of three categories,

In this experiment, we use a R18 based MemDPC model that takes 128×128128\times 128 resolution video frames as input. After MemDPC is trained on K400 and the Oops training set videos with self-supervised learning, we further train it for unintentional action classification with a linear probe, and end-to-end finetuning (as shown in Table 4). The training details are given in the appendix. State-of-the-art performance is demonstrated by our MemDPC on this unintentional action classification task, even outperforming the model pretrained on K700 with full supervision with finetuning.

Conclusion

In this paper, we propose a new architecture and learning framework (MemDPC) for self-supervised learning from video, in particular for representations for action recognition. With the novel compressive memory, the model can efficiently handle the nature of multiple hypotheses in the self-supervised predictive learning procedure. In order to thoroughly evaluate the quality of the learnt representation, we conduct experiments on four different downstream tasks, namely action recognition, video retrieval, learning with scarce annotations, and unintentional action classification. In all cases, we demonstrate state-of-the-art or competitive performance over other approaches that use orders of magnitude more training data. Above all, for the first time, we show that it is possible to learn high-quality video representations with self-supervised learning, from the visual stream alone (without additional audio or text streams).

Funding for this research is provided by a Google-DeepMind Graduate Scholarship, and by the EPSRC Programme Grant Seebibyte EP/M013774/1. We would like to thank João F. Henriques, Samuel Albanie and Triantafyllos Afouras for helpful discussions.

References

Appendix 0.A Architectures in detail

This section gives the architectural details of the MemDPC components, including the encoder f(.)f(.) and temporal aggregator g(.)g(.).

The detailed architecture of the encoder function f(.)f(.) is shown in Table 5. The size of the convolutional kernel is denoted by [temporal×spatial2, channel][\text{temporal}\times\text{spatial}^{2}\text{, channel}]. The strides are denoted by [temporal, spatial2][\text{temporal, }\text{spatial}^{2}]. We assume the input is 8 video blocks, 5 frames per video block, and the frame has a resolution of 128×128128\times 128. The column of ‘output size’ shows the tensor dimension after the current stage. Some pooling layers are omitted for clarity. We will release all the source code after the paper decision.

The detailed architecture of the temporal aggregation function g(.)g(.) is shown in Table 6. It aggregates the feature maps over the past TT time steps. Table 6 shows the case where g(.)g(.) aggregates the feature maps over the past 5 steps. We use the same convention as above to denote convolutional kernel size. The temporal aggregator is an one-layer ConvGRU that returns context feature with same number of channels as its input.

Appendix 0.B Details of unintentional action classification

Figure 6 shows the architecture of the unintentional action classification task. In detail, at each time step tt, the single-directional MemDPC framework produces a feature ztz_{t} from the current input xtx_{t}, and also a predicted feature z^t\hat{z}_{t} from the past input x1,...,xt−1x_{1},...,x_{t-1}. For each time step, two features ztz_{t} and z^t\hat{z}_{t} are concatenated and passed to a linear classifier ξ(.)\xi(.) to classify into one of three categories (intentional, transitioning or unintentional actions).

Naturally in the Oops dataset the distribution of the three classes is very unbalanced, e.g\xperiod transitioning actions are very rare comparing with the other two categories. To handle this we oversample transitioning actions during training. For testing, we take the same sequence length as training from the video with a temporal moving window, and summarize the prediction.

The classifier is trained with cross entropy loss and optimized by the Adam optimizer with a 10−310^{-3} learning rate. The learning rate is decayed once to 10−410^{-4} when the validation loss plateaus.

Appendix 0.C Video retrieval results

Appendix 0.D Pseudocode of MemDPC

In this section, we give the core implementation of MemDPC in PyTorch-like style, including the compressive memory bank and the computation of the contrastive loss. We will release all the source code.

Appendix 0.E Visualization of learned memory

This section visualizes the memory learned by MemDPC. In Figure 8, we use single memory entry as the query to retrieve videos in the feature space. It shows the memory may have captured certain features in the video like repetitive textures or wide background. Figure 9 shows the t-SNE clustering results of the memory addressing probability pt+1p_{t+1} from Equation 3. It shows that as the training progresses, the network attends to different memory entries to predict futures for different action categories, although category information is not involved during training.