Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
Mandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, João F. Henriques
Introduction
Transformers have become a popular architecture across NLP , vision and speech . The self-attention mechanism in the transformer works well for different types of data and across domains. However, its generic nature and its lack of inductive biases also mean that transformers typically require extremely large amounts of data for training , or aggressive domain-specific augmentations . This is particularly true for video data, for which transformers are also applicable , but where statistical inefficiencies are exacerbated. While videos carry rich temporal information, they can also contain redundant spatial information from neighboring frames. Vanilla self-attention applied to videos compares pairs of image patches extracted at all possible spatial locations and frames. This can lead it to focus on the redundant spatial information rather than the temporal information, as we show by comparing normalization strategies in our experiments.
We therefore contribute a variant of self-attention, called trajectory attention, which is better able to characterize the temporal information contained in videos. For the analysis of still images, spatial locality is perhaps the most important inductive bias, motivating the design of convolutional networks and the use of spatial encodings in vision transformers . This is a direct consequence of the local structure of the physical world: points that belong to the same 3D object tend to project to pixels that are close to each other in the image. By studying the correlation of nearby pixels, we can thus learn about the objects.
Videos are similar, except that 3D points move over time, thus projecting on different parts of the image along certain 2D trajectories. Existing video transformer methods disregard these trajectories, pooling information over the entire 3D space-time feature volume , or pooling axially across the temporal dimension . We contend that pooling along motion trajectories would provide a more natural inductive bias for video data, allowing the network to aggregate information from multiple views of the same object or region, to reason about how the object or region is moving (for example, the linear and angular velocities), and to be invariant to camera motion.
We leverage attention itself as a mechanism to find these trajectories. This is inspired by methods such as RAFT , which showed that excellent estimates of optical flow can be obtained from the correlation volume obtained by comparing local features across space and time. We observe that the joint attention mechanism for video transformers computes such a correlation volume as an intermediate result. However, subsequent processing collapses the volume without consideration for its particular structure. In this work, we seek instead to use the correlation volume to guide the network to pool information along motion paths.
We also note that visual transformers operate on image patches which, differently from individual pixels, cannot be assumed to correspond to individual 3D points and thus to move along simple 1D trajectories. For example, in Figure 1, depicting the action ‘kicking soccer ball’, the ball spans up to four patches, depending on the specific video frame. Furthermore, these patches contain a mix of foreground (the ball) and background objects, thus at least two distinct motions. Fortunately, we are not forced to select a single putative motion: the attention mechanism allows us to assemble a motion feature from all relevant ‘ball regions’.
Inspired by Nyströmformer , we also propose a principled approximation to self-attention, Orthoformer. Our approximation sets state-of-the-art performance on the recent Long Range Arena (LRA) benchmark for evaluating efficient attention approximations and generalizes beyond the video domain to long text and high resolution images, with lower FLOPS and memory requirements compared to alternatives, Nyströmformer and Performer . Combining our approximation with trajectory attention allows us to significantly improve its computational and memory efficiency. With our contributions, we set state-of-the-art results on four video action recognition benchmarks.
Related Work
Hand-crafted features were originally used to convert video data into a representation amenable to analysis by a shallow linear model. Such representations include SIFT-3D , HOG3D , and IDT . Since the breakthrough of AlexNet on the ImageNet classification benchmark , which demonstrated the empirical benefits of deep neural networks to learn representations end-to-end, there have been many attempts to do the same for video. Architectures with 3D convolutions—3D-CNNs—were originally proposed to learn deep video representations . Since then, improvements to this paradigm include the use of ImageNet-inflated weights , the space-time decomposition of 3D convolutions , channel-separated convolutions , non-local blocks , and attention layers . Optical flow-based pooling can be used instead of temporal convolutions to improve the representation’s robustness to camera and object motions . Our approach shares this motivation.
The transformer architecture , originally proposed for natural language processing, has recently gained traction in the computer vision domain. The vision transformer (ViT) decomposes an image into a sequence of words and uses a multi-layer transformer to perform image classification. To improve ViT’s data efficiency, DeiT used distillation from a strong teacher model and aggressive data augmentation. Transformers have also been used in a variety of vision image tasks, such as image representation learning , image generation , object detection , video question-answering , few-shot learning , and image–text , video-text , and video-audio representation learning.
The self-attention operation proposed in the transformer have been adapted to video recognition tasks. Wang et al. propose the non-local mean operation for video action recognition, which is equivalent to the standard transformer self-attention applied uniformly across space and time, while our proposed trajectory attention does not treat the space and time dimensions equivalently. Zhao et al. propose a CNN architecture that explicitly predicts trajectories and aggregates information along them using a convolution operation. In contrast, our transformer architecture does not explicitly predict trajectories, but instead provides an inductive bias that encourages the network to consider motion trajectories where useful. Concurrent works have also adapted the self-attention operation to the spatio-temporal nature of videos, however, these approaches do not have a mechanism for reasoning about motion paths, treating time as just another dimension, unlike our approach.
Due to the quadratic complexity of self-attention, there has been a significant amount of research on how to reduce its computational complexity with respect to time and memory use. Sparse attention mechanisms were used to reduce self-attention complexity to , and locality-sensitivity hashing was used by Reformer to further reduce this to . More recently, linear attention mechanisms have been introduced, namely Longformer , Linformer , Performer and Nyströmformer . The Long Range Arena benchmark was recently introduced to compare these different attention mechanisms.
There are many approaches that aim to establish explicit correspondences between video frames as a way to reason about camera and object motion. For short-range correspondences across time, optical flow algorithms are highly effective. In particular, RAFT showed the effectiveness of an all-pairs inter-frame correlation volume as an encoding, which is essentially an attention map. All-pairs intra-frame correlations were subsequently shown to help resolve correspondence ambiguities . For longer-range correspondences, object tracking by repeated detection and data association can be used. In contrast to these approaches, our work does not explicitly establish temporal correspondences, but facilitates implicit correspondence learning via trajectory attention. Jabri et al. estimate correspondences in a similar way, framing the problem as a contrastive random walk on a graph and apply explicit guidance via a cycle consistency loss. Incorporating such guidance into a video transformer is an interesting direction.
Trajectory Attention for Video Data
We now have a set of tokens that form the input to a sequence of transformer layers that, as in ViT , consist of Layer Norm (LN) operations , multi-head attention (MHA) , residual connections , and a feed-forward network (MLP):
In the next section, we shall focus on a single head of the attention operation, and demonstrate how self-attention can realize a suitable inductive bias for video data. For clarity of exposition, we abuse the notation slightly, neglecting the layer norm operation and using the same dimensions for single-head attention as for multi-head attention.
In this way, each query is compared to all keys using dot products, the results are normalized using the softmax operator, and the weights thus obtained are used to average the values corresponding to the keys. Compared to a standard transformer, we have omitted for brevity the softmax temperature parameter and instead assume that the queries and keys have been divided by .
One issue with this formulation is that it has quadratic complexity in both space and time, i.e., . An alternative is to restrict attention to either space or time (called divided space-time attention):
This reduces the complexity to and , respectively, but only allows the model to analyse time and space independently. This is usually addressed by interleaving or stacking the two attention modules in a sequence.
Note that the attention in this formula is applied spatially (index ) and independently for each frame. Intuitively, this pooling operation implicitly seeks the location of the trajectory at time by comparing the trajectory query to the keys at time .
Once the trajectories are computed, we further pool them across time to reason about intra-frame information/connections. To do so, the trajectory tokens are projected to a new set of queries, keys and values as usual:
Like joint space-time attention, our approach has quadratic complexity in both space and time, , so has no computational advantage and is in fact slower than divided space-time attention. However, we demonstrate better accuracy than both joint and divided space-time attention mechanisms. We also provide fast approximations in Section 3.2. A flowchart of the full trajectory attention operation is shown in tensor form in Figure 2.
2 Approximating attention
The aim for prototype-based attention approximation schemes is to use as few prototypes as possible while reconstructing the attention operation as accurately as possible. As such, it behooves us to select prototypes efficiently. We have two priorities for the prototypes: to dynamically adjust to the query and key vectors so that their region of space is well-reconstructed, and to minimize redundancy. The latter is important because the relative probability of a query–key pair may be over-estimated if many prototypes are clustered near that query and key. To address these criteria, we incrementally build a set of prototypes from the set of queries and keys such that a new prototype is maximally orthogonal to the prototypes already selected, starting with a query or key at random. This greedy strategy is dynamic, since it selects prototypes from the current set of queries and keys, and has high entropy, since it preferences well-separated prototypes. Moreover, it balances speed and performance by using a greedy strategy, rather than finding a globally-optimal solution to the maximum entropy sampling problem , making it suitable for use in a transformer.
Naïvely applying prototype-based attention approximation techniques to video transformers would involve creating a unique set of prototypes for each frame in the video. However, additional memory savings can be realized by sharing prototypes across time. Since there is significant information redundancy between frames, video data is opportune for compression via temporally-shared prototypes.
The proposed approximation algorithm is outlined in Algorithm 1. The attention matrix is approximated using intermediate prototypes, selected as the most orthogonal subset of the queries and keys, given a desired number of prototypes . To avoid a linear dependence on the sequence length , we first randomly subsample queries and keys, for a constant , before selecting the most orthogonal subset, resulting in a complexity quadratic in the number of prototypes . The algorithm then computes two attention matrices, much smaller than the original problem, and multiplies them with the values. The most related approach in the literature is Nyströmformer attention, outlined in Algorithm 2. This approach involves a pseudoinverse to attenuate the effect of near-parallel prototypes, has more operations, and a greater memory footprint.
3 The Motionformer model
Our full video transformer model builds on previous work, as shown in Table 1. In particular, we use the ViT image transformer model as the base architecture, the separate space and time positional encodings of TimeSformer , and the cubic image tokenization strategy as in ViViT . These design choices are ablated in Section 4. The crucial difference for our model is the trajectory attention mechanism, with which we demonstrate greater empirical performance than the other models.
Experiments
Kinetics is a large-scale video classification dataset consisting of short clips collected from YouTube, licensed by Google under Creative Commons. As it is a dataset of human actions, it potentially contains personally identifiable information such as faces, names and license plates. Something–Something V2 is a video dataset containing more than 200,000 videos across 174 classes, with a greater emphasis on short temporal clips. In contrast to Kinetics, the background and objects remain consistent across different classes, and therefore models have to reason about fine-grained motion signals. We verified the importance of temporal reasoning on this dataset by showing that a single frame model gets significantly worse results, a decrease of top-1 accuracy. In contrast, a drop of only is seen on the Kinetics-400 dataset, showing that temporal information is much less relevant there. We obtained a research license for this data from https://20bn.com; the data was collected with consent. Epic Kitchens-100 is an egocentric video dataset capturing daily kitchen activities. The highest scoring verb and action pair predicted by the network constitutes an action, for which we report top-1 accuracy. The data is licensed under Creative Commons and was collected with consent by the Epic Kitchens teams.
We follow a standard training and augmentation pipeline , as detailed in the appendix. For ablations, our default Motionformer model is the Vision Transformer Base architecture (ViT/B), pretrained on ImageNet-21K , patch-size with central frame initialization , separate space-time positional embedding and our trajectory attention. The base architecture has 12 layers, 12 attention heads, and an embedding dimension of 768. Our default Motionformer model operates on videos with temporal stride 4 i.e. temporal extent of 2s. For comparisons with state-of-the-art, we report results on two additional variants: Motionformer-HR, which has a high spatial resolution ( videos with temporal stride 4 i.e. temporal extent of 2s), and Motionformer-L, which has a long temporal range ( videos with temporal stride 3 i.e. temporal extent of 3s). Experiments with the large ViT architecture are deferred to the appendix.
1 Ablation studies
We consider the effect of different input tokenization approaches for both joint and trajectory attention on Kinetics-400 (K-400) and Something–Something V2 (SSv2) in Table 2(b). For patch tokenization (), we use inputs of size , while for cubic tokenization (), we use inputs of size to ensure that the model has the same number of input tokens over the same temporal range of 2 seconds. For both attention types, we see that cubic tokenization gives a accuracy improvement over square tokenization on SSv2, a dataset for which temporal information is critical. Furthermore, our proposed trajectory attention using cubic tokenization outperforms joint space-time attention on both datasets.
Here, we ablate using a joint or separate (default) space-time positional encoding in Table 2(b). Similar to the results for input tokenization, the choice of positional encoding is particularly important for the fine-grained motion dataset, SSv2. Since joint space-time attention treats tokens in the space-time volume equally, it benefits particularly from separating the positional encodings, allowing it to differentiate between space and time dimensions, with a improvement on SSv2 over joint space-time encoding. Our proposed trajectory attention elicits a more modest improvement of from using separated positional encodings on SSv2, and outperforms joint space-time attention in both settings on both datasets.
We compare our proposed trajectory attention to joint space-time attention , and divided space-time attention in Table 4. Our trajectory attention (bottom row) outperforms both alternatives on the K-400 and SSv2 datasets. While we see only modest improvements on the appearance cue-reliant K-400 dataset, our trajectory attention significantly outperforms () the other approaches on the motion cue-reliant SSv2 dataset. This dataset requires fine-grained motion understanding, something explicitly singled out by previous video transformer works as a challenge for their models. In contrast, our trajectory attention excels on this dataset, indicating that its motion-based design is able to capture some of this information.
We ablate two design choices for our trajectory attention: the per-frame softmax normalization and the 1D temporal attention. Unlike joint space-time attention, which normalizes the attention map over all tokens in space and time, trajectory attention normalizes independently per frame, allowing us to implicitly track the trajectories of query patches in time. In row 5 of Table 4, we ablate the benefits of this design choice. We observe a reduction of on K-400 and on SSv2 by normalizing over space and time (NormST) compared with normalizing over space alone (NormS). In row 4, we show the benefit of using 1D temporal attention (AttT) to aggregate temporal features, compared to average pooling (AvgT). We observe reductions of on K-400 and on SSv2 when using average pooling instead of temporal attention applied to the motion trajectories, although it saves computing the additional query/key/value projections.
2 Orthoformer approximated attention
In Table 3(a), we compare our Orthoformer algorithm to alternative strategies: Nyströmformer and Performer . Our algorithm performs comparably with Nyströmformer with a reduced memory footprint. In Table 5, we also compare these attention mechanisms on the Long Range Arena benchmark to show applicability to other tasks and data types. Orthoformer is able to effectively approximate self-attention, outperforming the state-of-the-art despite using far fewer prototypes (64) and so gaining significant computational and memory benefits.
A key part of our Orthoformer algorithm is the prototype selection procedure. In Table 3(b), we ablate three prototype selection strategies: segment-means, random, and greedy most-orthogonal selection. Segment-means, the strategy used in Nyströmformer, performs poorly because it can generate multiple parallel prototypes, which will over-estimate the relative probability of query–key pairs near those redundant prototypes. In contrast, our proposed strategy of selecting the most orthogonal prototypes from the query and key set works the best across both datasets, because it explicitly minimises prototype redundancy with respect to direction.
In Table 3(c), we show that Orthoformer improves monotonically as the number of prototypes is increased. In particular, we see an average performance improvement of on both datasets as we increase the number of prototypes from 16 to 128.
In Table 3(d), we demonstrate the memory savings and performance benefits of sharing prototypes across time. On SSv2, we observe a improvement in performance and a decrease in memory usage. These gains may be attributed to the regularization effect of having prototypes leverage redundant information across frames.
The Orthoformer attention approximation algorithm allows us to train larger models and higher resolution inputs for a given GPU memory budget. Here, we verify this, by training a large vision transformer model (ViT-L/16) with a higher resolution input ( pixels) on the Kinetics-400 dataset, using the Orthoformer approximation with 196 temporally-shared prototypes and the same schedule as the base model. We use a fixed patch size (in pixels) for all models, and so the number of input tokens to the transformer scales with the square of the image resolution. As shown in Table 7, this model achieves a competitive accuracy without fine-tuning the training schedule, hyperparameters or data augmentation strategy. We expect that fine-tuning these on a validation set would greatly improve the model’s performance, based on results from contemporary work . Obviously such a parameter sweep is more time-consuming for these large models, however, these preliminary results are indicative that higher accuracies are attainable if these parameters were to be optimized.
3 Comparison to the state-of-the-art
In Table 6, we compare our method against the current state-of-the-art on four common benchmarking datasets: Kinetics-400, Kinetics-600, Something–Something v2 and Epic-Kitchens. We find that our method performs favorably against current methods, even when compared against much larger models such as ViViT-L. In particular, it achieves strong top-1 accuracy improvements of and for SSv2 and Epic-Kitchen Nouns, respectively. These datasets require greater motion reasoning than Kinetics and so are a more challenging benchmark for video action recognition.
Conclusion
We have presented a new general-purpose attention block for video data that aggregates information along implicitly determined motion trajectories, lending a realistic inductive bias to the model. We further address its quadratic dependence on the input size with a new attention approximation algorithm that significantly reduces the memory requirements, the largest bottleneck for transformer models. With these contributions, we obtain state-of-the-art results on several benchmark datasets. Nonetheless, our approach inherits many of the limitations of transformer models, including poor data efficiency and slow training. Specific to this work, trajectory attention has higher computational complexity than alternative attention operations used for video data. This is attenuated by the proposed approximation algorithm, with significantly reduced memory and computation requirements. However, its runtime is bottlenecked by prototype selection, which is not easily parallelized.
There are many applications of trajectory attention beyond video action classification, such as those tasks where temporal context is highly important. We see significant potential for using trajectory attention for tracking , temporal action localization and online action detection , among other settings, and leave these as avenues for future work.
One negative impact of this research is the significant environmental impact associated with training transformers, which are large and compute-expensive models. Compared to 3D-CNNs where the compute scales linearly with the sequence length, video transformers scale quadratically. To mitigate this, we proposed an approximation algorithm with linear complexity that greatly reduces the computational requirements. There is also potential for video action recognition models to be misused, such as for unauthorized surveillance.
Acknowledgments and Disclosure of Funding
We are grateful for support from the Rhodes Trust (M.P.), the European Research Council Starting Grant (IDIU 638009, D.C.), Qualcomm Innovation Fellowship (Y.A.), the Royal Academy of Engineering (RF201819/18/163, J.H.), and EPSRC Centre for Doctoral Training in Autonomous Intelligent Machines & Systems (EP/L015897/1, M.P. and Y.A.). Funding for M.P. was received under his Oxford affiliation. We thank Bernie Huang, Dong Guo, Rose Kanjirathinkal, Gedas Bertasius, Mike Zheng Shou, Mathilde Caron, Hugo Touvron, Benjamin Lefaudeux, Haoqi Fan, and Geoffrey Zweig from Facebook AI for their help, support, and discussion around this project. We also thank Max Bain and Tengda Han from VGG for fruitful discussions.
References
Appendix
In the main paper (and below in Section 6.1.2), we provide evidence that action classification on the Something–Something V2 (SSv2) dataset is more reliant on motion cues than the Kinetics dataset , where appearance cues dominate and a single-frame model achieves high accuracy. Improved performance on SSv2 is one way to infer that our model makes better use of temporal information, however, here we consider another way. We artificially adjust the speed of the video clips by changing the temporal stride of the input. A larger stride simulates faster motions, with adjacent frames being more different. If our trajectory attention is able to make better use of the temporal information in the video than the other attention mechanisms, we expect the margin of improvement to increase as the temporal stride increases. As shown in Figure 3, this is indeed what we observe, with the lines diverging as temporal stride increases, especially for the motion cue-reliant SSv2 dataset. Since the same number of frames are used as input in all cases, the larger the stride, the more of the video clip is seen by the model. This provides additional confirmation that seeing a small part of a Kinetics video is usually enough to classify it accurately, as shown on the bottom left, where the absolute accuracy is reported.
1.2 How important are motion cues for classifying videos from the Kinetics-400 and Something–Something V2 datasets?
To determine the relative importance of motion cues compared to appearance cues for classifying videos on two of the major video action recognition datasets (Kinetics-400 and Something–Something V2), we trained a single frame vision transformer model and compare the results to a multi-frame model that can reason about motion. The single frame was sampled from the video at random. Table LABEL:abl_table:ssv2_k400 shows that single-frame action classifiers can do almost as well as video action classifiers on the Kinetics-400 dataset, implying that the motion information is much less relevant. In contrast, classifying videos from the Something-Something V2 dataset clearly requires this motion information. Therefore, to excel on the SSv2 dataset, a model must reason about motion information. Our model, which introduces an inductive bias that favors pooling along motion trajectories, is able to do this and sees corresponding performance gains.
1.3 Which classes is the performance difference larger with and without the trajectory attention?
The class labels with the largest performance increase (given in parentheses) on the Something-Something v2 dataset are: “Spilling [something] next to [something]” (18%), “Pretending to put [something] underneath [something]” (15%), and “Trying to pour [something] into [something], but missing so it spills next to it” (14%). The classes with the largest performance decrease are: “Putting [something] that can’t roll onto a slanted surface, so it stays where it is” (10%), “Putting [something] on a flat surface without letting it roll” (9%), and “Showing a photo of [something] to the camera” (8%). It is apparent that classes involving predominantly stationary objects do not benefit from trajectory attention, as we would expect.
1.4 Trajectory attention maps
In Figure 4, we show qualitative results of the intermediate attention maps of our trajectory attention operation. The learned attention maps appear to implicitly track the query points across time, a strategy that is easier to learn with the inductive bias instilled by trajectory attention.
1.5 How long does it take to train Motionformer model?
For Table 3c, using Motionformer with orthoformer approximation, 16 prototypes take 384 GPU hours, 64 prototypes take 800 GPU hours, and 128 prototypes take 1216 GPU hours. For the Kinetics-400 state-of-the-art table, the Mformer-B model took 384 GPU hours, the Mformer-L took 1334 GPU hours, and Mformer-HR model took 1376 GPU hours to train. Our baseline Mformer-B model, which outperforms TimeSformer-B by over , takes similar GPU hours (416 (ours) vs the 416 GPU hours reported in Table 2 of TimeSformer). We cannot directly compare to ViViT because they didn’t report training time, but they used a very large transformer (24 layers) compared to ours (12 layers) and so we expect the training time for their approach to be significantly greater.
1.6 Semi-supervised Video Object Segmentation on DAVIS 2017
We evaluate our baseline Motionformer Kinetics-pretrained model (16x16 with Trajectory Attention) on the semi-supervised video object segmentation task on the DAVIS 2017 dataset as in Jabri et al. in Table 9. We directly use the attention maps of our Motionformer model in the label propagation setting, as in . We report mean (m) of standard boundary alignment (F) and region similarity (J) metrics. We attain a competitive J&F-Mean of 60.6. For comparison, DINO obtains J&F-Mean of 62.3 with the same architecture (ViT-B/16x16), but by using a self-supervised learning task on IM-1K. We expect that we could significantly improve the performance by using an 8x8 patch size, as this was shown to be highly effective for the task .
2 Implementation details
During training, we randomly sample clips of size at a rate of from FPS videos, thereby giving an effective temporal resolution of just over seconds. We normalize the inputs with mean and standard deviation 0.5, rescaling in the range $$. We use standard video augmentations such as random scale jittering, random horizontal flips and color jittering. For smaller datasets such as Something–Something V2 and Epic-Kitchens, we additionally apply rand-augment . During testing, we uniformly sample 10 clips per video and apply a 3 crop evaluation .
For all datasets, we use the AdamW optimizer with weight decay , a batch size per GPU of 4, label smoothing with alpha and mixed precision training . For Kinetics-400/600 and Something-Something V2, we train for 35 epochs, with an initial learning rate of , which we decay by 10 at epochs 20, 30. As Epic-Kitchens is a smaller dataset, we use a longer schedule and train for epochs with decay at and .
For the Long-Range Arena benchmark , we used the training, validation, and testing code and parameters from the Nyströmformer Github repository. The Performer implementation was ported over to PyTorch from the official Github repo, and the Nyströmformer implementation was used directly from its Github repository.
Ablation experiments were run on a GPU cluster using 4 nodes (32 GPUs) with an average training time of 12 hours. Experiments for comparing with state-of-the-art models used 8 nodes (64 GPUs), with an average training time of 7 hours.
For our code implementation, we used the timm library for our base vision transformer implementation, and the PySlowFast library for training, data processing, and the evaluation pipeline.