Video Transformer Network
Daniel Neimark, Omri Bar, Maya Zohar, Dotan Asselmann
Introduction
Attention matters. For almost a decade, ConvNets have ruled the computer vision field . Applying deep ConvNets produced state-of-the-art results in many visual recognition tasks, i.e., image classification , object detection , semantic segmentation , object instance segmentation , face recognition and video action recognition . But, recently this domination is starting to crack as transformer-based models are showing promising results in many of these tasks .
Video recognition tasks also rely heavily on ConvNets. In order to handle the temporal dimension, the fundamental approach is to use 3D ConvNets . In contrast to other studies that add the temporal dimension straight from the input clip level, we aim to move apart from 3D networks. We use state-of-the-art 2D architectures to learn the spatial feature representations and add the temporal information later in the data flow by using attention mechanisms on top of the resulting features. Our approach input only RGB video frames and without any bells and whistles (e.g., optical flow, streams lateral connections, multi-scale inference, multi-view inference, longer clips fine-tuning, etc.) achieves comparable results to other state-of-the-art models.
Video recognition is a perfect candidate for Transformers. Similar to language modeling, in which the input words or characters are represented as a sequence of tokens , videos are represented as a sequence of images (frames). However, this similarity is also a limitation when it comes to processing long sequences. Like long documents, long videos are hard to process. Even a 10 seconds video, such as those in the Kinetics-400 benchmark , are processed in recent studies as short, seconds, clips.
But how does this clip-based inference would work on much longer videos (i.e., movie films, sports events, or surgical procedures)? It seems counterintuitive that the information in a video of hours, or even a few minutes, can be grasped using only a snippet clip of a few seconds. Nevertheless, current networks are not designed to share long-term information across the full video.
VTN’s temporal processing component is based on a Longformer . This type of transformer-based model can process a long sequence of thousands of tokens. The attention mechanism proposed by the Longformer makes it feasible to go beyond short clip processing and maintain global attention, which attends to all tokens in the input sequence.
In addition to long sequence processing, we also explore an important trade-off in machine learning – speed vs. accuracy. Our framework demonstrates a superior balance of this trade-off, both during training and also at inference time. In training, even though wall runtime per epoch is either equal or greater, compared to other networks, our approach requires much fewer passes of the training dataset to reach its maximum performance; end-to-end, compared to state-or-the-art networks, this results in a faster training. At inference time, our approach can handle both multi-view and full video analysis while maintaining similar accuracy. In contrast, other networks’ performance significantly decreases when analyzing the full video in a single pass. In terms of GFLOPS x Views, their inference cost is considerably higher than those of VTN, which concludes to a fewer GFLOPS and a faster validation wall runtime.
Our framework’s structure components are modular (Fig. 1). First, the 2D spatial backbone can be replaced with any given network. The attention-based module can stack up more layers, more heads or can be set to a different Transformers model that can process long sequences. Finally, the classification head can be modified to facilitate different video-based tasks, like temporal action localization.
Related Work
Most recent studies in video recognition suggested architectures that are based on 3D ConvNets . In , a two-stream architecture was used, one stream for RGB inputs and another for Optical Flow (OF) inputs. Residual connections are inserted into the two-stream architecture to allow a direct link between RGB and OF layers. The idea of inflating 2D ConvNets into their 3D counterpart (I3D) was introduced in . I3D takes 2D ConvNets and expands its layers into 3D. Therefore it allows to leverage pre-trained state-of-the-art image recognition architectures in the spatial-temporal domain and apply them for video-based tasks.
Non-local Neural Networks (NLN) introduced a non-local operation, a type of self-attention, that computes responses based on relationships between different locations in the input signal. NLN demonstrated that the core attention mechanism in Transformers can produce good results on video tasks, however it is confined to processing only short clips. In order to extract long temporal context, introduced a long-term feature bank that acts as the entire video memory and a Feature Bank Operator (FBO) that computes interactions between short-term and long-term features. However, it requires precomputed features, and it is not efficient enough to support end-to-end training of the feature extraction backbone.
SlowFast explored a network architecture that operates in two pathways and different frame rates. Lateral connections fuse the information between the slow pathway, focused on the spatial information, and the fast pathway focused on temporal information.
The X3D study builds on top of SlowFast. It argues that in contrast to image classification architectures, which have been developed via a rigorous evolution, the video architectures have not been explored in detail, and historically are based on expanding image-based networks to fit the temporal domain. X3D introduces a set of networks that progressively expand in different axes, e.g., temporal, frame rate, spatial, width, bottleneck width, and depth. Compared to SlowFast, it offers a lightweight network (in terms of GFLOPS and parameters) with similar performance.
Transformers in computer vision.
The Transformers architecture reached state-of-the-art results in many NLP tasks, making it the de-facto standard. Recently, Transformers are starting to disrupt the field of computer vision, which traditionally depends on deep ConvNets. Studies like ViT and DeiT for image classification , DETR for object detection and panoptic segmentation , and VisTR for video instance segmentation are some examples showing promising results when using Transformers in the computer vision field. Binding these results with the sequential nature of video makes it a perfect match for Transformers.
Applying Transformers on long sequences.
BERT and its optimized version RoBERTa are transformer-based language representation models. They are pre-trained on large unlabeled text and later fine-tuned on a given target task. With minimal modification, they achieve state-of-the-art results on a variety of NLP tasks.
One significant limitation of these models, and Transformers in general, is their ability to process long sequences. This is due to the self-attention operation, which has a complexity of per layer ( is sequence length) .
Longformer addresses this problem and enables lengthy document processing by introducing an attention mechanism with a complexity of . This attention mechanism combines a local-context self-attention, performed by a sliding window, and task-specific global attention.
Similar to ConvNets, stacking up multiple windowed attention layers results in a larger receptive field. This property of Longformer gives it the ability to integrate information across the entire sequence. The global attention part focuses on pre-selected tokens (like the token) and can attend to all other tokens across the input sequence.
Video Transformer Network
Video Transformer Network (VTN) is a generic framework for video recognition. It operates with a single stream of data, from the frames level up to the objective task head. In the scope of this study, we demonstrate our approach using the action recognition task by classifying an input video to the correct action category.
The architecture of VTN is modular and composed of three consecutive parts. A 2D spatial feature extraction model (spatial backbone), a temporal attention-based encoder, and a classification MLP head. Fig. 1 demonstrates our architecture layout.
VTN is scalable in terms of video length during inference, and enables the processing of very long sequences. Due to memory limitation, we suggest several types of inference methods. (1) Processing the entire video in an end-to-end manner. (2) Processing the video frames in chunks, extracting features first, and then applying them to the temporal attention-based encoder. (3) Extracting all frames’ features in advance and then feed them to the temporal encoder.
The spatial backbone operates as a learned feature extraction module. It can be any network that works on 2D images, either deep or shallow, pre-trained or not, convolutional- or transformers-based. And its weights can be fixed (pre-trained) or trained during the learning process.
2 Temporal attention-based encoder
As suggested by , we use a Transformer model architecture that applies attention mechanisms to make global dependencies in a sequence data. However, Transformers are limited by the number of tokens they can process at the same time. This limits their ability to process long inputs, such as videos, and incorporate connections between distant information.
In this work, we propose to process the entire video at once during inference. We use an efficient variant of self-attention, that is not all-pairwise, called Longformer . Longformer operates using sliding window attention that enables a linear computation complexity. The sequence of feature vectors of dimension (Sec. 3.1) is fed to the Longformer encoder. These vectors act as the 1D tokens embedding in the standard Transformer setup.
Like in BERT we add a special classification token () in front of the features sequence. After propagating the sequence through the Longformer layers, we use the final state of the features related to this classification token as the final representation of the video and apply it to the given classification task head. Longformer also maintains global attention on that special token.
3 Classification MLP head
Similar to , the classification token (Sec. 3.2) is processed with an MLP head to provide a final predicted category. The MLP head contains two linear layers with a GELU non-linearity and Dropout between them. The input token representation is first processed with a Layer normalization.
4 Looking beyond a short clip context
The common approach in recent studies for video action recognition uses 3D-based networks. During inference, due to the addition of a temporal dimension, these networks are limited by memory and runtime to clips of a small spatial scale and a low number of frames. In , the authors use the whole video during inference, averaging predictions temporally. More recent studies that achieved state-of-the-art results processed numerous, but relatively short, clips during inference. In , inference is done by sampling ten clips evenly from the full-length video and average the softmax scores to achieve the final prediction. SlowFast follows the same practice and introduces the term “view” – a temporal clip with a spatial crop. SlowFast uses ten temporal clips with three spatial crops at inference time; thus, 30 different views are averaged for the final prediction. X3D follows the same practice, but in addition, it uses larger spatial scales to achieve its best results on 30 different views.
This common practice of multi-view inference is somewhat counterintuitive, especially when handling long videos. A more intuitive way is to “look” at the entire video context before deciding on the action, rather than viewing only small portions of it. Fig. 2 shows 16 frames extracted evenly from a video of the abseiling category. The actual action is obscured or not visible in several parts of the video; this might lead to a false action prediction in many views. The potential in focusing on the segments in the video that are most relevant is a powerful ability. However, full video inference produces poor performance in methods that were trained using short clips (Table 3 and 4). In addition, it is also limited in practice due to hardware, memory, and runtime aspects.
Video Action Recognition with VTN
In order to evaluate our approach and the impact of context attention on video action recognition, we use several spatial backbones pre-trained on 2D images.
Combining the state-of-the-art image classification model, ViT-Base , as the backbone in VTN. We use a ViT-Base network that was pre-trained on ImageNet-21K. Using ViT as the backbone for VTN produces an end-to-end transformers-based network that uses attention both for the spatial and temporal domains.
R50/101-VTN.
As a comparison, we also use a standard 2D ResNet-50 and ResNet-101 networks , pre-trained on ImageNet.
DeiT-B/BD/Ti-VTN.
Since ViT-Base was trained on ImageNet-21K we also want to compare VTN by using similar networks trained on ImageNet. We use the recent work of and apply DeiT-Tiny, DeiT-Base, and DeiT-Base-Distilled as the backbone for VTN.
1 Implementation Details
The spatial backbones we use were pre-trained on either ImageNet or ImageNet-21k. The Longformer and the MLP classification head were randomly initialized from a normal distribution with zero mean and 0.02 std. We train the model end-to-end using video clips. These clips are formed by choosing a random frame as the starting point, then sampling 2.56 or 5.12 seconds as the video’s temporal footprint. The final clip frames are subsampled uniformly to a fixed number of frames , depending on the setup.
For the spatial domain, we randomly resize the shorter side of all the frames in the clip to a $224\times 224$. Horizontal flip is also applied randomly on the entire clip.
The ablation experiments were done on a 4-GPU machine. Using a batch size of 16 for the ViT-VTN (on 16 frames per clip input) and a batch size of 32 for the R50/101-VTN. We use an SGD optimizer with an initial learning rate of and a different learning rate reduction policy, steps-based for the ViT-VTN versions and cosine schedule decay for the R50/101-VTN versions. In order to report the wall runtime, we use an 8-V100-GPU machine.
For the Longformer, we use an effective attention window of size 32, which was applied for each layer. Two other hyperparameters are the dimensions set for the Hidden size and the FFN inner hidden size. These are a direct derivative of the spatial backbone. Therefore, in R50/101-VTN we use 2048 and 4096, respectively, and for ViT-B-VTN we use 768 and 3072, respectively. In addition, we apply Attention Dropout with a probability of 0.1. We also explore the impact of the number of Longformer layers.
The positional embedding (PE) information is only relevant for the temporal attention-based encoder (Fig. 1). We explore three positional embedding approaches (Table 2b): (1) Learned positional embedding - since a clip is represented using frames taken from the full video sequence, we can learn an embedding that uses as input the frame location (index) in the original video, giving the Transformer information regarding the position of the clip in the entire sequence; (2) Fixed absolute encoding - we use a similar method to the one in DETR , and modified it to work on the temporal axis only; and (3) No positional embedding - no information is added in the temporal dimension, but we still use the global position to mark the special token position.
Inference.
In order to show a comparison between different models, we use both the common practice of inference in multi-views and a full video inference approach (Sec. 3.4).
In the multi-view approach, we sample 10 clips evenly from the video. For each clip, we first resize the shorter side to 256, then take three crops of size from the left, center, and right. The result is 30 views per video, and the final prediction is an average of all views’ softmax scores.
In the full video inference approach, we read all the frames in the video. Then, we align them for batching purposes, by either sub- or up-sampling, to 250 frames uniformly. In the spatial domain, we resize the shorter side to 256 and take a center crop of size .
Experiments
The original Kinetics-400 dataset consists of 246,535 training videos and 19,761 validation videos. Each video is labeled with one of 400 human action categories, curated from YouTube videos. Since some YouTube links are expired, we could only download 234,584 of the original dataset, thus missing 11,951 videos from the training set, which are about 5%. This leads to a slight drop in performance of about 0.5%https://github.com/facebookresearch/video-nonlocal-net/blob/master/DATASET.md.
In the validation set, we are missing one video. To test our data’s validity and compare it to previous studies, we evaluated the SlowFast-8X8-R50 model, published in PySlowFast, on our validation data. We got 76.45% top1-accuracy vs. the reported 77%, thus a drop of 0.55%. This drop might be related to different FFmpeg encoding and rescaling of the videos. From this point forward, when comparing to other networks, we report results taken from the original studies except when we evaluate them on the full video inference in which we use our validation set. All our approach results are reported based on our validation set.
Spatial backbone variations.
We start by examining how different spatial backbone architectures impact VTN performance. Table 1 shows a comparison of different VTN variants and the pretrain dataset the backbone was first trained on. ViT-B-VTN is the best performing model and reaches 78.6% top-1 accuracy and 93.7% top-5 accuracy. The pretraining dataset is important. Using the same ViT backbone, only changing between DeiT (pre-trained on ImageNet) and ViT (pre-trained on ImageNet-21K) we get an improvement in the results.
Longformer depth.
Next, we explore how the number of attention layers impacts the performance. Each layer has 12 attention heads and the backbone is ViT-B. Table 2a shows the validation top-1 and top-5 accuracy for 1, 3, 6, and 12 attention layers. The comparison shows that the difference in performance is small. This is counterintuitive to the fact that deeper is better. It might be related to the fact that Kinetics-400 videos are relatively short, around 10 seconds. We believe that processing longer videos will benefit from a large receptive field obtained by using a deeper Longformer.
Longformer positional embedding.
In Table 2b we compare three different positional embedding methods, focusing on learned, fixed, and no positional embedding. All versions are done with a ViT-B-VTN, a temporal footprint of 5.12 seconds, and a clip size of 16 frames. Surprisingly, the one without any positional embedding achieved slightly better results than the fixed and learned versions.
As this is an interesting result, we also use the same trained models and evaluate them after randomly shuffling the input frames only in the validation set videos. This is done by first taking the unshuffled frame embeddings, then shuffle their order, and finally add the positional embedding. This raised another surprising finding, in which the shuffle version gives better results, reaching 78.9% top-1 accuracy on the no positional embedding version. Even in the case of learned embeddings it does not have a diminishing effect. Similar to the Longformer depth, we believe that this might be related to the relatively short videos in Kinetics-400, and longer sequences might benefit more from positional information. We also argue that this could mean that Kinetics-400 is primarily a static frame, appearance based classification problem rather than a motion problem .
Temporal footprint and number of frames in a clip.
We also explore the effect of using longer clips in the temporal domain and compare a temporal footprint of 2.56 vs. 5.12 seconds. And also how the number of frames in the clip impact the network performance. The comparison is done on a ViT-B-VTN with one attention layer in the Longformer. Table 2c shows that top-1 and top-5 accuracy are similar, implying that VTN is agnostic to these hyperparameters.
Finetune the 2D spatial backbone.
Instead of fine-tuning the spatial backbone, by continuing the back-propagation process, when training VTN, we can use a frozen 2D network solely for feature extraction. Table 2d shows the validation accuracy when training a ViT-B-VTN with three attention layers with and without also training the backbone. Fine-tuning the backbone improves the results by 7% in Kinetics-400 top-1 accuracy.
Does attention matter?
A key component in our approach is the impact of attention functionally on the way VTN perceives the full video sequence. To convey this impact we train two VTN networks, using three layers in the Longformer, but with a single head for each layer. In one network the head is trained as usual, while in the second network instead of computing attention based on query/key dot products and softmax, we replace the attention matrix with a hard-coded uniform distribution that is not updated during back-propagation.
Fig. 4 shows the learning curves of these two networks. Although the training has a similar trend, the learned attention performs better. In contrast, the validation of the uniform attention collapses after a few epochs demonstrating poor generalization of that network. Further, we visualize the token attention weights by processing the same video from Fig. 2 with the single-head trained network and depicted, in Fig. 3, all the weights of the first attention layer aligned to the video’s frames. Interestingly, the weights are much higher in segments related to the abseiling category. In Appendix A. we show a few more examples.
Training and validation runtime.
An interesting observation we make concerns the training and validation wall runtime of our approach. Although our networks have more parameters, and therefore, are longer to train and test, they are actually much faster to converge and reach their best performance earlier. Since they are evaluated using a single view of all video frames, they are also faster during validation. Table 3 shows a comparison of different models and several VTN variants. Compared to the state-of-the-art SlowFast model, our ViT-B-VTN with one layer achieves almost the same results but completes an epoch faster while requiring fewer epochs. This accumulates to a faster end-to-end training. The validation wall runtime is also faster due to the full video inference approach.
To better demonstrate the fast convergence of our approach, we wanted to show an apples-to-apples comparison of different training and evaluating curves for various models. However, since other methods use the multi-view inference only post-training, but use a single view evaluation while training their models, this was hard to achieve. Thus, to show such comparison and give the reader additional visual information, we trained a NL I3D (pre-trained on ImageNet) with a full video inference protocol during validation (using our codebase and reproduced the original model results). We compare it to DeiT-B-VTN which was also pre-trained on ImageNet. Fig. 5 shows that the VTN-based network converges to better results much faster than the NL I3D and enables a much faster training process compared to 3D-based networks.
Data augmentation.
Recent studies showed that data augmentation significantly improves the performance of transformers-based models . To demonstrate its impact on VTN, we apply extensive data augmentation as suggested in DeiT and RandAugment . Table 3 shows that our method reaches 79.8% top-1 accuracy, a 1.2% improvement vs. the same model trained without such augmentations. Training with augmentations requires 10 more epochs but didn’t impact the training wall runtime.
Final inference computational complexity.
Finally, we examine what is the final inference computational complexity for various models by measuring GFLOPs. Although other models need to evaluate multiple views to reach their highest performance, ViT-B-VTN performs almost the same for both inference protocols. Table 4 shows a significant drop of about 8% when evaluating the SlowFast-8X8-R50 model using the full video approach. In contrast, ViT-B-VTN maintains the same performance while requiring, end-to-end, fewer GFLOPs at inference.
2 Experiments on Moments in Time
The Moments in Time (MiT) dataset is a large-scale collection of short (3 seconds) videos . MiT is a challenging dataset, with state-of-the-art results just above 34% top-1 accuracy . In this work, we use MiT-v2, consisting of 727,305 training videos and 30,500 validation videos. Each video is labeled with one of 305 classes of dynamic events. Although previous studies worked on MiT-v1 (802,264 training videos, 33,900 validation videos, 339 classes), this dataset is no longer available. In Table 3, we show the results of various models on MiT-v1 and MiT-v2. Since the relation between v1 and v2 in terms of performance was not established and thus unknown, we also trained our implementation of NL I3D on MiT-v2 using RGB inputs and achieved comparable results to those of I3D (RGB+OF) published on MiT-v1 . Furthermore, ViT-B-VTN achieves the highest top-1 accuracy on MiT-v2 while using only RGB frames as input.
Conclusion
We presented a modular transformer-based framework for video recognition tasks. Our approach introduces an efficient way to evaluate videos at scale, both in terms of computational resources and wall runtime. It allows full video processing during test time, making it more suitable for dealing with long videos. Although current video classification benchmarks are not ideal for testing long-term video processing ability, hopefully, in the future, when such datasets become available, models like VTN will show even larger improvements compared to 3D ConvNets.
We thank Ross Girshick for providing valuable feedback on this manuscript and for helpful suggestions on several experiments.