VidTr: Video Transformer Without Convolutions

Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, Joseph Tighe

Introduction

We introduce Video Transformer (VidTr) with separable-attention, one of the first transformer-based video action classification architecture that performs global spatio-temporal feature aggregation. Convolution-based architectures have dominated the video classification literature in recent years , and although successful, the convolution-based approaches have two drawbacks: 1. they have limited receptive field on each layer and 2. information is slowly aggregated through stacked convolution layers, which is inefficient and might be ineffective . Attention is a potential candidate to overcome these limitations as it has a large receptive field which can be leveraged for spatio-temporal modeling. Previous works use attention to modeling long-range spatio-temporal features in videos but still rely on convoluational backbones . Inspired by recent successful applications of transformers on NLP and computer vision , we propose a transformer-based video network that directly applies attentions on raw video pixels for video classification, aiming at higher efficiency and better performance.

We evaluated our VidTr on 6 most commonly used datasets, including Kinetics 400/700, Charades, Something-something V2, UCF-101 and HMDB-51. Our model achieved state-of-the-art (SOTA) or comparable performance on five datasets with lower computational requirements and latency compared to previous SOTA approaches. Our error analysis and ablation experiments show that the VidTr works significantly better than I3D on activities that requires longer temporal reasoning (e.g. making a cake vs. eating a cake), which aligns well with our intuition. This also inspires us to ensemble the VidTr with the I3D convolutional network as features from global and local modeling methods should be complementary. We show that simply combining the VidTr with a I3D50 model (8 frames input) via ensemble can lead to roughly a 2% performance improvement on Kinetics 400. We further illustrate how and why the VidTr works by visualizing the separable-attention using attention rollout , and show that the spatial-attention is able to focus on informative patches while temporal attention is able to reduce the duplicated/non-informative temporal instances. Our contributions are:

Video transformer: We propose to efficiently and effectively aggregate spatio-temporal information with stacked attentions as opposed to convolution based approaches. We introduce vanilla video transformer as proof of concept with SOTA comparable performance on video classification.

VidTr: We introduce VidTr and its permutations, including the VidTr with SOTA performance and the compact-VidTr with significantly reduced computational costs using the proposed standard deviation based pooling method.

Results and model weights: We provide detailed results and analysis on 6 commonly used datasets which can be used as reference for future research. Our pre-trained model can be used for many down-streaming tasks.

Related Work

The early research on video based action recognition relies on 2D convolutions . The LSTM was later proposed to model the image feature based on ConvNet features . However, the combination of ConvNet and LSTM did not lead to significantly better performance. Instead of relying on RNNs, the segment based method TSN and its permutations were proposed with good performance.

Although 2D network was proved successful, the spatio-temporal modeling was still separated. Using 3D convolution for spatio-temporal modeling was initially proposed in and further extended to the C3D network . However, training 3D convnet from scratch was hard, initializing the 3D convnet weights by inflate from 2D networks was initially proposed in I3D and soon proved applicable with different type of 2D network . The I3D was used as backbone for many following work including two-stream network , the networks with focus on temporal modeling , and the 3D networks with refined 3D convolution kernels .

The 3D networks are proved effective but often not efficient, the 3D networks with better performance often requires larger kernels or deeper structures. The recent research demonstrates that depth convolution significantly reduce the computation , but depth convolution also increase the network inference latency. TSM and TAM proposed a more efficient backbone for temporal modeling, however, such design couldn’t achieve SOTA performance on Kinetics dataset. The neural architecture search was proposed for action recognition recently with competitive performance, however, the high latency and limited generalizability remain to be improved.

The previous methods heavily rely on convolution to aggregate features spatio-temporally, which is not efficient. A few previous work tried to perform global spatio-temporal modeling but still limited by the convolution backbone. The proposed VidTr is fundamentally different from previous works based on convolutions, the VidTr doesn’t require heavily stacked convolutions for feature aggregation but efficiently learn feature globally via attention from first layer. Besides, the VidTr don’t rely on sliding convolutions and depth convolutions, which runs at less FLOPs and lower latency compared with 3D convolutions .

2 Vision Transformer

The transformers was previously proposed for NLP tasks and recently adopted for computer vision tasks. The transformers were roughly used in three different ways in previous works: 1.To bridge the gap between different modalities, e.g. video captioning , video retrieval and dialog system . 2. To aggregate convolutional features for down-streaming tasks, e.g. object detection , pose estimation , semantic segmentation and action recognition . 3. To perform feature learning on raw pixels, e.g. most recently image classification .

Action recognition with self-attention on convolution features is proved successful, however, convolution also generates local feature and gives redundant computations. Different from and inspired by very recent work on applying transformer on raw pixels , we pioneer the work on aggregating spatio-temporal feature from raw videos without relying on convolution features. Different from very recent work that extract spatial feature with vision transformer on every video frames and then aggregate feature with attention, our proposed method jointly learns spatio-temporal feature with lower computational cost and higher performance. Our work differs from the concurrent work , we present a split attention with better performance without requiring larger video resolution nor extra long clip length. Some more recent work further studied the multi-scale and different attention factorization methods.

Video Transformer

We introduce the Video Transformer starting with the vanilla video transformer (section 3.1) which illustrates our idea of video action recognition without convolutions. We then present VidTr by first introducing separable-attention (section 3.2), and then the attention pooling to drop non-representative information temporally (section 3.2).

2 VidTr

To address these memory constraints, we introduce a multi-head separable-attention (MSA) by decoupling the 3D self-attention to a spatial attention MSA⁡s\operatorname{MSA}_{s} and a temporal attention MSA⁡t\operatorname{MSA}_{t} (Figure 1):

2.2 Temporal Down-sampling method

Video content usually contains redundant information , with multiple frames depicting near identical content over time. We introduce compact VidTr (C-VidTr) by applying temporal down-sampling within our transformer architecture to remove some of this redundancy. We study different temporal down-sampling methods (poolpool in Eq. 3) including temporal average pooling and 1D convolutions with stride 2, which reduce the temporal dimension by half (details in Table 5d).

A limitation of these pooling the methods is that they uniformly aggregate information across time but often in video clips the informative frames are not uniformly distributed. We adopted the idea of non-uniform temporal feature aggregation from previous work . Different from previous work that directly down-sample the query using average pooling, we found that in our proposed network, the temporal attention highly activates on a small set of temporal features when the clip is informative, while the attention equally distributed over the length of the clip when the clip caries little additional semantic information. Building on this intuition, we propose a topK based pooling (topK_stdtopK\_std pooling) that orders instances by the standard deviation of each row in the attention matrix:

3 Implementation Details

Model Instantiating: Based on the input clip length and sample rate, we introduce three base VidTr models (VidTr-S,VidTr-M and VidTr-L). By applying the different pooling strategies we introduce two compact VidTr permutations (C-VidTr-S, and C-VidTr-M). To normalize the feature space, we apply layer normalization before and after the residual connection of each transformer layer and adopt the GELU activation as suggested in . Detailed configurations can be found in Table 1. We empirically determined the configuration for different clip length to produce a set of models from low FLOPs and low latency to high accuracy (details in Compact VidTr down-sampling twice starting from layer 11 and skipping different number of layers.).

During training we initialize our model weights from ViT-B . To avoid over fitting, we adopted the commonly used augmentation strategies including random crop, random horizontal flip (except for Something-something dataset). We trained the model using 64 Tesla V100 GPUs, with batch size of 6 per-GPU (for VidTr-S) and weight decay of 1e-5. We adopted SGD as the optimizer but found the Adam optimizer also gives us the same performance. We trained our network for 50 epochs in total with initial learning rate of 0.01, and reduced it by 10 times after epochs 25 and 40. It takes about 12 hours for VidTr-S model to converge, the training process also scales well with fewer GPUs (e.g. 8 GPUs for 4 days). During inference we adopted the commonly used 30-crop evaluation for VidTr and compact VidTr, with 10 uniformly sampled temporal segments and 3 uniformly sampled spatial crop on each temporal segment . It is worth mentioning that we can further boost the inference speed of compact VidTr by adopting a single pass inference mechanise, this is because the attention mechanism captures global information more effectively than 3D convolution. We do this by training a model with frames sampled in TSN style, and uniformly sampling NN frames in inference (details in supplemental materials).

Experimental Results

We evaluate our method on six of the most widely used datasets. Kinetics 400 and Kinetics 700 consists of approximately 240K/650K training videos and 20K/35K validation videos trimmed to 10 seconds from 400/700 human action categories. We report top-1 and top-5 classification accuracy on the validation sets. Something-Something V2 dataset consists of 174 actions and contains 168.9K training videos and 24.7K evaluation videos. We report top-1 accuracy following previous works evaluation setup. Charades has 9.8k training videos and 1.8k validation videos spanning about 30 seconds on average. Charades contains 157 multi-label classes with longer activities, performance is measured in mean Average Precision (mAP). UCF-101 and HMDB-51 are two smaller datasets. UCF-101 contains 13320 videos with an average length of 180 frames per video and 101 action categories. The HMDB-51 has 6,766 videos and 51 action categories. We report the top-1 classification on the validation videos based on split 1 for both dataset.

2 Kinetics 400 Results

We report results on the validation set of Kinetics 400 in Table 2, including the top-1 and top-5 accuracy, GFLOPs (Giga Floating-Point Operations) and latency (ms) required to compute results on one view.

As shown in Table 2, the VidTr achieved the SOTA performance compared to previous I3D based SOTA architectures with lower GFLOPs and latency. The VidTr significantly outperform previous SOTA methods at roughly same computational budget, e.g. at 200 GFLOPs, the VidTr-M outperform I3D50 by 3.6%3.6\%, NL50 by 2.1%2.1\%,and TPN50 by 0.9%0.9\%. At similar accuracy levels, VidTr is significantly more computationally efficient than other works, e.g. at 78%78\% top-1 accuracy, the VidTr-S has 6×\times fewer FLOPs than NL-101, 2×\times fewer FLOPs than TPN and 12%12\% fewer FLOPs than Slowfast-101. We also see that our VidTr outperforms I3D based networks at higher sample rate (e.g. s=8s=8, TPN achieved 76.1%76.1\% top-1 accuracy), this denotes, the global attention learns temporal information more effectively than 3D convolutions. X3D-XXL from architecture search is the only network that outperforms our VidTr. We plan to use architecture search techniques for attention based architecture in future work.

2.2 Compact VidTr

We evaluate the effectiveness of our compact VidTr with the proposed temporal down-sampling method (Table 1). The results (Table 1) show that the proposed down-sampling strategy removes roughly 56%56\% of the computation required by VidTr with only 2%2\% performance drop in accuracy. The compact VidTr complete the VidTr family from small models (only 39GFLOPs) to high performance models (up to 79.1% accuracy). Compared with previous SOTA compact models , our compact VidTr achieves better or similar performance with lower FLOPs and latency, including: TEA (+0.6% with 16% fewer FLOPs) and TEINet (+0.5% with 11% fewer FLOPs).

2.3 Error and Ensemble Analysis

We compare the errors made by VidTr-S and the I3D50 network to better understand the local networks’ (I3D) and global networks’ (VidTr) behavior. We provide the top-5 activities that our VidTr-S gain most significant improvement over the I3D50. We find that our VidTr-S outperformed the I3D on the activities that requires long-term video contexts to be recognized. For example, our VidTr-S outperformed the I3D50 on “making a cake” by 26% in accuracy. The I3D50 overfits to “cakes” and often recognize making a cake as eating a cake. We also analyze the top-5 activities where I3D does better than our VidTr-S (Table 4). Our VidTr-S performs poorly on the activities that need to capture fast and local motions. For example, our VidTr-S performs 21% worse in accuracy on “shaking head”.

Inspired by the findings in our error analysis, we ensembled our VidTr with a light weight I3D50 network by averaging the output values between the two networks. The results (Table 2) show that the the I3D model and transformer model complements each other and the ensemble model roughly lead to 2% performance improvement on Kinetics 400 with limited additional FLOPs (37G). The performance gained by ensembling the VidTr with I3D is significantly better than the improvement by combine two 3D networks (Table 2).

2.4 Ablations

We perform all ablation experiments with our VidTr-S model on Kinetics 400. We used 8×224×2248\times 224\times 224 input with a frame sample rate of 8, and 30-view evaluation. Patching strategies: We first compare the cubic patch (4×1624\times 16^{2}), where the video is represented as a sequence of spatio-temporal patches, with the square patch (1×1621\times 16^{2}), where the video is represented as a sequence of spatial patches. Our results (Table 5a) show that the model using cubic patches with longer temporal size has fewer FLOPs but results in significant performance drop (73.1 vs. 75.5). The model using square patches significantly outperform all cubic patch based models, likely because the linear embedding is not enough to represent the shot-term temporal association in the cubic. We further compared the performance of using different patch sizes (1×1621\times 16^{2} vs. 1×3221\times 32^{2}), using 32232^{2} patches lead to 4×4\times decreasing of the sequence length, which decreases memory consumption of the affinity matrices by 16×16\times, however, using 16216^{2} patches significantly outperform the model using 32232^{2} patches (77.7 vs. 71.2). We did not evaluate the model using smaller patching sizes (e.g., 8×88\times 8) because of the high memory consumption. Attention Factorization: We compare different factorization for attention design, including spatial modeling only (WH), jointly spatio-temporal modeling module (WHT, vanilla-Tr), spatio-temporal separable-attention (WH + T, VidTr), and axial separable-attention (W + H + T). We first evaluate an spatio-only transformer. We average the class token for each input frame for our final output. Our results (Table 5b) show that the spatio-only transformer requires less memory but has worse performance compare with spatio-temporal attention models. This shows that temporal modeling is critical for attention based architectures. The joint spatio-temporal transformer significantly outperforms the spatio-only transformer but requires a restrictive amount of memory (T2T^{2} times for the affinity matrices). Our VidTr using spatio-temporal separable-attention requires 3.3×3.3\times less memory with no accuracy drop. We further evaluate the axial separable-attention (W + H + T), which requires the least memory. The results (Table 5b) show that the axial separable-attention has a significant performance drop likely due to breaking the X and Y spatial dimensions. Sequence down-sampling comparison: We compare different down-sampling strategy including temporal average pooling, 1D temporal convolution and the proposed STD-based topK pooling method. The results (Table 5d) show that our proposed STD-based down-sampling method outperformed the temporal average pooling and the convolution-based down-sampling strategies that uniformly aggregate information over time.

Backbone generalization: We evaluate our VidTr initialized with different models, including T2T , ViT-B, and ViT-L. The results on Table 5c show that our VidTr achieves reasonable performance across all backbones. The VidTr using T2T as the backbone has the lowest FLOPs but also the lowest accuracy. The Vit-L-based VidTr achieve similar performance with the Vit-B-based VidTr even with 3×3\times FLOPs. As showed in previous work , transformer-based network are more likely to over-fit and Kinetics-400 is relatively small for Vit-L-based VidTr. Where to down-sample: Finally we study where to perform temporal down-sampling. We perform temporal down-sampling at different layers (Table 5e). Our results (Table 5e) show that starting to perform down-sampling after the first encoder layer has the best trade-off between the performance and FLOPs. Starting to perform down-sampling at very beginning leads to the fewest FLOPs but has a significant performance drop (72.9 vs. 74.9). Performing down-sampling later only has slight performance improvement but requires higher FLOPs. We then analyze how many layers to skip between two down-sample layers. Based on the results in Table 5f, skipping one layer between two down-sample operations has the best trade-off. Performing down-sampling on consecutive layers (0 skip layers) has lowest FLOPs but the performance decreases (73.9 vs. 74.9). Skipping more layers did not show significant performance improvement but does have higher FLOPs.

2.5 Run-time Analysis

We further analyzed the trade-off between latency, FLOPs and accuracy. We note that the VidTr achieved the best balance between these factors (Figure 2). The VidTr-S achieve similar performance but significantly fewer FLOPs compare with I3D101-NL (5×5\times fewer FLOPs), Slowfast101 8×88\times 8 (12%12\% fewer FLOPs), TPN101 (2×2\times fewer FLOPs), and CorrNet50 (20×20\times fewer FLOPs). Note that the X3D has very low FLOPs but high latency due to the use of depth convolution. Our experiments show that the X3D-L has about 3.6×3.6\times higher latency comparing with VidTr-S (Figure 2).

3 More Results

Our experiments show a consistent performance trend on Kinetics 700 (Table 6). The VidTr-S significantly outperformed the baseline I3D model (+9%), the VidTr-M achieved the performance comparable to Slowfast101 8×88\times 8 and the VidTr-L is comparable to previous SOTA slowfast101-nonlocal. There is a small performance gap between our model and Slowfast-NL , because Slowfast is pre-trained on both Kinetics 400 and 600 while we only pre-trained on Kinetics 400. Previous findings that VidTr and I3D are being complementary is consistent on Kinetics 700, ensemble VidTr-L with I3D leads to +0.6% performance boost. Charades Results: We compare our VidTr with previous SOTA models on Charades. Our VidTr-L outperformed previous SOTA methods LFB and NUTA101, and achieved the performance comparable to Slowfast101-NL (Table 6). The results on Charades demonstrates that our VidTr generalizes well to multi-label activity datasets. Our VidTr performs worse than the current SOTA networks (X3D-XL) on Charades likely due to overfitting. As discussed in previous work , the transformer-based networks overfit easier than convolution-based models, and Charades is relatively small. We observed a similar finding with our ensemble, ensembling our VidTr with a I3D network (40.3 mAP) achieved SOTA performance. Something-something V2 Results: We observe that the VidTr does not work well on the something-something dataset (Table 6), likely because pure transformer based approaches do not model local motion as well as convolutions. This aligns with our observation in our error analysis. Further improving local motion modeling ability is an area of future work. UCF and HMDB Results: Finally we train our VidTr on two small dataset UCF-101 and HMDB-51 to test if VidTr generalizes to smaller datasets. The VidTr achieved SOTA comparable performance with 6 epochs of training (96.6% on UCF and 74.4% on HMDB), showing that the model generalize well on small dataset (Table 6).

Visualization and Understanding VidTr

We first visualized the VidTr’s separable-attention with attention roll-out method (Figure 3a). We find that the spatial attention is able to focus on informative regions and temporal attention is able to skip the duplicated/non-representative information temporally. We then visualized the attention at 4th, 8th and 12th layer of VidTr (Figure 3b), we found the spatial attention is stronger on deeper layers. The attention does not capture meaningful temporal instances at early stages because the temporal feature relies on the spatial information to determine informative temporal instances. Finally we compared the I3D activation map and rollout attention from VidTr (Figure 3c). The I3D mis-classified the catching fish as sailing, as the I3D attention focused on the people sitting behind and water. The VidTr is able to make the correct prediction and the attention showed that the VidTr is able to focus on the action related regions across time.

Conclusion

In this paper, we present video transformer with separable-attention, an novel stacked attention based architecture for video action recognition. Our experimental results show that the proposed VidTr achieves state-of-the-art or comparable performance on five public action recognition datasets. The experiments and error analysis show that the VidTr is especially good at modeling the actions that requires long-term reasoning. Further combining the advantage of VidTr and convolution for better local-global action modeling and adopt self-supervised training on large-scaled data will be our future work.

References