BEVT: BERT Pretraining of Video Transformers
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, Lu Yuan
Introduction
Transformers have become the dominant network structures in the natural language processing (NLP) field and made tremendous success in different NLP tasks. Recently, the pioneering work ViT proposes to tokenize one image into a series of patch-based tokens and apply the transformer architecture for image recognition. Many approaches further demonstrate the power of transformers as generic vision backbones and achieve state-of-the-art performance on various vision tasks. Beyond image tasks, there are also a few studies showing the promise of transformers for video understanding .
The key to the success of Transformers in NLP is BERT pretraining , one of the most successful pretraining tasks, which predicts masked tokens in corrupted texts. This motivates a few recent studies to explore the BERT-style pretraining for image representation learning by recovering raw pixels or latent codes of masked image patches. However, how to leverage such a training strategy for video understanding has never been explored before.
In this paper, we study BERT pretraining of video transformers. Unlike static images, videos depict how objects move and interact over time. Such dynamic nature brings additional difficulty for representation learning. It is often found that learning representations from scratch on videos is computationally expensive and requires extremely large-scale datasets with millions of samples , if not hundreds of millions of samples . Instead of training from scratch, a few methods demonstrate that self-supervised models pretrained on image datasets benefit video recognition under both supervised and unsupervised settings . These approaches simply leverage pretrained models as better initializations to learn spatial-temporal features in videos. While widely used and sometimes effective, the spatial context relationships learned from the image prertaining phase are likely to be drastically modified during video feature learning.
We argue that spatial priors encoded in pretrained self-supervised models should be explicitly preserved when performing video representation learning. The intuition behind is that there are large inter-class variations among different videos and their dependencies on what discriminative information to use (i.e., spatial and temporal clues) to make correct predictions differ. For instance, for actions like “applying lipstick”, spatial knowledge is generally sufficient, as evidenced by the fact simply using 2D features offers decent results on datasets like Kinetics . On the other hand, temporal dynamics are crucial for differentiating actions between two fine-grained diving sequences . This highlights the importance of considering the differences among video samples during feature learning.
In light of this, we introduce BEVT, which decouples video representation learning into spatial representation learning and temporal dynamics learning. More specifically, BEVT builds upon the Video Swin Transformer due to their computationally efficient architectures Note that we only use the architecture and do not load the pretrained weights., and is trained with a BERT-style objective to fully unleash the power of transformers for representation learning. BEVT contains an image stream for spatial modeling and a video stream for temporal modeling, interacting with each other for video modeling. In particular, the image stream, operating on RGB images, learns spatial priors first on ImageNet in an unsupervised fashion by predicting masked image patches in the form of latent codes derived from a pretrained VQ-VAE as in . It is then used to initialize the attention weight matrices of the video stream, whose inputs are sampled video clips, so as to save computation for video transformers. The video stream, on the other hand, learns temporal dynamics in videos through predicting masked 3D tubes represented by latent codes. The two streams, taking image and video pairs as inputs, are then jointly trained on video data through a weight sharing strategy. Such a design not only maintains spatial knowledge learned from image datasets to ensure decent results for static video samples but also learns temporal information to guarantee correct predictions for samples that contain dynamic movements. Finally, BEVT is finetuned on targeted datasets for downstream evaluation.
We conduct extensive experiments on three challenging video datasets, i.e., Kinetics-400 (K400) , Something-Something-v2 (SSv2) , and Diving-48 (Diving48) . On K400, BEVT offers Top-1 accuracy, which is better than the strong supervised baseline . On SSv2 and Diving48, BEVT achieves and Top-1 accuracy outperforming state-of-the-art methods by clear margins. To further analyze the performance difference among these three datasets, we further provide the temporal dependency analysis and demonstrate that videos in K400 mainly rely on spatial clues for correct predictions while videos from SSv2 and Diving48 require more temporal information.
Our main contributions are summarized as follows: (1) We explore the BERT-style training objective to fully unleash the power of transformers to learn discriminative video representations; (2) We introduce a novel two-stream network that decouples spatial representation learning and temporal dynamics learning; (3) We demonstrate different video samples have different preferences towards spatial and temporal clues; (4) We conduct extensive experiments on three challenging video benchmarks and achieve comparable or better results with state-of-the-art methods.
Related Work
Video understanding with CNNs. There is a plethora of work on video understanding with CNNs, most of which focus on learning spatial-temporal features . These approaches can be divided into two categories: (1) temporal aggregation and (2) 3D CNNs. In particular, temporal aggregation methods typically extract image features/scores frame-by-frame and then combine frame-level information to achieve video-level predictions through recurrent networks or average pooling . On the other hand, 3D CNNs extend 2D convolutions into the time domain by using 3D convolutions on stacked RGB frames for the joint-modeling of spatial-temporal relationships . 3D CNNs are generally computationally expensive, and this motivates a line of research on efficient video recognition . Instead of using CNNs, we explore transformers for video understanding due to their strong results on image recognition tasks.
Vision transformers. Motivated by the impressive performance of transformers that are able to capture long-range dependencies in a wide range of natural language processing tasks, there is a growing interest in using transformers for computer vision tasks . More specifically, ViT is the first work that generalizes transformers to the image domain by splitting images into patches that are further embedded with a linear layer as inputs to a Transformer. While demonstrating great potential in image recognition tasks, ViT relies on pretraining on substantially large-scale datasets like ImageNet-21K and the training process is computationally expensive. To mitigate these issues, extensive studies have been introduced. For example, DeiT uses a distillation loss to speed up the training process. PiT incorporates the pooling-based design into transformers. CSwin Transformer proposes multi-head grouping and performs attention within cross-shaped windows. And the recent work Mobile-Former further extends transformer at low FLOP regime. There are also a few very recent studies that extend image transformers for video understanding tasks . Fan et al. use a multi-scale design to generate spatial-temporal tokens in different sizes for action recognition . Liu et al. extend Swin Transformers into the video domain . In this paper, we focus on studying BERT pretraining of video transformer in a self-supervised manner, which is orthogonal to such transformer design efforts.
Self-supervised representation learning. At the core of many computer vision tasks is how to learn discriminative features for targeted datasets. Since collecting labeled datasets is labor-intensive and costly, there is an ever-increasing trend in learning representations in a self-supervised manner . The main idea is to design surrogate tasks including inpainting , colorization , jigsaw predictions , rotation predictions , etc., as a form of supervisory signals in lieu of manual labels. More recently, contrastive learning has been a popular paradigm for feature learning by forcing images to be closer to their augmented copies than other samples . In contrast to these approaches using CNNs as backbones, there are a few very recent studies leveraging contrastive learning for transformers.
BERT pretraining. In contrast to contrastive learning widely used in vision, BERT pretraining is extremely popular and extensively studied in NLP. As an effort that unifies vision and NLP under the same BERT pretraining framework, the recent work BEiT and ICT utilizes the masked image modeling task to do BERT pretraining of image transformers and achieves great success for different tasks. And one concurrent work PeCo further proposes a perceptual codebook to improve the performance. Another concurrent work extends it from recovering patch tokens to raw pixels. In this paper, we study BERT pretraining for video transformers as an orthogonal unifying effort. Different from BERT pretraining of image transformers and the concurrent effort , we decouple video pretraining into spatial representation learning and temporal dynamics learning so as to accommodate the varying need of distinct salient clues for different videos.
Method
The goal of BEVT is to learn video representations effectively for both relatively static videos and dynamic videos in a self-supervised manner. Here, “relatively static videos” mean the videos only requiring discriminative spatial representation for recognition, while “dynamic videos” mean that videos that also require temporal dynamics for recognition. Besides the effectiveness, another key problem to consider in video pretraining is efficiency. Compared to image pretraining, video pretraining is more computationally expensive, thus making pretraining on large-scale video data from scratch inefficient or even inapplicable without massive computational resources.
To this end, BEVT decouples the video pretraining into spatial representation learning and temporal dynamics learning. And the spatial representation learning is only conducted on image data, while the temporal dynamics learning is conducted on video data. To implement this idea, our BEVT contains two streams, operating on images and videos, respectively. In the following, we introduce different components of our framework. Figure 2 gives an overview of our framework.
Masked image and video tokens. Motivated by the great success of BERT in NLP tasks, BEVT is optimized to simultaneously perform masked image modeling (MIM) and masked video modeling (MVM) by predicting “corrupted” image and video tokens, respectively. The MIM is designed to capture spatial priors while the MVM is used to capture temporal dynamics in videos. In particular, for the image stream, since input images are divided into non-overlapping patches, we randomly mask several patches and the image stream is trained to recover them as in . More specifically, the embedded feature of each masked patch is replaced by a learnable mask token embedding. For the video stream, we randomly mask 3D tokens and train the video stream to predict those masked tokens. The set of masked image and video tokens and the remaining patch features are sent to encoders, as will be introduced below.
Mask strategy. For masked image modeling, following , we use blockwise masking instead of randomly selecting each masked patch. When generating masked positions for an image, we mask a block of patches each time and set the minimum number of patches for each block. The position, the aspect ratio and the size of each block are randomly selected under a preset range. We repeat masking blocks until the ratio of masked patches exceeds the preset lower bound. For masked video modeling, we employ a tube masking strategy that is a straightforward extension of blockwise masking. Given an input video clip of length , we first randomly choose the number of masked frames (tube length) and the start frame . Then we employ blockwise masking to generate a 2D mask, and apply this 2D mask to each frame from to . In other words, for each masked frame, the set of masked positions is the same and the shape of the whole 3D mask is a tube. The range of the masked tube length is and the masking ratio of each masked frame is .
BEVT encoders. BEVT contains two encoders, one for the image stream and one for the video stream. Both encoders are instantiated with the Video Swin Transformer due to its strong performance with a moderate computational cost. Note that in contrast to that performs fully-supervised training, we use the Video Swin Transformer as our backbone for self-supervised learning. In particular, Video Swin Transformer follows the design of Swin Transformer and is a hierarchical architecture consisting of four stages. Between every two stages, spatial downsampling is performed by patch merging layers, which concatenates the features of each group of spatially neighboring patches. After downsampling, a linear layer maps the features of each concatenated token to half of their dimension. A series of Swin attention blocks comes after to apply feature transformation.
Given a sequence of tokens as inputs, the video encoder outputs a feature map with the size of . Since Video Swin Transformer only performs temporal downsampling in the beginning linear embedding layer, it degrades to a 2D architecture when the temporal dimension of the input is 1. As a result, for the image encoder, the output feature map has a size of .
BEVT decoders. To learn meaningful representations by predicting the tokens for the masked image and video patches in inputs, BEVT has an image decoder and a video decoder as the auxiliary prediction heads, which will be discarded in finetuning stage. Existing modern vision transformers including Swin Transformer follow the hierarchical design and downsample the input into decreased spatial/temporal resolutions. Taking the VideoSwin of video stream shown in Figure 3 as an example, it consists of four stages, and the feature maps in the last stage have the dimension of . In order to match the dimension of feature maps to the number of groundtruth visual tokens, we design a lightweight decoder for the video stream in BEVT. As shown in Figure 3, it first spatially upsamples the stage-4 feature by using a transposed convolutional layer, and then concatenate the upsampled stage-4 feature with stage-3 feature together and fuse them with a simple linear layer. Finally, the fused feature will be temporally upsampled with another transposed convolution layer.
To predict the token for each position , a simple softmax based classifier is applied upon :
where is the feature vector of the output feature map at position , denotes the corresponding probability vector. and are the weight and the bias of a linear layer. For the decoder in the image stream, it follows a similar design, and the only difference is without the temporally upsampling part.
Training objectives. Denote the positions of masked patches in the input images and videos as and , the objective of masked image modeling is to maximize the log-likelihood of the groundtruth token for each mask position .
The objective of two-stream joint training is a simple combination of two objectives:
where is the hyper-parameter that balances the weights of the image stream and the video stream.
Training strategies. Following our decoupled design, we first train the image stream on ImageNet with the masked image modeling task to learn discriminative spatial representation. The resulting model is then used to initialize the video stream, and both streams are jointly trained by optimizing Equation 5 such that the objective preserves spatial information while learns to capture temporal dynamics in videos. Such a strategy not only makes BEVT much more efficient than pretraining video transformers on large-scale video data from scratch, but also satisfies the need of learning different discriminative clues for different types of video samples.
Weight sharing between streams. When jointly training the image and video stream, instead of learning two sets of model weights for the two streams independently, we design a weight sharing strategy so that they can share model weights for the encoder except some image/video specific parts. This is motivated by the good property of transformer networks, i.e., most operators (including multi-head attention and FFN) are oriented to tokens but not specific input types. Taking the Video Swin transformer as an example, we use the following strategies for weight sharing: (1) We use independent 2D patch partitioning layers instead of 3D patch partitioning, and add a linear embedding layer in the first stage for projecting image tokens to the same dimension as the original 3D video patch embedding; (2) We adapt 3D shifted local window to the 2D scenario. This is fulfilled by reusing the submatrix of the original 3D relative positional embedding where the relative temporal distance is as the 2D relative positional embedding. With such a design, the image stream and the video stream can help each other by optimizing one “mostly-unified” encoder.
Finetuning and inference. Once pretrained, BEVT provides decent video representations that can be transferred for downstream tasks. On targeted datasets, we simply use the 3D patch embedding layers and the video encoder, to which a few task-specific layers (e.g., classification head for video recognition) are appended, for finetuning. The resulting model can then be readily used for inference.
Experiments
Datasets and evaluation metrics. We evaluate our method on three representative video recognition datasets: Kinetics-400 (K400) , Something-Something-v2 (SSv2) , and Diving-48 (Diving48) . K400 contains videos clips from YouTube with an average duration of 10 seconds and the videos are manually labeled into 400 categories. Following , we use videos for training and videos for testing. SSv2 is also a large-scale video dataset that contains videos for training and videos for testing. The videos in SSv2 are labeled into 174 classes the average duration is 4 seconds. Diving48 contains fine-grained diving sequences, which are further split into a training set with around clips and a testing set with clips. Compared to K400, recognizing videos in SSv2 and Diving48 requires more temporal information, as will be introduced below. Following official instructions, we report Top-1 accuracy on all three datasets. And the default resolution is used.
Implementation Details. We use Video Swin-Base for experiments throughout the paper unless mentioned otherwise. For pretraining the image stream BEVT-I alone, we train the model for 800 epochs on ImageNet-1K with a batch size of 2048. For pretraining the video stream BEVT-V alone or the two-stream BEVT, we train the model for 150 epochs on K400 with a batch size of 256 and the clip length is 16. When performing two stream pretraining, ImageNet images are employed to train the image stream with a batch size of 2048, and the loss weight is simply set to . We use the DALL-E tokenizer unless explicitly stated. For downstream tasks, we finetune the pretrained models for 60 epochs with a batch size of 64 and the clip length is 32. We use the AdamW optimizer with a linear warm-up and a cosine learning rate schedule for both pretraining and finetuning. The pretraining on the image stream takes about 8 days on 16 V100 GPUs. The two-stream joint pretraining of 150 epochs takes about 4 days on 32 V100 GPUs.
2 Main Results
Effectiveness of BEVT for video transformer pretraining. To demonstrate the effectiveness of BEVT, we compare it with the four image transformer pretraining baselines: (1) Image Sup: pretraining the image Swin Transformer on Imagenet-1K in a supervised way. Similar strategies are commonly used in existing video transformer papers . (2) Image CL: pretraining the image Swin Transformer on Imagenet-1K with a self-supervised contrastive learning method . (3) BEVT-I: pretraining the image Swin Transformer only with the image stream, which is similar to BEiT . (4) BEVT-V: pretraining the video Swin Transformer only with the video stream. The pretrained weights from Image Sup, Image CL and BEVT-I used as the initialization of the Video Swin Transformer for finetuing. For the video stream, we design two baselines by conducting BERT pretraining on K400 and HowTo100M from scratch, i.e., BEVT-V, which is our framework without the decoupled design. As emphasized before, because video pretraining is more computationly expensive than image pretraining, pretraining on HowTo100M dataset with many epochs is not applicable. For fair comparisons, we also use 32 V100 GPUs to pretrain HowTo100M about 8 days (about 2 epochs).
The comparison results are summarized in Table 1. We observe that: (1) BEVT outperforms the Image Sup baseline by clear margins (4.3% and 2.7%) on SSv2 and Diving48, respectively. This not only suggests that learning representations with BEVT using a BERT-style training objective is promising without the need for manual labels, but also shows that only image-based pretraining is not enough for these two datasets. On K400, the performance of BEVT is on par with the Image Sup baseline. (2) We also see that BEVT offers comparable or better results compared to Image CL on these three datasets. (3) Compared to BEVT-I, BEVT is better on SSv2 and Diving48 by 1.4% and 5.5% respectively, highlighting the gains brought by the video stream. Similarly, BEVT obtains similar results as BEVT-I on K400. (4) Compared to BEVT, BEVT-V pretraining on K400 or HowTo100M from scratch under the similar computation budget achieves much worse results. We hypothesize it may be because the data diversity of K400 is not as good as ImageNet. And for HowTo100M, pretraining much more epochs may help learn better video representation, but it is too costly. This also further justifies the decoupled design in our BEVT.
Deeper dataset analysis. To further understand the performance variations of BEVT among three datasets, we perform a temporal dependency study to investigate the amount of temporal information required for correct predictions. Specifically, we use the following two testing strategies: (1) Single-frame, where we randomly sample a frame and replace all other frames with this one, leading to a static video; (2) Random-Shuffling, where a random shuffling is performed along the temporal axis. The results are summarized in 2. We observe that both strategies have a relatively small impact on K400 compared to SSv2 and Diving48 where there is a 60% and 70% performance drop when using the Single-frame strategy. This suggests most videos in K400 can be recognized by discriminative spatial clues whereas temporal dynamics is particularly important for SSv2 and Diving48. Comparing across Table 1 and Table 2, we make the following conclusions: (1) On datasets like K400 where spatial clues are dominant, finetuning a model with spatial priors, e.g. pretrained on ImageNet, can achieve decent performance. Additional video modeling brings little effect to the overall performance; (2) The use of the video stream in BEVT is crucial to learn necessary temporal information for datasets like SSv2 and Diving48. This confirms our hypothesis that different videos rely on different discriminative clues for accurate predictions due to the large intra-class and inter-class variations among videos.
Comparisons with State-of-the-art methods. We compare BEVT with state-of-the-art methods on SSv2, Diving48 and K400. On SSv2 and Diving48, we see from Table 3 and Table 4 that our approach achieves the best performance by clear margins when compared to existing SOTA methods, including supervised models. It is worth mentioning that, on SSv2, a common practice to achieve better performance is to perform two rounds of pretraining—a model is pretrained on both ImageNet and K400 in a fully supervised fashion before finetuning the model on SSv2. Instead, we pretrain on ImageNet and K400 without using any manual labels, yet our performance is still better. On K400, we see from Table 5 that BEVT achieves competitive results with state-of-the-art methods using similar or less computation measured by GFLOPs.
3 Ablation Study
Below, we provide a set of ablation studies to justify the contribution of different components in our framework.
Importance of image stream pretraining. In our BEVT, we first conduct the image stream pretraining only on the large-scale image data to efficiently learn spatial representations, then use it as the initialization for joint pretraining. To show its importance, we provide some ablation results in Table 6, where the column “Init” means whether to use the image stream pretrained weights as initialization or not. We have some interesting findings: (1) Using the image stream pretrained weights as initialization can benefit both pure video stream pretraining (i.e.,“BEVT-V”) and the following joint pretraining of image stream and video stream (i.e.,“BEVT”). (2) Even with the initializations, jointly training the image stream with the video stream is still necessary and can bring desirable performance gain.
Image data in joint training. By default, when jointly learning spatial and temporal representations, the image stream in BEVT continues to use the ImageNet images as training images. In this ablation, we also experiment with a variant that uses K400 frames for the image stream. The results are shown in Table 7. We see that images from ImageNet are slightly better than those from K400, i.e. less than 0.3% on all three datasets. This suggests that the image stream, designed to preserve spatial knowledge, is not very sensitive to data domains.
Different Pretrained Tokenizer. We also experiment with the PeCo tokenizer instead of the DALL-E tokenizer in BEVT. PeCo is only pretrained on ImageNet-1K and uses the same codebook size as in DALL-E. As the results shown in Table 3-5, PeCo tokenizer outperforms the DALL-E tokenizer on all three datasets and pushes BEVT to higher state-of-the-art performance with 71.4% and 87.2% Top-1 accuracy on SSv2 and Diving48 respectively. This demonstrates that better performance can be achieved with a better visual tokenizer.
Effect of masking strategies. We evaluate the BEVT-V with different masking strategies for the video stream, i.e. the temporal length to be masked and the ratio of masking. The experiments are conducted with Video Swin Tiny for time consideration. In addition to tube masking strategies, we compare with: (1) Random-3D: which samples random patches and masks them following a uniform distribution. (2) Frame-Diff: which uses the same strategy to choose masked frames as tube masking, but applies blockwise masking independently for each frame. 2D masks may be different for different masked frames. (3) Random-Frame: which samples random frames and masks them with the same 2D mask generated by blockwise masking. The temporal positions of masked frames may not be consecutive. The results are summarized in Table 8. We have several observations: (1) Masking tubes offers the better results compared to other masking methods like Random-3D and Random-Frame. (2) Setting too small tube temporal length (e.g., [0.25T, 0.75T]) or too large temporal length (e.g., T) will both incur inferior results on SSv2. We guess it is because the former setting will make the masked video modeling too easy while the later will degrade to masked image model to some extent.(3) Applying different block masks for different frames (“Frame-Diff”) is also not good, which possibly shares the similar reason as the small temporal length, i.e., making masked video modeling too easy because information can be easily borrowed from adjacent/short-term frames.
Conclusion and Discussion
In the NLP field, transformers have become the de-facto standard architecture and reshaped varieties of NLP tasks in the past several years. This is largely driven by the widely used BERT pretraining strategy that demonstrates scaling abilities for pretraining large models on large-scale data. Recent success of transformers on a variety of computer vision tasks motivates a line of work to explore BERT pretraining in vision.
Unlike a few concurrent studies on images, we take a step forward and study how to explore BERT pretraining for video transformers. This may look very straightforward, but worth being extensively studied. We introduced BEVT that learns both discriminative spatial representation and temporal dynamics. We demonstrate that decoupling video pretraining into spatial and temporal representation learning is not only efficient but also effective. Empowered by the simple design of BEVT, we achieved SOTA performance on three video recognition datasets. We hope our study will inspire more research efforts in this direction.
Limitations. Although our decoupled design has significantly improved the video pretraining efficiency, BEVT still requires pretty high computational resources (e.g., more than one week on 32 V100 GPUs). This would be the biggest obstacle in scaling the pretraining on larger video datasets and larger models. The idea proposed in the concurrent work may help, i.e., ignoring the masked tokens in the transformer computation during BERT pretraining, which is left as the future work.
References
Appendix A Details on Pretrained Tokenizer
In this paper, we use the visual tokenizer of a pretrained image VQ-VAE from , which is also called discrete VAE . The tokenizer of the VQ-VAE is trained to transform each image into a image token map according to a visual codebook, while the decoder of the VQ-VAE is trained to reconstruct each input image from its tokens. The vocabulary size of the visual tokens is .