SVFormer: Semi-supervised Video Transformer for Action Recognition
Zhen Xing, Qi Dai, Han Hu, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang
Introduction
Videos have gradually replaced images and texts on Internet and grown at an exponential rate. On video websites such as YouTube, millions of new videos are uploaded every day. Supervised video understanding works have achieved great successes. They rely on large-scale manual annotations, yet labeling so many videos is time-consuming and labor-intensive. How to make use of unlabeled videos that are readily available for better video understanding is of great importance .
In this spirit, semi-supervised action recognition explores how to enhance the performance of deep learning models using large-scale unlabeled data. This is generally done with labeled data to pretrain the networks , and then leveraging the pretrained models to generate pseudo labels for unlabeled data, a process known as pseudo labeling. The obtained pseudo labels are further used to refine the pretrained models. In order to improve the quality of pseudo labeling, previous methods use additional modalities such as optical flow and temporal gradient , or introduce auxiliary networks to supervise unlabeled data. Though these methods present promising results, they typically require additional training or inference cost, preventing them from scaling up.
Recently, video transformers have shown strong results compared to CNNs . Though great success has been achieved, the exploration of transformers on semi-supervised video tasks remains unexplored. While it sounds appealing to extend vision transformers directly to SSL, a previous study shows that transformers perform significantly worse compared to CNNs in the low-data regime due to the lack of inductive bias . As a result, directly applying SSL methods, e.g., FixMatch , to ViT leads to an inferior performance .
Surprisingly, in the video domain, we observe that TimeSformer, a popular video Transformer , initialized with weights from ImageNet , demonstrates promising results even when annotations are limited . This encourages us to explore the great potential of transformers for action recognition in the SSL setting.
Existing SSL methods generally use image augmentations (e.g., Mixup and CutMix ) to speed up convergence under limited label resources. However, such pixel-level mixing strategies are not perfectly suitable for transformer architectures, which operate on tokens produced by patch splitting layers. In addition, strategies like Mixup and CutMix are particularly designed for image tasks, which fail to consider the temporal nature of video data. Therefore, as will be shown empirically, directly using Mixup or CutMix for semi-supervised action recognition leads to unsatisfactory performance.
In this work, we propose SVFormer, a transformer-based semi-supervised action recognition method. Concretely, SVFormer adopts a consistency loss that builds two differently augmented views and demands consistent predictions between them. Most importantly, we propose Tube TokenMix (TTMix), an augmentation method that is naturally suitable for video Transformer. Unlike Mixup and CutMix, Tube TokenMix combines features at the token-level after tokenization via a mask, where the mask has consistent masked tokens over the temporal axis. Such a design could better model the temporal correlations between tokens.
Temporal augmentations in literatures (e.g. varying frame rates) only consider simple temporal scaling or shifting, neglecting the complex temporal changes of each part in human action. To help the model learn strong temporal dynamics, we further introduce the Temporal Warping Augmentation (TWAug), which arbitrarily changes the temporal length of each frame in the clip. TWAug can cover the complex temporal variation in videos and is complementary to spatial augmentations . When combining TWAug with TTMix, significant improvements are achieved.
As shown in Fig. 1, SVFormer achieves promising results in several benchmarks. (i) We observe that the supervised Transformer baseline is much better than the Conv-based method , and is even comparable with the 3D-ResNet state-of-the-art method on Kinetics400 when trained with 1% of labels. (ii) SVFormer-S significantly outperforms previous state-of-the-arts with similar parameters and inference cost, measured by FLOPs. (iii) Our method is also effective for the larger SVFormer-B model. Our contributions are as follows:
We are the first to explore the transformer model for semi-supervised video recognition. Unlike SSL for image recognition with transformers, we find that using parameters pretrained on ImageNet is of great importance to ensure decent results for action recognition in the low-data regime.
We propose a token-level augmentation Tube TokenMix, which is more suitable for video Transformer than pixel-level mixing strategies. Coupled with Temporal Warping Augmentation, which improves temporal variations between frames, TTMix achieves significant boost compared with image augmentation.
We conduct extensive experiments on three benchmark datasets. The performances of our method in two different sizes (i.e., SVFormer-B and SVFormer-S) outperform state-of-the-art approaches by clear margins. Our method sets a strong baseline for future transformer-based works.
Related Works
Deep Semi-supervised Learning Deep learning relies on large-scale annotated data, however collecting these annotations is labor-intensive. Semi-supervised learning is a natural solution to reduce the cost of labeling, which leverages a few labeled samples and a large amount of unlabeled samples to train the model. The research and application of SSL mainly focus on image recognition with a two-step process: data augmentation and consistency regularization. Concretely, different data augmentations views are input to the model, and their output consistencies are enforced through a consistency loss. Another line of work generates new data and labels using mixing to train the network. Among these state-of-the-art methods, FixMatch have been widely used due its effective and its variants have been extended to many other applications, such as object detection , semantic segmentation , 3D reconstruction , etc. Although FixMatch has achieved good performance in many tasks, it may not achieve satisfactory results when directly transferred to video action recognition due to the lack of temporal augmentation. In this paper, we introduce temporal augmentation TWAug with mixing based method TTMix, which is suitable for video transformers at SSL settings.
Semi-supervised Action Recognition VideoSSL presents a comparative study of applying 2D SSL methods to videos, which verifies the limitations of the direct extension of pseudo labeling method. TCL explores the effect of a group contrastive loss and self-supervised tasks. MvPL and LTG introduce optical flow or temporal gradient modal to generate high quality pseudo labels for training, respectively. CMPL introduce an auxiliary network, which requires more frames in training, increasing the difficulty of application. Besides, previous methods are all based on 2D or 3D convolutional networks, which require more training epochs. Our approach is the first to make the exploration of Video Transformer for SSL action recognition and achieves the best performance with the least training cost.
Video Transformer The great success of vision transformer in image recognition leads to the development of exploring the transformer-base architecture for video recognition tasks. VTN uses additional temporal attention on the top of the pretrained ViT . TimeSformer investigates different spatial-temporal attention mechanisms and adopts factored space time attention as a trade-off of speed and accuracy. ViViT explores four different types of attention, and selects the global spatio-temporal attention as the default to achieve promising performance. In addition, MviT , Video Swin , Uniformer and Video Mobile-Former incorporate the inductive bias in convolution into transformers. While these methods focus on fully-supervised setting, limited effort has been made for transformers in the semi-supervised setting.
Data Augmentation Data augmentation is an essential step in modern deep networks to improve the training efficiency and performance. Cutout removes random rectangle regions in images. Mixup performs image mixing by linearly interpolating both the raw image and labels. In CutMix , patches are cut and pasted among image pairs. AutoAugment automatically searches for augmentation strategies to improve the results. PixMix explores the natural structural complexity of images when performing mixing. TokenMix mixes images at the token-level and allows the region to be multiple isolated parts. Though these methods have achieved good results, they are all specially designed for pure image and most of them are pixel-level augmentations. In contrast, our TTMix coupled with TWAug is devised for video.
Method
In this section, we first introduce the preliminaries of SSL in Sec. 3.1. The pipeline of our proposed SVFormer is described in Sec. 3.2. Then we detail the proposed Tube TokenMix (TTMix) in Sec. 3.3, as well as the effective albeit simple temporal warping augmentation. Finally, we show the training paradigm in Sec. 3.4.
Suppose we have training video samples, including labeled videos and unlabeled videos , where is the labeled video sample with a category label , and is the unlabeled video sample. In general, . The aim of SSL is to utilize both and to train the model.
2 Pipeline
SVFormer follows the popular semi-supervised learning framework FixMatch that use a consistency loss between two differently augmented views. The training paradigm is divided into two parts. For the labeled set , the model optimizes the supervised loss :
where refers to the predictions produced by the model and is the standard cross entropy loss.
For unlabeled samples , we first use weak augmentations (e.g., random horizontal flipping, random scaling, and random cropping) and strong augmentations (e.g., AutoAugment or Dropout ) to generate two views separately, , . Then the pseudo label of the weak view , which is produced by the model, is utilized to supervise the strong view, with the following unsupervised loss:
In FixMatch, the two augmented inputs share the same model, which tends to cause model collapsing easily . Therefore, we adopt the exponential moving average (EMA)-Teacher in our framework, which is an improved version of FixMatch. The pseudo labels are generated by the EMA-Teacher model, whose parameters are updated by exponential moving average of the student parameters, formulated as:
where is a momentum coefficient, and are the parameters of teacher and student model, respectively. EMA has achieved success in many tasks, such as self-supervised learning , SSL of image classification , and object detection . Here we are the first to adopt this method in semi-supervised video action recognition.
3 Tube TokenMix
One of the core problems in semi-supervised frameworks is how to enrich the dataset with high-quality pseudo labels. Mixup is a widely adopted data augmentation strategy, which performs convex combination between pairs of samples and labels as follows:
where the ratio is a scalar that conforms to the beta distribution. Mixup and its variants (e.g. CutMix ) have achieved success in many tasks in the low-data regime, such as long-tail classification , domain adaptation , few-shot learning , etc. For SSL, Mixup also performs well by mixing the pseudo labels of unlabeled samples in image classification .
While directly applying Mixup or CutMix to video scenarios results in clear improvements in Conv-based methods , these methods show unsatisfactory performance in our method. The reason is that our method adopts the Transformer backbone, where the pixel-level mixing augmentation (Mixup or CutMix) may be not perfectly suitable for such token-level models . To narrow the gap, we propose 3 token-level mixing augmentation methods for video data, namely, Rand TokenMix, Frame TokenMix, and Tube TokenMix.
where is element-wise multiplication, and 1 is a binary mask with all ones.
The mask M differs in the three augmentation methods, as demonstrated in Fig. 3. For Rand TokenMix, the masked tokens are randomly selected from the whole video clip (from tokens). For Frame TokenMix, we randomly select frames from the frames and mask all the tokens in these frames. For Tube TokenMix, we adopt the tube-style masking strategy, that is, different frames share the same spatial mask matrix. In this case, the mask M has consistent masked tokens over the temporal axis. While our mask design shares similarity with the recent masked image/video modeling , our motivation is totally different. They focus on removing certain regions and making the model predict the masked areas for feature learning. In contrast, we leverage the mask to mix two clips and synthesize a new data sample.
The mixed sample is then fed to the student model , obtaining the model prediction . In addition, the pseudo labels for are produced by inputting the weak augmented samples to the teacher model :
Note that if , the pseudo label remains the soft label . The pseudo label for is generated by mixing and with mask ratio :
Finally, the student model is optimized by the following consistency loss:
where is the number of mixed samples. The algorithm of consistency loss for TTMix is shown in Algorithm 1.
Temporal Warping Augmentation
Most existing augmentation methods are designed for image tasks, which focus more on the spatial augmentation. They manipulate single or a pair of images to generate new image samples, without considering any temporal changes. Even the commonly adopted temporal augmentations, including varying temporal locations and frame rates , only consider simple temporal shift or scaling, that is, changing the holistic location or play speed. However, human actions are very complex and can have different temporal variation at every timestamp. To cover such challenging cases, we propose to distort the temporal duration of each frame, thus introducing higher randomness into the data.
Our Temporal Warping Augmentation (TWAug) can stretch one frame to various temporal length. Given an extracted video clip of frames (e.g., 8 frames), we randomly determine to keep all the frames, or select a small portion of frames (e.g., 2 or 4 frames) while masking the others. The masked frames are then padded with random neighbouring visible (unmasked) frames. Note that after temporal padding, the frame order is still retained. Fig. 5 shows three examples of selecting 2, 4, and 8 frames, respectively. The proposed TWAug can help the model learn the flexible temporal dynamics during training.
The Temporal Warping Augmentation serves as a strong augmentation in TTMix. Typically, we combine TWAug with the conventional spatial augmentation to perform the mixing. As shown in Fig. 4, the two input clips are first transformed by spatial augmentation and TWAug separately, after which the two clips are mixed through TTMix. We verify the effectiveness of our TWAug in Sec. 4.
4 Training Paradigm
The training of SVFormer consists of three parts: supervised loss formulated by Eq. (1), unsupervised pseudo-label consistency loss Eq. (2), and TTMix consistency loss Eq. (10). The final loss function is as follows:
where and are the hyperparameters for balancing the loss items.
Experiment
In this section, we first introduce the experimental settings in Sec. 4.1. Following previous work , we conduct experiments under different labeling rates in Sec 4.2. In addition, we also perform ablation experiments and empirical analysis in Section 4.3. If not emphasized, we only use RGB modal for inference with the official validation set.
Datasets Kinetics-400 is a large-scale human action video dataset, with up to 245k training samples and 20k validation samples, covering 400 different categories. We follow the state-of-the-art methods MvPL and CMPL to sample 6 or 60 labeled training videos per category, i.e. at 1% or 10% labeling rates. UCF-101 is a dataset with 13,320 video samples, which consists of 101 categories. We also sample 1 or 10 samples in each category as the labeled set following CMPL . As for HMDB-51 , it is a small-scale dataset with only 51 categories composed of 6,766 videos. Following the division of LTG and VideoSSL , we conduct experiments at three different labeling rates: 40%, 50%, and 60%.
Evaluation Metric We show the accuracy of Top-1 in main results, and also present the accuracy of Top-5 in some ablation experiments.
Baseline We utilize the ViT extended video TimeSformer as the backbone of our baseline. The hyperparameters are mostly kept the same as the baseline, and we adopt the divided space-time attention as in TimeSformer . Since TimeSformer only have ViT-Base models, we implement SVFormer-Small model from DeiT-S with the dimension of 384 and 6 heads, in order to have comparable number of parameters with other Conv-based methods . For fair comparisons, we train 30 epochs for TimeSformer as the supervised baseline.
Training and Inference For training, we follow the setting of TimeSformer . The training uses 8 or 16 GPUs, with a SGD optimizer using a momentum of 0.9 and a weight decay of 0.001. For each setting, the basic learning rate is set to 0.005, which is divided by 10 at epochs 25, and 28. As for the confidence score threshold, we search for the optimal from {0.3, 0.5, 0.7, 0.9}. and are set to 2. The masking ratio is sampled from beta distribution Beta(), where . In the testing phase, following the inference strategies in MvPL and CMPL , we uniformly sample five clips from the entire video, and make three different crops to get resolution to cover most of the spatial areas of the clips. The final prediction is the average of the softmax probabilities of these predictions. We also conduct a comparison of inference setting in the ablation study in Sec. 4.
2 Main Results
The main results of Kinetics-400 and UCF-101 are shown in Table 1. Compared with previous methods, our model SVFormer-S achieves the best performance with the fewest training epochs among the methods that only use RGB data. In particular, at the labeling rate of 1% setting, SVFormer-S improves previous approach by 6.3% in UCF-101 and 15.0% in Kinetics-400. In addition, when adopting larger models, SVFormer-B significantly outperforms the state-of-the-art methods.
Specifically, in Kinetics-400, SVFormer-B can achieve 69.4% with only 10% labeled data, which is comparable to 77.9% of fully-supervised setting in TimeSformer . Moreover, as shown in Table 2, for the small-scale dataset HMDB-51 , our SVFormer-S and SVFormer-B have also improved by about 10% and 15% compared with the previous method .
3 Ablation Studies
To understand the effect of each part of the design in our method, we conduct extensive ablation studies on the Kinetics-400 and UCF-101 at the 1% labeling ratio setting with SVFormer-S.
Analysis of SSL framework The comparison of FixMatch and EMA-Teacher is shown in Table 3. It is clear that the two methods have significantly improved the baseline approach. In addition, EMA-Teacher has exhibited considerable gains over FixMatch in both datasets with very few labeled samples, probably because it has improved the stability of training. FixMatch may lead to model collapse with limited labels as shown in .
Analysis of different mixing strategies We now compare the Tube TokenMix strategy with three pixel-level mixing methods, CutMix , Mixup , PixMix , as well as the other two token-level mixing methods, i.e., Frame TokenMix and Rand TokenMix. The examples of different mixing methods are shown in Fig. 6. The quantitative results are shown in Table 4. Compared with these alternative methods, all mixing methods can improve the performance, which proves the effectiveness of mixing-based consistency losses. In addition, we observe that the token-level methods (Rand TokenMix and Tube TokenMix) perform better than the pixel-level mixing methods. This is not surprising since transformers operate on tokens, and thus token-level mixing has inherent advantages.
The performance of Frame TokenMix is even worse than that of pixel-level mixing methods, which is also expected. We hypothesize that replacing entire frames in video clip will scramble up the temporal reasoning, thus leading to poor temporal attention modeling. In addition, Tube TokenMix achieves the best results. We suppose the consistent masked tokens over temporal axis can prevent information leaky between adjacent frames in the same spatial locations, especially in such a short-term clip. Therefore, Tube TokenMix could better model the spatio-temporal correlations.
Analysis of Augmentations The effects of spatial augmentation and temporal warping augmentation are evaluated in Table 5. The baseline indicates removing both strong spatial augmentation (e.g. AutoAugment and Dropout ) and TWAug in all branches. In this case, the experimental performance drops dramatically. When spatial augmentation or temporal warping augmentation is incorporated into the baseline separately, the performance is improved. The best practice is to perform data augmentations in both spatial and temporal.
Analysis of Inference We evaluate the effect of frame rate sampling and different inference schemes, as shown in Table 6. Previous methods, i.e. CMPL and MvPL , utilize the clip-based sparse sampling, which samples 8 frames at frame rate 8 as a clip. For each video, 10 clips are sampled, where each clip is cropped 3 times according to different spatial positions. Finally, the predictions of samples are averaged. TimeSformer adopts the video-based sparse sampling, that is, 8 frames are sampled at frame rate 32 as the representation of the whole video. Then 3 different cropped views are used, namely . Applying the video-based sparse sampling scheme as in TimeSformer can reduce the inference cost, but the performance is worse than that of clip-based sampling. Experiments at sampling schemes of and demonstrate better performance. For the trade-off between efficiency and accuracy, we use as the default setting following CMPL and MvPL .
Analysis of hyperparameters Here we explore the effect of different hyperparameters. We conduct experiments under 1% setting of Kinetics-400. We first explore the effect of different threshold values of . As shown in Fig. 7(a), we can observe that when labeled samples are extremely scarce, best results are achieved by setting a small value (). We then evaluate how the ratio between labeled samples and unlabeled samples in a mini-batch affect the result. We fix the labeled sample number to 1, and sample unlabeled samples to form a mini-batch, where is in {1, 2, 3, 5, 7}. The results are shown in Fig. 7(b). When , the model produces the highest result. Finally, we explore the choice of momentum coefficient of EMA and the loss weights and , as shown in Fig. 7(c) and Fig. 7(d). We thus set and as default setting in all the experiments.
Conclusion
In this paper we present SVFormer, a transformer-based semi-supervised video action recognition method. We propose Tube TokenMix, a data augmentation method that is specially designed for video transformer models. Coupled with the temporal warping augmentation, which covers the complex temporal variations by arbitrarily changing the frame length, TTMix achieves significant improvement compared with conventional augmentations. SVFormer outperforms the state-of-the-art with a large margin on UCF-101, HMDB-51 and Kinetics-400 without increasing overheads. Our work establishes a new benchmark for semi-supervised action recognition and encourages future work to adopt Transformer architecture.
Acknowledgement This project was supported by NSFC under Grant No. 62032006 and No. 62102092.