Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning
AJ Piergiovanni, Weicheng Kuo, Anelia Angelova
Introduction
Visual Transformers (ViT) have been an ubiquitous backbone for visual representation learning, leading to many advances in image understanding , multimodal tasks and self-supervised learning , etc. However, adaptations to video are both challenging and computationally intensive, so video versions have been been specially designed to handle the larger number of frames, for example, ViViT , MultiView , TimeSFormer and others .
Video understanding is an essential computer vision task, and a large number of successful video architectures have been developed . Previous video 3D CNNs were designed to handle videos by learning spatio-temporal information; they often borrow from mechanisms for learning on images, for example use pre-trained image CNN weights by inflating the kernels to 3D. However, once adapted to videos, these kernels are no longer applicable to images.
Furthermore, most previous works treat image and video as entirely different inputs, providing independent methods for either videos or images, since designing a model capable of handling both is challenging. At the same time, image and video inputs are inherently related and a single visual backbone should be able to handle either or both inputs. Previous methods for co-training image and video adapt the architectures to do so with significant portions of the network designed for each input. Works such as Perceiver and Flamingo address this by resampling the input and compressing it into a fixed number of features. However, this resampling can still be expensive for long videos, and, in the case of Flamingo, it treats videos as individual frames sampled at 1 FPS, which limits the temporal information. Such low FPS sampling and per-frame modeling would often be insufficient for datasets which rely on motion and temporal understanding, e.g., SomethingSomething , or for recognizing quick and short actions. On the other hand, using one of the above-mentioned approaches with dense frames is computationally infeasible.
To address these limitations, we propose a simple but effective model, named TubeViT, to utilize a standard ViT model seamlessly for both image and videos. We introduce Sparse Video Tubes, a lightweight approach for joint image and video learning. Our method works by sparsely sampling various sized 3D space-time tubes from the video to generate learnable tokens, which are used by the vision transformer (Figure 1). With sparse video tubes, the model is easily applicable to either input, and can better leverage either or both sources of data for training and fine-tuning. The sparse video tubes naturally handle raw video signals and image signals which is crucial to understanding actions and other spatio-temporal information in videos.
Video models are also expensive to train, and previous works have studied ways to leverage already trained models, such as using frozen ones or adapting them to videos . We expand on these ideas, and use the Sparse Video Tubes to adapt much larger ViT models to videos with lightweight training (Sec. 3.6). Thus we create powerful large video models with less resources.
We evaluate the approach across many standard video datasets: Kinetics-400, Kinetics-600, Kinetics-700, and SomethingSomething V2, outperforming the state-of-the-art (SOTA). Our methods are trained from scratch or on ImageNet-1k and Kinetics datasets and outperform even methods additionally pre-trained from very large datasets (e.g., JFT ). Our work also outperforms models targeting video pretraining, such as recent video Masked Auto-Encoder (MAE) works .
Our key findings are that by using the sparse video tubes, we are able to better share the weights learned for both images and videos. This is in contrast to prior works that either inflate kernels or add new temporal-specific layers. Further, due to the sparse sampling, the number of tokens remains low, which we also find is important, both for reducing FLOPs and improving performance.
Our contribution is construction of sparse video tubes, obtained by sparsely sampling videos with various sized 3D space-time tubes. With that we accomplish the following: (1) a universal visual backbone which easily adapts a ViT architecture to videos; (2) joint image and video understanding which seamlessly uses either input; (3) an easy-to-scale approach for video understanding, which can also leverage already trained (large) ViT models.
Related work
Video understanding is an important topic in computer vision. Early works hand-designed trajectory features to understand motion and time . With the success of neural networks, many different approaches have been developed, such as two-stream CNNs taking image frames plus optical flow for motion information as input , finding a clear benefit from adding the flow information. Works studying 3D CNNs found the learning of temporal kernels to be important , but also required much more data in order to be effective . Many of the existing video CNN approaches, have been specialized to handle videos, either with flow streams or 3D kernels and thus have not been applicable to images.
With the introduction of transformer models and self-attention , vision transformers have been very effective for image-based tasks. However, due to the quadratic cost of self-attention and the dense sampling, their use for videos has required different elements, such as space-time factorized attention . However, these video transformers have not really been tested on longer videos and are mostly evaluated on short clips. The ability to handle larger number of input frames and understand long-term actions and their relationships is of key importance, but becomes computationally prohibitive with current models.
Previous works have found that transformers focus on only a few tokens and works have been designed to pool or reorganized tokens effectively . Many video works have found that frames contain redundant information, and thus propose strategies to sample frames . Other works have studied ways to reduce the number of tokens in video transformer models . However, all these works still use an initial dense sampling of the video, then some heuristics to reduce the number of inputs. In this work, we more sparsely sample the input initially, increasing efficiency.
Other recent works have studied video MAE tasks as pretraining , they similarly treat videos as tubes, and study the sparseness in terms of the masking, having similar findings that sparseness is beneficial. However, they use a single tube shape and create non-overlapping patches and have not been studied when joint training with images.
This work is also related to approaches which use multiple views or streams from the input data, e.g., Multi-View Transformers , SlowFast Networks and others , all have found benefits from multiple input views or streams. MultiView Transformers , similarly to us, is using tubes of varying shapes. The key difference is the sparse sampling we use enables the use of a single ViT encoder model, rather than multiple smaller, per-view encoders. This further unifies the approach with images.
Another line of work in video understanding is leveraging image datasets during pre-training . This is valuable as image-only datasets are better annotated and provide richer semantic information. One approach is to bootstrap the video models from image-pretrained models, often by inflating kernels. The model is first pre-trained on image data, and then only trained on video. Other works proposed to co-train image and video jointly . These approaches adapt the architectures to handle both inputs which might be inefficient, e.g., treating an image input as a video of 1 frames or using separate networks to first encode the inputs.
In contrast to all the previous works, our method is simple and straightforward. One crucial set of differences is that the tubes are sparsely applied to the raw input, consists of different shaped, possibly overlapping tubes, and uses a single, shared backbone network, different from all previous approaches (). This leads to both more efficient and accurate models. Secondly, and more importantly, the model is entirely shared between the image and video modalities. This is an important distinction as it not only improves performance for both tasks, but is also more generally applicable to vision tasks.
Method
The standard ViT architecture takes an image and converts it into patch embedding, for example, by using a 2D convolutional kernel, with a stride. This results in a sequence of patches as the image representation, e.g., 196 for a input image. Given a video , prior approaches either used the same, dense 2D patches (e.g., TimeSFormer ) or used dense 3D kernels, e.g., 2 or as in ViViT . In both cases, this results in significantly more tokens, e.g., , where is the number of frames. These tubes or patches are then linearly projected into an embedding space, . This sequence of tokens is then processed by a transformer encoder, using standard components, MSA - the multi-head self attention and MLP - the standard transformer projection layer (LN denotes Layer Norm). For a sequence of layers , we compute the representation and next token features for all the tokens:
To reduce the computational cost, prior approaches factorize the attention mechanism, to have a spatial and temporal attention or use multiple views with smaller, view level transformers .
2 Sparse Video Tubes
We propose a simple and straightforward method which is seamlessly applicable to both images and videos. Our approach follows the standard ViT tokenization approach for images: a 2D convolution with a kernel. We build on the observation that sparseness is effective for videos. Rather than following the prior works that densely tokenize the video, we instead use the same 2D kernel, but with a large temporal stride, for example, applied to every 16th frame. Thus for an input video clip of , this results in only 392 tokens, rather than the 6k in TimeSFormer or 1-2k in ViViT.
However, this sparse spatial sampling might lose information, especially for quick or short actions. Thus, we create sparse tubes of different shapes, for example, a tube to obtain information from many frames at low spatial resolution. These tubes can have any shape, and we experimentally explore the effect of these. Importantly, these tubes also have large strides, sparsely sampling the video in different views. We also optionally add an offset to the start location, so that the patches do not always start at and this allows a reduction in the overlap between the tubes. This is illustrated in Figure 2. Tubes of various sizes are also used in the MultiView approach for video classification , however there they are densely sampled and processed by multiple transformers, resulting in a more computationally intensive approach.
Furthermore, in contrast to prior works, we also allow for overlap between the tubes. Specifically, we can represent a tube as for the kernel shape, for the spatio-temporal stride applied to the kernel, and as the offset of the starting point of the convolution.
With the proposed design, our approach enables seamless fusion of the image- and video- visual information. The sparse spatial sampling allows sharing the image and frame tokens and the sparse video tubes create a low number of video-specific tokens. This enables better sharing of the ViT model between images and videos.
3 Positional embedding for sparse video tubes
A key aspect of our approach is the implementation of the positional embedding. In language models, relative positional embeddings are a common and effective approach . However, here, the relative position between two tokens has minimal meaning, and no real reference to where the patch/tube came from in the original video or image. The ViT model and similarly TimeSFormer and ViViT used learnable positional embeddings for the patches. Here, such an approach can be hard for the model, as these learned embeddings do not necessarily reflect where the patches came from in the original video, especially in the case where patches overlap.
Instead, we use a fixed sine/cosine embedding. Importantly, we take into account the stride, kernel shape and offsets of each tube when applying the positional embeddings. This ensures that the positional embedding of each patch and tube has the global spatio-temporal location of that tube.
Specifically, we compute the embeddings as follows. Here is a constant hyperparameter (we used 10,000). For from 0 to ( is the number of features), and for from 0 to , :
This adds each spatio-temporal position embedding to the feature dimension of the token . Following previous work , this is done for different wavelengths for each channel. is used since we have 6 elements (a sine and cosine value for each ), this creates a position value for each channel of the representation.
Importantly, here represents the center of the tube, taking into account any strides and offsets used in the tube construction (the channel dimension is not shown here).
After the tokenization step, we concatenate all the tokens together and apply a standard transformer model. This simple structure lets the model share the majority of the weights between all inputs, which we find to be quite beneficial.
4 Sparse Tube Construction
We explore several methods to create the visual tubes. Our core approach consist of 2 tubes: the tube used to tokenize the image and a tube additionally used for the video. Both have strides of . This base tokenizer provides strong performance, but we explore several variations on it.
Multi-Tube. We add multiple tubes to the core approach of various sizes. For example, we can add temporally long and spatially small tubes, such as to learn long actions, or more spatially focused tubes such as a tube. There are many variations of tube shape and stride, which we experimentally explore.
Space-to-Depth Another way to extend the core approach is a method inspired by depth-to-space . Here, we reduce the number of channels in a tube, e.g., by a factor of 2. Thus the tube shape becomes . Next, we concatenate 2 tokens along the channel axis. We can then also reduce the stride of the tube. This results in the same number of tokens and dimensions as the original, but effectively increases the kernel size without changing the number of parameters. I.e., when the stride is reduced on the time axis, the token now represents locations, but only uses parameters. In the experiments, we explore different settings: e.g., more temporal dense vs more spatially dense and the depth to space factor (2, 4, 8, etc.).
Interpolated Kernels. For this setting, rather than having a unique kernel for each tube, we learn 1 3D kernel of shape . We then use tri-linear interpolation to reshape the kernel to various sizes, e.g., 4x16x16 or 32x4x4, etc. depending on the tube configuration. Any sized kernel can be created from this single kernel. This method has several advantages. (1) It reduces the number of learned parameters that are only used on the video stream. (2) It enables more flexible usage of the kernels, e.g., it can be made longer to handle longer videos, or spatially larger to find small objects.
The TubeViT approach consists of the union of the above-mentioned Multi-Tube and Space-to-Depth, the exact settings are provided in the supplemental materials. We experiment with Interpolated Kernels in ablations.
5 Image and Video Joint Training
As described above, our approach seamlessly adapts to either image, video or both inputs. While image+video joint inputs are rare, the ability to use them together while training is very important as many datasets with valuable annotations (e.g., ImageNet, Kinetics) come from either image sources or video sources but not both. Jointly training with our approach is easy – the image is tokenized by the 2D kernel and the video is tokenized by both the 2D patches (with large temporal stride) and Sparse Tubes. Both are then passed into a standard ViT; the position embedding will be supplied in either case. The position embedding approach is also needed for the joint training to be effective. We demonstrate the benefits of our approach for joint training in the experiments, Section 4.
6 Image-To-Video Scaling Up of Models
We also propose a method for a more efficient way of scaling up the models (Figure 3). Training large ViT models is computationally expensive, especially for videos. Since nearly all the components of our model are shared between the both images and videos, we explore a method to utilize large models without having heavy fine-tuning.
First, we train a smaller model jointly on images and videos. This gives us a set of weights for the tubes. Then we take a large pre-trained image ViT, but further add the tubes. These tubes use the same kernel weights as the smaller model, and so we can avoid further training them. Since larger ViTs generally use more channel dimensions than smaller ones, we use the space-to-depth transform again here to create tokens with the proper channel dimensions without needing new weights.
Next, we pick a point in the network and freeze all the layers before it, for example, the 26th of 32 layers in ViT-H. At this point, we add a gated connection to the network:
where is the layer the network is frozen at (e.g., 26) of the ViT model and is the raw input tokens from the tubes. is the learned gating parameter, initialized at 0. In the first steps of training, this gate has no effect on the representation, and thus the ViT is unchanged. However, it can learn to incorporate the raw tubes at this point and further refine the later weights.
Experiments
We evaluate the approach on several popular datasets: Kinetics 400, Kinetics 600, Kinetics 700 , and SomethingSomething V2 . These datasets cover a wide variety of video understanding challenges and are well established in the literature. The main results are trained jointly on ImageNet-1k (of 1.2M images) and the video data, please see the supplemental materials for full details. We use standard Top 1 and Top 5 evaluation metrics and report FLOPs of ours and previous works, when available. Our model sizes are 90M Base (B), 311M Large (L). A 635M Huge (H) is ‘created’ with Image-to-Video scaling.
For the main results, we use 4 tubes with the following configuration (order of ): (1) with a stride of ; (2) with a stride of and an offset of ; (3) with a stride of and an offset of ; and (4) with a stride of . For an input of , this results in only 559 tokens, significantly less than other approaches. In the supplemental material, we have detailed experiments over many tube configurations, as well as the space-to-depth settings used.
We would like to note that with data augmentation such as random spatial and temporal cropping, over multiple training epochs the model will see different parts of the video, even with sparse sampling.
Comparison to SOTA. First, we compare our final approach to previous state-of-the-art (SOTA) methods. Tables 1, 2 and 3 shows the performance of our model compared to the state-of-the-art on the Kinetics-400 Kinetics-600 and Kinetics-700 datasets. Table 1 shows additional information (e.g. views, pre-training datasets) which applies to the other tables as well. These results show our approach outperforms SOTA, both in terms of accuracy and efficiency. We also outperform methods on co-training of images and videos, and methods with strong video pre-training.
We note that all the sizes of our model perform well, despite the fact that others are much larger or use significantly larger pre-training data (e.g., CoCa with 1B params and 1.8B examples, MerlotReserve has 644M params and uses YT-1B dataset). Table 4 shows our results on the Something-Something dataset (SSv2). This dataset is often used to evaluate more dynamic activities. Our approach outperforms SOTA on it as well.
Joint image+video training. We further explore the effects of co-training on image+video datasets, finding this to be highly effective as also shown above. Table 5 evaluates this in a side-by-side experiment of using Kinetics (video) only vs Kinetics and ImageNet datasets for pre-training. We see that there is a large gain from the co-training of our approach. We see that two-stage training, i.e., first training on one dataset and then training on a second one, is also weaker than the joint training, as the two datasets cannot interact during training. We also compare to prior methods such as TimeSFormer only using dense 2D patches, or using inflated 3D kernels (e.g., ViViT ). In both cases, we see a clear benefit from the proposed approach. We also note that these prior approaches have significantly more FLOPs, due to the large number of tokens from the dense sampling. Our observations that image and video co-training is beneficial are consistent with prior works ; here the difference is that we have a single compact model to do that.
As a sanity check, we also compare our performance on ImageNet-1k, without any hyperparameter tuning or additions: our ViT-B model only trained on ImageNet has 78.1 accuracy, similar to the ViT-B in . When joint training with Kinetics-600, the model gets 81.4, a gain of 3.4%, showing the benefits of joint training for image-only tasks too. While other works achieve higher performance on ImageNet, they often use specialized data augmentation, learning schedules, and other tricks which we are not using. Instead, we are purely studying the benefit from using both videos and images.
Scaling video training with sparse video tubes. In Table 6 we demonstrate how a small TubeViT model can be adapted leveraging a large and (often independently) pre-trained model on images only. We start by leveraging a large, image-pretrained ViT, here ViT-H. We then take the learned tubes from TubeViT-B and use them along with the ViT-H image tokenizer to generate a set of tokens from a video, same as before. Then these are used as input to ViT-H, and we finetune only the latter parts of the model on the video data. These results suggests that this is an effective way to scale and utilize giant ViT models without needing the high compute cost to fully finetune the model. We also see that the gating in Eq. 8 is effective. We also found that in this setting, training time was reduced by 43%, as it has fewer weights to update.
Detrimental Effects of Too Many Tokens. Next we study the effect of number of tokens used in the model, shown in Figure 4. This result is another key insight as to why our approach works so well: with too many tokens, the performance drops, especially when only using Kinetics data. There are a number of possible reasons for why this occurs, for example, the self-attention mechanism could be struggling to learn for longer sequences, or there may not be sufficient data to learn the longer sequences, or perhaps the model is overfitting with longer sequences. This result indicates that for current datasets, the sparse sampling is an effective and efficient way to process videos. Further, it is possible that existing using long, densely sampled sequences are effected by this, and perhaps another reason the factorized attention modules are needed.
2 Ablations
In this section, we present a number of ablation studies to determine why this method is effective. For these experiments we use Kinetics 600.
Main ablations. First, we study the effect of the choice of position biases (Table 7a). We find that adding fixed cosine position embedding performs best and much better than other embeddings. Intuitively, this makes sense, since we are sparsely sampling potentially overlapping tokens, this method is able to best capture the token location.
Next in Table 7b, we study the number of tubes used. This finding, which is consistent with previous multi-view observations , shows that having a variety of tubes is beneficial to video understanding.
Next, in Table 7c, we study the depth-to-space versions of the network. Here, we reduce the channels of the generated tokens from , e.g., by a factor of 2 or 4. Then after generating the tokens, we concatenate them along the channel axis. We study both increasing the number of tokens along the spatial and temporal dimensions. We find this to be an effective method, as it enables more dense samples without increasing the number of parameters or tokens.
Table 7d compares evaluating with more patches than the model was trained with. To do this we reduce the strides of the kernel. Initially this improves results, but after increasing 2x, the performance begins to drop, likely because the evaluation data is too different from the training one.
In Table 7e, we study the ability of the interpolated single kernel. I.e., rather than having 3D convolutional kernels, one for each tube, we build 1 3D kernel and use interpolation to generate the different tube shapes. Somewhat surprisingly, we find this works fairly well, while also reducing the number of learnable parameters in the network.
In Table 7f, we compare the approach with different number of temporal and spatial crops. We find that even a single crop gives strong performance, and the standard performs nearly the same as the setting, suggesting that the sparse samples are quite suitable and further information is not as beneficial.
Factorized attention ablations. In Table 8, we further study the effect of adding a new attention layer to an ImageNet pre-trained ViT model. Here, we are using the tube method to tokenize the inputs, but instead of using a factorized attention module, we simply add an additional self-attention layer. This has a similar effect of the factorized attention approaches that add new, uninitialized projections to a pre-trained ViT (e.g., TimeSFormer and ViViT). These results indicate that such methods are not able to best utilize the image pre-trained weights of the network due to these new layers. Since the sparse tubes yield few additional tokens, they can directly use the same ViT model without factorized attention and are thus able to better utilize the image trained weights. Note that there are still differences between the works, e.g., the reduced number of tokens, etc. However, we believe this observation holds, and is a possible explanation for why the spatio-temporal attention in ViVit performed better for some datasets.
Model scaling ablations. Table 9 provides ablations on scaling to create TubeViT Base from a Tiny one. Even just training the final few layers is effective (4 of 12), and can nearly match the performance of full finetuning. This is consistent with our observations in Table 6 for ViT-H.
Figure 5 visualizes the learned 2D patches and 3D tubes.
Conclusion
We proposed sparse video tubes for video recognition. With sparse video tubes, a ViT encoder can be transformed into an efficient video model. The approach is simple, enables seamless joint training with images and videos and improves video recognition across multiple datasets. We also demonstrate an elegant scaling of video models with our proposed method. We conduct extensive ablation experiments to determine why the approach works, finding the a combination of the joint training, reduced tokens, and better utilization of shared image+video weights led to the improvements. We obtain SOTA or above performance.
Appendix A Implementation Details
Our hyperparameters are summarized in Table 10. For all datasets, we employ random spatial and temporal cropping. For most datasets, these settings were the same. For Charades, we decreased the batch size but used longer, 128 frame clips, as Charades videos are roughly 30 seconds long, compared to 10 seconds for Kinetics.
We also found some training instability when using larger ViT models. When using ViT-L or ViT-H models, we had to decrease the weight decay value as well as the learning rate, otherwise we found the training accuracy dropped to 0 and the loss stayed flat.
For smaller datasets, such as Charades and SSv2, we had to increase the data augmentation settings, as done in previous works, e.g., . We added Mixup and label smoothing and dropout to them.
For all the datasets, we applied RandAugment , as we found this to be beneficial. We also kept the number of steps the same for all datasets.
Joint ImageNet and Kinetics Training. When jointly training on the two (or more) datasets, we use a separate fully connected layer to output the class predictions. E.g., for ImageNet and Kinetics-600, we use an FC layer with 1,000 and 600 outputs. We then compute the loss for the relevant head and backpropagate it. During the joint training, we use the same settings as listed in Table 10. We use the joint training for Kinetics 400, 600 and 700. For Charades and SSv2, we use the Kinetics-600+ImageNet pretrained model and finetune it on the dataset.
Full Model Settings. Our model is based on the standard ViT models, thus the core of the approach is the same as previous ViTs . We summarize those settings in Table 11.
In Table 12, we detail the settings for each tube.
Appendix B Additional Experiments on Charades
We include results on Charades to show the effectiveness of this approach on longer videos, since Charades videos are on average 30 seconds long. However, Charades is also a multi-label dataset, and we found it required different settings to effectively train, so we include all those details here.
First, we found that the core multi-tube approach were not performing as well as some prior work (e.g., AssembleNet ). Since Charades has a lot of object-related actions and contains longer videos with more temporal information, we modified the core model to make it more suitable for this data. First, we used the interpolation method to increase the tube shapes to:
We note two important factors. First, since we use interpolation to create the larger kernels, the number of learned parameters is the same, and initialized from the same kernels for the other datasets. Second, since the number of strides is unchanged, this results is the same number of tokens. Critically, this change has very little effect on the network and its parameters, but enables the model to better capture the information for Charades.
In Table 13, we report the results. The core MultiTube approach performs quite well, but with the interpolated kernels, is able to perform on par with TokenLearner , the state-of-the-art, while still sparsely sampling the video. We also perform similarly using significantly less data, e.g., JFT-300M was used to pre-trained TokenLearner, we accomplish the same performance without such large scale data.
Appendix C Ablations on Tube Shapes.
In Table 14, we provide a detailed study on Kinetics-600 of various tube configurations. We observe that the model isn’t overly sensitive to tube shapes, at least on Kinetics-600, but having multiple, different tubes, as well as variation in their shapes is generally beneficial. We use the following tubes in these experiments: