Multiview Transformers for Video Recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, Cordelia Schmid
Introduction
Vision architectures based on convolutional neural networks (CNNs), and now more recently transformers, have made great advances in numerous computer vision tasks. A central idea, that has remained constant across classical methods based on handcrafted features to CNNs and now transformers , has been to analyze input signals at multiple resolutions.
In the image domain, multiscale processing is typically performed with pyramids as the statistics of natural images are isotropic (all orientations are equally likely) and shift invariant . To model multiscale temporal information in videos, previous approaches such as SlowFast have processed videos with two streams, using a “Fast” stream operating at high frame rates and a “Slow” stream at low frame rates, or employed graph neural networks to model long-range interactions .
When creating a pyramidal structure, spatio-temporal information is partially lost due to its pooling or subsampling operations. For example, when constructing the “Slow” stream, SlowFast subsamples frames, losing temporal information. In this work, we propose a simple transformer-based model without relying on pyramidal structures or subsampling the inputs to capture multi-resolution temporal context. We do so by leveraging multiple input representations, or “views” of the input video. As shown in Fig. 1, we extract tokens from the input video over multiple temporal durations. Intuitively, tokens extracted from long time intervals capture the gist of the scene (such as the background where the activity is taking place), whilst tokens extracted from short segments can capture fine-grained details (such as the gestures performed by a person).
We propose a multiview transformer (Fig. 1) to process these tokens, and it consists of separate transformer encoders specialized for each “view”, with lateral connections between them to fuse information from different views to each other. We can use transformer encoders of varying sizes to process each view, and find that it is better (in terms of accuracy/computation trade-offs) to use a smaller encoder (e.g. smaller hidden sizes and fewer layers) to represent the broader view of the video (Fig. 1 left) while an encoder with larger capacity is used to capture the details (Fig. 1 right). This design therefore poses a clear contrast to pyramid-based approaches where model complexity increases as the spatio-temporal resolution decreases. Our design is verified by our experiments which show clear advantages over the former approach.
Our proposed method, of processing different “views” of the input video is simple, and in contrast to previous work generalizes readily to a variable number of views. This is significant, as our experiments show that accuracy increases as the number of views grows. Although our proposed architecture increases the number of tokens processed by the network according to the number of input views, we show that we can consistently achieve superior accuracy/computation trade-offs compared to the current state of the art , across a spectrum of model sizes, ranging from “Small” to “Huge”. We show empirically that this is because processing more views in parallel enables us to achieve larger accuracy improvements than increasing the depth of the transformer network. We perform thorough ablation studies of our design choices, and achieve state-of-the-art results on six standard video classification datasets. Moreover, we show that these results can be further improved with large-scale pretraining.
Related Work
Early works relied on hand-crafted features to encode motion and appearance information. With the emergence of large labelled datasets like ImageNet , Convolutional Neural Networks (CNNs) showed their superiority over the classic methods. Since AlexNet won the ImageNet challenge by a large margin, CNNs have been quickly adopted to various vision tasks, their architectures have been refined over many generations and later improved by Neural Architecture Search (NAS) . At the same time, CNNs and RNNs have quickly become the de-facto backbones for video understanding tasks . Since the release of the Kinetics dataset , 3D CNNs have gained popularity, and many variants have been developed to improve the speed and accuracy. Convolution operations can only process one local neighborhood at a time, and consequently, transformer blocks have been inserted into CNNs as additional layers to improve modeling of long range interactions among spatio-temporal features . Although achieving great success in natural language , pure transformer architectures had not gained the same popularity in computer vision until Vision Transformers (ViT) . Inspired by ViT, ViViT and Timesformer were the first two works that successfully adopted a pure transformer architecture for video classification, advancing the state of the art previously set by 3D CNNs.
“Pyramid” structures are one of the most popular multiscale representations for images and have been key in the early computer vision works, where their use has been widespread in multiple domains including feature descriptors , feature tracking , image compression , etc. This idea has also been successfully adopted for modern CNNs where the spatial dimension of the network is gradually reduced while the network “depth” is gradually increased to encode more semantically rich features. Also, this technique has been used to produce higher resolution output features for downstream tasks. Multiscale processing is necessary for CNNs because a convolution operation only operates on a sub-region of the input and a hierarchical structure is required to capture the whole view of the image or video. In theory, such a hierarchy is not required for transformers as each token “attends” to all other positions. In practice, due to the limited amount of training data, applying similar multiscale processing in transformers to reduce complexity of the model has proven to be effective.
Our model does not follow the pyramid structure but directly takes different views of the video and feeds them into cross-view encoders. As our experiments validate, this alternative multiview architecture has consistently outperformed its single-view counterpart in terms of accuracy/FLOP trade-offs. This is because processing more views in parallel gives us larger accuracy improvements than increasing the depth of the transformer network. Significantly, such improvement persists as we scale the model capacity to over a billion parameters (e.g., our “Huge” model), which has not been shown by the previous pyramid-structured transformers . Conceptually, our method is most comparable to SlowFast where a two-stream CNN is used to process two views of the same video clip (densely sampled and sparsely sampled frames). Instead of sampling the input video at different frame rates, we obtain different view by linearly projecting spatio-temporal “tubelets” of varying sizes for each view. Furthermore, we empirically show that our proposed method outperforms when using transformer backbones.
Multiview Transformers for Video
We begin with an overview of vision transformer, ViT , and its extension to video, ViViT , which our model is based on, in Sec. 3.1. As shown in Fig. 1, our model constructs different “views” of the input video by extracting tokens from spatio-temporal tubelets of varying dimensions (Sec. 3.2). These tokens are then processed by a multiview transformer, which incorporates lateral connections to efficiently fuse together information from multiple scales (Sec. 3.3).
Note that the linear projection can also be seen as a 3D convolution with a kernel of size and stride of in the time, height and width dimensions respectively.
where MSA denotes multi-head self-attention , LN is layer normalization and MLP consists of two linear projections separated by GeLU non-linearity.
2 Multiview tokenization
3 Multiview transformer
After extracting tokens from multiple views, we have from the input, which are processed with a multiview transformer as shown in Fig. 1. As self-attention has quadratic complexity , processing tokens from all views jointly is not computationally feasible for video. As a result, we first use a multiview encoder, comprising of separate transformer encoders (consisting of transformer layers) for the tokens between views, with lateral connections between these encoders to fuse information from each view (Fig. 2). Finally, we extract a token representation from each view, and process these jointly with a final global encoder to produce the final classification token, which we linearly read-off to obtain the final classification.
Our multiview encoder consists of separate transformer encoders for each view which are connected by lateral connections to fuse cross-view information. Each transformer layer within the encoders follows the same design as the original transformer of Vaswani et al. , except for the fact that we optionally fuse information from other streams within the layer as described in Sec. 3.3.2. Note that our model is agnostic to the exact type of transformer layer used. Furthermore, within each transformer layer, we compute self-attention only among tokens extracted from the same temporal index, following the Factorised Encoder of . This significantly reduces the computational cost of the model. Furthermore, self-attention along all spatio-temporal tokens is unnecessary, as we fuse information from other views within the multiview encoder, and also because of the subsequent global encoder which aggregates tokens from all streams.
3.2 Cross-view fusion
We consider the following three cross-view fusion methods. Note that the hidden dimensions of the tokens, , can vary between views.
A straight-forward method of combining information between different views is to perform self-attention jointly on all tokens where is the number of tokens in the view. However, due to the quadratic complexity of self-attention, this is prohibitive computationally for video models, and hence we perform a more efficient alternative.
We sequentially fuse information between all pairs of two adjacent views, and , where the views are ordered in terms of increasing numbers of tokens (i.e. ). Concretely, to update the tokens from the larger view, , we compute attention where the queries are , and the keys and values are (the tokens from the smaller view). As the hidden dimensions of the tokens between the two views can be different, we first project the keys and values to the same dimension, as denoted by
Note that , and are the query-, key- and value-projection matrices used in the attention operation . As shown in Fig. 2(a), we also include a residual connection around the cross-view attention operation, and zero-initialize the parameters of this operation, as this helps when using image-pretrained models as is common practice . Similar studies on cross stream attention have been done by for images.
An efficient method of transferring information between tokens from two views, and , is by an intermediate set of bottleneck tokens. Once again, we sequentially fuse information between all pairs of two adjacent views, and , where the views are ordered in terms of increasing numbers of tokens.
As with cross-view attention, we sequentially perform fusion between all pairs of adjacent views, beginning from the view with the largest number of tokens, and proceeding in order of decreasing token numbers. Intuitively, this allows the view with the fewest tokens to aggregate fine-grained information from all subsequent views.
Note that the only parameters introduced into the model from this fusion method are the linear projections of bottleneck tokens from one view to the next, and the bottleneck tokens themselves which are learned from random initialization. We also note that “bottleneck” tokens have also been used by .
Recall that each transformer encoder layer consists of a multi-head self attention operation (Eq. 2), followed by an MLP block (Eq. 3). A simple method is to fuse before the MLP block within each encoder layer.
Concretely, as shown in Fig. 2(c), tokens from view , with hidden dimension are concatenated with tokens from view along the hidden dimension. These tokens are then fed into the MLP block of layer and linearly projected to the depth . This process is repeated between adjacent views of the network, where once again, views are ordered by increasing number of tokens per view.
We note that it is not necessary to perform cross-view fusion at each layer of the cross-view encoder to transfer information among the different views, since each fusion operation has a global “receptive field” that considers all the tokens from the previous views. Furthermore, it is also possible for the encoders for each individual view to have different depths, meaning that fusion can occur between layer of view and layer of view where . Therefore, we consider the fusion locations as a design choice which we perform ablation studies on.
3.3 Global encoder
Finally, we aggregate the tokens from each of the views with the final global encoder, as shown in Fig. 1, effectively fusing information from all views after the cross-view transformer. We extract the classification token from each view, , and process them further with another transformer encoder, following Vaswani et al. , that aggregates information from all views. The resulting classification token is then mapped to one of classification outputs, where is the number of classes.
Experiments
For the backbone of each view, we consider five ViT variants, “Tiny”, “Small”, “Base”, “Large”, and “Huge”. Their settings strictly follow the ones defined in BERT and ViT , i.e. number of transformer layers, number of attention heads, hidden dimensions. See Appendix A.4 for the detailed settings. For convenience, each model variant is denoted with the following abbreviations indicating the backbone size and tubelet length. For example, B/2+S/4+Ti/8 denotes a three-view model, where a “Base”, “Small”, and “Tiny” encoders are used to processes tokens from the views with tubelets of sizes , , and , respectively. Note that we omit 16 in our model abbreviations because all our models use as the spatial tubelet size except for the “Huge” model, which uses , following ViT . All model variants use the same global encoder which follows the “Base” architecture, except that the number of heads is set to 8 instead of 12. The reason is that the hidden dimension of the tokens should be divisible by the number of heads for multi-head attention, and the number of hidden dimensions across all standard transformer architectures (from “Tiny” to “Huge” ) is divisible by 8.
We follow the training settings of ViViT reported in the paper and public code , unless otherwise stated. Namely, all models are trained on 32 frames with a temporal stride of 2. We train our model using synchronous SGD with momentum of 0.9 following a cosine learning rate schedule with a linear warm up. The input frame resolution is set to be in both training and inference. We follow and apply the same data augmentation and regularization schemes , which were used by to train vision transformers more effectively. During inference, we adopt the standard evaluation protocol by averaging over multiple spatial and temporal crops. The number of crops is given in the results tables. For reproducibility, we include exhaustive details in Appendix A.3.
Following previous works , we initialize our model from a corresponding ViT model pretrained on large-scale image datasets obtained from the public code of . The initial tubelet embedding operator, , and positional embeddings, , have different shapes in the pretrained model and we use the same technique as to adapt them to initialize each view of our multiview encoder (Sec. 3.3.1). The final global encoder (Sec. 3.3.3) is randomly initialized.
We report the performance of our proposed models on a diverse set of video classification datasets:
Kinetics is a collection of large-scale, high-quality datasets of 10s video clips focusing on human actions. We report results on Kinetics 400, 600, and 700, with 400, 600, and 700 classes, respectively.
Moments in Time is a collection of 800,000 labeled 3 second videos, involving people, animals, objects or natural phenomena, that capture the gist of a dynamic scene.
Epic-Kitchens-100 consists of 90,000 egocentric videos, totaling 100 hours, recorded in kitchens. Each video is labeled with a “noun” and a “verb” and therefore we predict both categories using a single network with two “heads”. Three accuracy scores (“noun”, “verb”, and “action”) are commonly reported for this dataset with action accuracy being the primary metric. The “action” label is formed by selecting the top-scoring noun and verb pair.
Something-Something V2 consists of more than 220,000 short video clips that show humans interacting with everyday objects. Similar objects and backgrounds appear in videos across different classes. Therefore, in contrast to other datasets, this one challenges a model’s capability to distinguish classes from motion cues.
2 Ablation study
We conduct ablation studies on the Kinetics 400 dataset. In all cases, the largest backbone in the multiview encoder is “Base” for faster experimentation. We report accuracies when averaging predictions across multiple spatio-temporal crops, as standard practice . In particular, we use crops, that is 4 temporal crops, with 3 spatial crops for each temporal crop. We used a learning rate of 0.1 for all experiments for 30 epochs, and used no additional regularization as done by .
Recall that a view is a video representation in terms of tubelets, and that a larger view equates to larger tubelets (and hence fewer transformer tokens) and smaller views correspond to smaller tubelets (and thus more tokens).
We considered two model-view assignment strategies: larger models for larger views (e.g., B/8+Ti/2, the larger “Base” model is used to encode tubelets and the smaller “Tiny” model encodes tubelets) and smaller models for larger views (e.g., B/2+Ti/8). Table 1(b) shows that assigning a larger model to smaller views is superior. For example, B/2+S/4+Ti/8 scores 81.8% while B/8+S/4+Ti/2 only scores 78.5%. One may argue that this is due to the increase of the FLOPs but B/4+S/8+Ti/16 still outperforms B/8+S/4+Ti/2 by a large margin under similar FLOPs. Our explanation is that larger views capture the gist of the scene, which requires less complexity to learn while the details of the scene are encapsulated by smaller views so a larger-capacity model is needed.
Another strategy is to assign the same model to all views. Table 1(b) shows that in all three examples there is little difference between assigning a “Base” model and assigning a “Small” or “Tiny” model to larger views. This result is surprising yet beneficial since we can reduce the complexity of the model at almost no cost of accuracy.
Table 1(d) shows the comparison of different fusion methods on a three-view model. We use one late fusion and an ensemble approach as the baselines. “Ensemble” simply sums the probabilities produced from each view, where the models from each view are trained separately. We also tried summing up the logits and majority voting but both obtained worse results. This method actually decreases the performance compared to the B/4 model since “Small” and “Tiny” models perform not comparably well. “Late fusion” concatenates the final embeddings produced by the transformer encoder from each view without any cross-view operations before feeding it into the global encoder. It improves the B/4 model from 78.3% to 80.6%. All of our fusion methods except MLP outperform the baselines while CVA is the best overall. Based on this observation, we choose CVA as the fusion method for all subsequent experiments. MLP fusion is the worst performing method of the three and we think it is because concatenation in the MLP blocks introduces additional channels that have to be randomly initialized, making model optimization more difficult.
Table 1(f) shows performance on Kinetics-400 as we increase the number of views. With two views we achieve a +2.5% in Top-1 accuracy over the baseline B/4 model. As we increase to three views, the improvement widens to 2.8%. Furthermore, we show that such improvement is non-trivial. For example, we also train a 14-layer and a 17-layer variants of the “Base” model. They share similar FLOPs with our two-view and three-view counterparts but their performance remains similar to that of the baseline.
Motivated by Tab. 1(d), we fix the fusion method to CVA, and vary the locations and number of layers where we apply CVA, when using a three-view B+S+Ti model (each encoder thus has 12 layers) in Tab. 1(f). The choices are in the early-, mid-, and late-stages of the transformer encoders and the number of fusion layers is set to be one and two. When using one fusion layer, the best location for fusion is mid followed by late, then early. Adding more fusion layers in the same stage does not improve the performance but combining mid and late fusion improves the performance. For example, fusion at 5th and 11th layers achieve the best result. Based on this observation, we set the fusion layers to be {11, 23} for L+B+S+Ti and {11, 23, 31} for H+B+S+Ti model variants, respectively, in subsequent experiments.
SlowFast proposes a two-stream CNN architecture that takes frames sampled at two different frame rates. The “Slow” pathway, built with a larger encoder, processes the low frame rate stream to capture the semantics of the scene while the “Fast” pathway that takes in high frame rate inputs is used to capture motion information. To make a fair comparison, we implement in the context of transformers where we use “Base” and “Tiny” models as the encoders for the Slow and Fast paths respectively and use CVA for lateral connections. The Slow path takes four frames as inputs sampled with a temporal stride of 16 and the Fast path takes 16 frames sampled with a stride of 4. As SlowFast captures multiscale temporal information by varying the frame rate to the two streams, the temporal duration for the tubelets is set to 1 in this case. Table 1(d) shows that our method is significantly more accurate than the SlowFast method whilst also using fewer FLOPs.
3 Comparison to the state of the art
We compare to the state-of-the-art across six different datasets. We evaluate models with four temporal- and three spatial-views per video clip, following . To make the notation more concise, we now use MTV-B to refer to B/2+S/4+Ti/8, MTV-L to refer to L/2+B/4+S/8+Ti/16 and MTV-H to refer to H/2+B/4+S/8+Ti/16. Except for Kinetics, all our models start from a Kinetics 400 checkpoint and then are fine-tuned on the target datasets following .
Figure 3 compares our proposed MTV to its single-view counterpart, ViViT Factorized Encoder (FE) at every model scale on Kinetics 400. We compare to ViViT-FE using tubelets with a temporal dimension of , as the authors obtained the best performance with this.
We can control the complexity of MTV by increasing or decreasing used in each view. For example, increasing from 2 to 4 for the smallest view (and proportionally increasing for all other views) will roughly reduce the input tokens by half for each view, and thus halve the total FLOPs for processing each input. Our method with for the smallest view consistently achieves higher accuracy than ViViT-FE at every complexity level while using fewer FLOPs, indicated by the green arrows pointing to the upper-left in Fig. 3(a). This further validates that processing more views in parallel enables us to achieve larger accuracy improvements than increasing the number of input tokens. If we set as in ViViT-FE, we use additional FLOPs, but increase significantly in accuracy too, as indicated by the green arrow pointing to the upper-right in Fig. 3(a).
Furthermore, note how our B/2 model (transformer depth of 12 layers) outperforms ViViT-L/2 (24 layers), whilst using less FLOPs. Similarly, our L/2 model outperforms ViViT-H/2. This shows that we can achieve greater accuracy improvements by processing multiple views in parallel than by increasing the depth for processing a single view.
Finally, note that Fig. 3(b) shows that our conclusions are also consistent when using the inference time to measure our model’s efficiency. Appendix A.1 also shows that these trends also hold when using an unfactorized backbone architecture of ViViT and MTV.
We compare to methods that are pretrained on ImageNet-1K, ImageNet-21K and those that do not utilize pretraining at all in the first part of Tab. 2(f). In the second part of the tables, we compare to methods that are pretrained on web-scale datasets such as Instagram 65M , JFT-300M , JFT-3B , WTS , Florence or HowTo100M . Observe that we achieve state-of-the-art results both with and without web-scale pretraining.
On Kinetics 400, our ImageNet-21K pretrained “Base” model improves the “Large” ViViT-FE model , which corresponds to a deeper, single-view equivalent of our model by 0.1% and 1.2% in Top-1 and Top-5 accuracy, whilst using 40% of the total FLOPs. Our higher resolution version improves further by 0.7% and 1.4% while still using slightly fewer FLOPs. On Kinetics 600, our “Base” model scores second to whose model structure is derived using architecture search on Kinetics 600 itself. We show significant improvements over on both Kinetics 400 and 700 for which the architecture of was not directly optimized for.
When using additional JFT-300M pretraining, our “Huge” model outperforms other recent transformer models using the same pretraining dataset . And when we utilize the Weak Textual Supervision (WTS) dataset of for pre-training, we substantially advance the best reported results on Kinetics: On Kinetics 400, we achieve a Top-1 accuracy of 89.9%, which improves upon the previous highest result (CoVeR ) by 2.7%. Similarly, on Kinetics 600, we achieve a Top-1 of 90.3%, which is an absolute improvement of 2.4% on . On Kinetics 700, we achieve 83.4%, which improves even further by 3.6% over . We also improve upon R3D-RS , which also used WTS pretraining, by 6.4% and 6.0% on Kinetics-400 and -600.
Following the standard protocol , we report Top-1 action-, verb- and noun-accuracies with action accuracy being the primary metric. Our results are averaged over crops as additional spatial crops did not help. Both our MTV-B and MTV-B(320p) significantly improve the previous state-of-the-art on noun classes, and MTV-B(320p) achieves a new state-of-the-art of 48.6% on actions. With WTS pretraining and increasing resolution, we improved the results to 50.5%. We found that additional data augmentation (detailed in Appendix A.3) has to be used to achieve good performance (as also observed by ) as this is the smallest dataset of all six with 67,000 training examples.
This dataset consists of class labels such as “move to left” and “pointing to right” . As the model needs to explicitly reason about direction, we do not perform random horizontal or vertical flipping as data augmentation on this dataset as also done by . We improve substantially over ViViT-L-FE , which corresponds to a deeper single-view equivalent of our model by 2.6%, and also improve upon MFormer by 0.4%.
Our MTV-L model significantly improves over the previous state-of-the-art by 1.5% in Top-1 accuracy. Moreover, our model with ImageNet-21K pretraining even outperforms VATT , which was pretrained on HowTo100M , a dataset consisting of around 100M video clips. When using WTS pre-training, we improve our accuracy even further, achieving 47.2%.
Conclusion
We have presented a simple method for capturing multi-resolution temporal context in transformer architectures, based on processing multiple “views” of the input video in parallel. We have demonstrated that our approach performs better, in terms of accuracy/computation trade-offs than increasing the depth of current single-view architectures. Furthermore, we have achieved state-of-the-art results on six popular video classification datasets. These results were then further improved with large-scale pretraining .
Although we have improved upon the state-of-the-art, there is still a large room for improvement on datasets other than Kinetics. Furthermore, we have relied on models pretrained on large image- or video-datasets for initialization. Reducing this dependence on supervised pretraining is a clear avenue of future research. We have conducted thorough ablations on standard transformer architectures , and will investigate if our approach is complementary to recent, spatial-pyramid based multiscale transformer encoders such as MViT and Swin .
Video classification models can be used in a wide range of applications. We are unaware of all potential applications, but are mindful that each application has its own merits, and that also depends on the intentions of the individuals building and using these systems. We also note that training datasets may contain biases that models trained on them are unsuitable for certain applications.
References
Appendix A Additional experiments
In this Appendix, we provide additional experimental details. Section A.1 provides accuracy-FLOPs and accuracy-throughput comparison between two model variants of ViViT and MTV. Section A.2 provides the effect of spatial resolution of tubelets. Section A.3 and Section A.4 provides details of our training hyperparameters and model configurations used in our experiments.
We present additional results by changing the transformer architecture used within our multiview encoder. Specifically, we use the unfactorized ViViT transformer encoder (Model 1 of ). In this variant, each transformer encoder layer computes self-attention over all spatio-temporal tokens. This makes our multiview transformer encoder cover a wide range of spatial and temporal dimensions across different views. A one-layer MLP with hidden dimension of is used as the global encoder for our unfactorized MTV model.
As shown in Fig. 4, MTV (unfactorized) consistently outperforms its single-view counterpart (i.e. ViViT unfactorized) for every scale (see Fig. 4(a)) and corresponds to a better accuracy-throughput curve as shown in Fig. 4(b). Note how MTV can more than double the throughput of ViViT unfactorized, whilst still improving its accuracy, for each model scale. Specifically, MTV (unfactorized) H/4+B/8+S/16+Ti/32 model leads to a significant speed-up by while still keeping a higher accuracy of improvement compared to ViViT-H.
Moreover, we report the accuracy-throughput comparison between MTV and ViViT factorized model (ViViT-FE) in Fig. 4(d). Note that the accuracy-FLOPs comparison is already reported in paper Section 4.3. The improvements in accuracy-throughput, and accuracy-FLOPs remain significant in this setting.
Note that the unfactorized ViViT transformer encoder, which attends to all spatio-temporal tokens, is less efficient than the Factorized Encoder architecture that we used in the main paper. However, we achieve larger relative improvements in accuracy/computation trade-offs compared to the corrsponding single-view ViViT baseline when using this encoder architecture.
A.2 Spatial resolution of tubelets
We study the effect of the spatial resolution of tubelets in Tab. LABEL:tab:spatial_ablations. We use our B/4 + Ti/16 model variant, and vary the spatial resolution of the tubelets. Our results indicate that the accuracy is primarily impacted by the spatial resolution of the large encoder. We also note that processing more tokens, and thus using more computation, typically results in higher accuracies.
A.3 Hyperparameters for each datasets
Table 4 details the hyperparamters used in all of our experiments. We use synchronous SGD with momentum, a cosine learning rate schedule with linear warmup, and a batch size of 64 for all experiments on the Kinetic datasets. We found that larger batch size and additional regularization are helpful when training on the smaller Epic Kitchens and Something-Something v2 datasets, as also noted by .
A.4 Model configurations
Table 5 summarizes our model configurations of each view for our multiview transformer encoder. For the backbone of each view, we consider five ViT variants, “Tiny”, “Small”, “Base”, “Large”, and “Huge”. Their settings strictly follow the ones defined in BERT and ViT . For the global encoder, all model variants of MTV use the same global encoder which follows the “Base” architecture, except that the number of heads is set to 8 instead of 12. The reason is that the hidden dimension of the tokens should be divisible by the number of heads for multi-head attention, and the number of hidden dimensions across all backbone sizes is divisible by 8 (as shown in Tab. 5). All model variants of MTV (unfactorized) use a one-layer MLP with the same hidden dimension as the “Base” architecture.