MoViNets: Mobile Video Networks for Efficient Video Recognition
Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, Boqing Gong
Introduction
Efficient video recognition models are opening up new opportunities for mobile camera, IoT, and self-driving applications where efficient and accurate on-device processing is paramount. Despite recent advances in deep video modeling, it remains difficult to find models that run on mobile devices and achieve high video recognition accuracy. On the one hand, 3D convolutional neural networks (CNNs) offer state-of-the-art accuracy, but consume copious amounts of memory and computation. On the other hand, 2D CNNs require far fewer resources suitable for mobile and can run online using frame-by-frame prediction, but fall short in accuracy.
Many operations that make 3D video networks accurate (e.g., temporal convolution, non-local blocks , etc.) require all input frames to be processed at once, limiting the opportunity for accurate models to be deployed on mobile devices. The recently proposed X3D networks provide a significant effort to increase the efficiency of 3D CNNs. However, they require large memory resources on large temporal windows which incur high costs, or small temporal windows which reduce accuracy. Other works aim to improve 2D CNNs’ accuracy using temporal aggregation , however their limited inter-frame interactions reduce these models’ abilities to adequately model long-range temporal dependencies like 3D CNNs.
This paper introduces three progressive steps to design efficient video models which we use to produce Mobile Video Networks (MoViNets), a family of memory and computation efficient 3D CNNs.
We first define a MoViNet search space to allow Neural Architecture Search (NAS) to efficiently trade-off spatiotemporal feature representations.
We then introduce Stream Buffers for MoViNets, which process videos in small consecutive subclips, requiring constant memory without sacrificing long temporal dependencies, and which enable online inference.
Finally, we create Temporal Ensembles of streaming MoViNets, regaining the slightly lost accuracy from the stream buffers.
First, we design the MoViNet search space to explore how to mix spatial, temporal, and spatiotemporal operations such that NAS can find optimal feature combinations to trade-off efficiency and accuracy. Figure 1 visualizes the efficiency of the generated MoViNets. MoViNet-A0 achieves similar accuracy to MobileNetV3-large+TSM on Kinetics 600 with 75% fewer FLOPs. MoViNet-A6 achieves state-of-the-art 83.5% accuracy, 1.6% higher than X3D-XL , requiring 60% fewer FLOPs.
Second, we create streaming MoViNets by introducing the stream buffer to reduce memory usage from linear to constant in the number of input frames for both training and inference, allowing MoViNets to run with substantially fewer memory bottlenecks. E.g., the stream buffer reduces MoViNet-A5’s memory usage by 90%. In contrast to traditional multi-clip evaluation (test-time data augmentation) approaches which also reduce memory, a stream buffer carries over temporal dependencies between consecutive non-overlapping subclips by caching feature maps at subclip boundaries. The stream buffer allows for a larger class of operations to enhance online temporal modeling than the recently proposed temporal shift . We equip the stream buffer with temporally unidirectional causal operations like causal convolution , cumulative pooling, and causal squeeze-and-excitation with positional encoding to force temporal receptive fields to look only into past frames, enabling MoViNets to operate incrementally on streaming video for online inference. However, the causal operations come at a small cost, reducing accuracy on Kinetics 600 by 1% on average.
Third, we temporally ensemble MoViNets, showing that they are more accurate than single large networks while achieving the same efficiency. We train two streaming MoViNets independently with the same total FLOPs as a single model and average their logits. This simple technique gains back the loss in accuracy when using stream buffers.
Taken together, these three techniques create MoViNets that are high in accuracy, low in memory usage, efficient in computation, and support online inference. We search for MoViNets using the Kinetics 600 dataset and test them extensively on Kinetics 400 , Kinetics 700 , Moments in Time , Charades , and Something-Something V2 .
Related Work
Deep neural networks have made remarkable progress for video understanding . They extend 2D image models with a temporal dimension, most notably incorporating 3D convolution .
Improving the efficiency of video models has gained increased attention . Some works explore the use of 2D networks for video recognition by processing videos in smaller segments followed by late fusion . The Temporal Shift Module uses early fusion to shift a portion of channels along the temporal axis, boosting accuracy while supporting online inference.
WaveNet introduces causal convolution, where the receptive field of a stack of 1D convolutions only extends to features up to the current time step. We take inspiration from other works using causal convolutions to design stream buffers for online video model inference, allowing frame-by-frame predictions with 3D kernels.
The use of NAS with multi-objective architecture search has also grown in interest, producing more efficient models in the process for image recognition and video recognition . We make use of TuNAS , a one-shot NAS framework which uses aggressive weight sharing that is well-suited for computation intensive video models.
Deep ensembles are widely used in classification challenges to boost the performance of CNNs . More recent results indicate that deep ensembles of small models can be more efficient than single large models on image classification , and we extend these findings to video classification.
Mobile Video Networks (MoViNets)
This section describes our progressive three-step approach to MoViNets. We first detail the design space to search for MoViNets. Then we define the stream buffer and explain how it reduces the networks’ memory footprints, followed by the temporal ensembling to improve accuracy.
Following the practice of 2D mobile network search , we start with the TuNAS framework , which is a scalable implementation of one-shot NAS with weight sharing on a supernetwork of candidate models, and repurpose it for 3D CNNs for video recognition. We use Kinetics 600 as the video dataset to search over for all of our models, consisting of 10-second video sequences each at 25fps for a total of 250 frames.
We build our base search space on MobileNetV3 , which provides a strong baseline for mobile CPUs. It consists of several blocks of inverted bottleneck layers with varying filter widths, bottleneck widths, block depths, and kernel sizes per layer. Similar to X3D , we expand the 2D blocks in MobileNetV3 to deal with 3D video input. Table 1 provides a basic overview of the search space, detailed as follows.
We denote by and (5fps) the dimensions and frame stride, respectively, of the input to the target MoViNets. For each block in the network, we search over the base filter width and the number of layers to repeat within the block. We apply multipliers over the feature map channels within every block, rounded to a multiple of 8. We set blocks, with strided spatial downsampling for the first layer in each block except the 4th block (to ensure the last block has spatial resolution ). The blocks progressively increase their feature map channels: . The final convolution layer’s base filter width is , followed by a D dense layer before the classification layer.
With the new time dimension, we define the 3D kernel size within each layer, , to be chosen as one of the following: {1x3x3, 1x5x5, 1x7x7, 5x1x1, 7x1x1, 3x3x3, 5x3x3} (we remove larger kernels from consideration). These choices enable a layer to focus on and aggregate different dimensional representations, expanding the network’s receptive field in the most pertinent directions while reducing FLOPs along other dimensions. Some kernel sizes may benefit from having different numbers of input filters, so we search over a range of bottleneck widths defined as multipliers in relative to . Each layer surrounds the 3D convolution with two 1x1x1 convolutions to expand and project between and . We do not apply any temporal downsampling to enable frame-wise prediction.
Instead of applying spatial squeeze-and-excitation (SE) , we use SE blocks to aggregate spatiotemporal features via 3D average pooling, applying it to every bottleneck block as in . We allow SE to be searchable, optionally disabling it to conserve FLOPs.
Our base search space forms the basis for MoViNet-A2. For the other MoViNets, we apply a compound scaling heuristic similar to the one used in EfficientNet . The major difference in our approach is that we scale the search space itself rather than a single model (i.e., search spaces for models A0-A5). Instead of finding a good architecture and then scaling it, we search over all scalings of all architectures, broadening the range of possible models.
We use a small random search to find the scaling coefficients (with an initial target of 300 MFLOPs per frame), which roughly double or halve the expected size of a sampled model in the search space. For the choice of coefficients, we resize the base resolution , frame stride , block filter width , and block depths . We perform the search on different FLOPs targets to produce a family of models ranging from MobileNetV3-like sizes up to the sizes of ResNet3D-152 . Appendix Aprovides more details of the search space, the scaling technique, and a description of the search algorithm.
The MoViNet search space gives rise to a family of versatile networks, which outperform state-of-the-art efficient video recognition CNNs on popular benchmark datasets. However, their memory footprints grow proportionally to the number of input frames, making them difficult to handle long videos on mobile devices. The next subsection introduces a stream buffer to reduce the networks’ memory consumption from linear to constant in video length.
2 The Stream Buffer with Causal Operations
Suppose we have an input video with frames that may cause a model to exceed a set memory budget. A common solution to reduce memory is multi-clip evaluation , where the model averages predictions across overlapping subclips with frames each, as seen in Figure 2 (left). It reduces memory consumption to . However, it poses two major disadvantages: 1) It limits the temporal receptive fields to each subclip and ignores long-range dependencies, potentially harming accuracy. 2) It recomputes frame activations which overlap, reducing efficiency.
To overcome the above mentioned limitations, we propose stream buffer as a mechanism to cache feature activations on the boundaries of subclips, allowing us to expand the temporal receptive field across subclips and requiring no recomputation, as shown in Figure 2 (right).
Formally, let be the current subclip (raw input or activation) at step , where we split the video into adjacent non-overlapping subclips of length each. We start with a zero-initialized tensor representing our buffer with length along the time dimension and whose other dimensions match . We compute the feature map of the buffer concatenated () with the subclip along the time dimension as:
where represents a spatiotemporal operation (e.g., 3D convolution). When processing the next clip, we update the contents of the buffer to:
where we denote as a selection of the last frames of the concatenated input. As a result, our memory consumption is dependent on , which is constant as the total video frames or number of subclips increase.
The Temporal Shift Module (TSM) can be seen as a special case of the stream buffer, where and is an operation that shifts a proportion of channels in the buffer to the input before computing a spatial convolution at frame .
2.1 Causal Operations
A reasonable approach to fitting 3D CNNs’ operations to the stream buffer is to enforce causality, i.e., any features must not be computed from future frames. This has a number of advantages, including the ability to reduce a subclip down to a single frame without affecting activations or predictions, and enables 3D CNNs to work on streaming video for online inference. While it is possible to use non-causal operations, e.g., buffering in both temporal directions, we would lose online modeling capabilities which is a desirable property for mobile.
By leveraging the translation equivariant property of convolution, we replace all temporal convolutions with CausalConvs , effectively making them unidirectional along the temporal dimension. Concretely, we first compute padding to balance the convolution across all axes and then move any padding after the final frame and merge it with any padding before the first frame. See Appendix Cfor an illustration of how the receptive field differs from standard convolution, as well as a description of the causal padding algorithm.
When using a stream buffer with CausalConv, we can replace causal padding with the buffer itself, carrying forward the last few frames from a previous subclip and copying them into the padding of the next subclip. If we have a temporal kernel size of (and we do not use any strided sampling), then our padding and therefore buffer width becomes . Usually, which implies , resulting in a small memory footprint. Stream buffers are only required before layers that aggregate features across multiple frames, so spatial and pointwise convolutions (e.g., 1x3x3, 1x1x1) can be left as-is, further saving memory.
We use CGAP to approximate any global average pooling involving the temporal dimension. For any activations up to frame , we can compute this as a cumulative sum:
where represents a tensor of activations. To compute CGAP causally, we keep a single-frame stream buffer storing the cumulative sum up to .
We denote CausalSE as the application of CGAP to SE, where we multiply the spatial feature map at frame with the SE computed from . From our empirical results, CausalSE is prone to instability likely due to the SE projection layers have a difficult time determining the quality of the CGAP estimate, which has high variance early in the video. To combat this problem, we apply a sine-based fixed positional encoding (PosEnc) scheme inspired by Transformers . We directly use frame index as the position and sum the vector with CGAP output before applying the SE projection.
2.2 Training and Inference with Stream Buffers
To reduce the memory requirements during training, we use a recurrent training strategy where we split a given batch of examples into subclips, applying a forward pass that outputs a prediction for each subclip, using stream buffers to cache activations. However, we do not backpropagate gradients past the buffer so that the memory of previous subclips can be deallocated. Instead, we compute losses and accumulate computed gradients between subclips, similar to batch gradient accumulation. This allows the network to account for all frames, performing forward passes before applying the gradients. This training strategy allows the network to learn longer term dependencies thus results in better accuracy than a model trained with shorter video length (see Appendix C).
We can set to any value without affecting accuracy. However, ML accelerators (e.g., GPUs) benefit from multiplying large tensors, so for training we typically set a value of . This accelerates training while allowing careful control of memory cost.
One major benefit of using causal operations like CausalConv and CausalSE is to allow a 3D video CNN to work online. Similar to training, we use the stream buffer to cache activations between subclips. However, we can set the subclip length to a single frame () for maximum memory savings. This also reduces the latency between frames, enabling the model to output predictions frame-by-frame on a streaming video, accumulating new information incrementally like a recurrent network (RNN) . But unlike traditional convolutional RNNs, we can input a variable number of frames per step to produce the same output. For streaming architectures with CausalConv, we predict a video’s label by pooling the frame-by-frame output features using CGAP.
3 Temporal Ensembles
The stream buffers can reduce MoViNets’ memory footprints up to an order of magnitude in the cost of about 1% accuracy drop on Kinetics 600. We can restore this accuracy using a simple ensembling strategy. We train two MoViNets independently with the same architecture, but halve the frame-rate, keeping the temporal duration the same (resulting in half the input frames). We input a video into both networks, with one network having frames offset by one frame and apply an arithmetic mean on the unweighted logits before applying softmax. This method results in a two-model ensemble with the same FLOPs as a single model before halving the frame-rate, providing prediction with enriched representations. In our observations, despite the fact that both models in the ensemble may have lower accuracy than the single model individually, together when ensembled they can have higher accuracy than the single model.
Experiments on Video Classification
In this section, we evaluate MoViNets’ accuracy, efficiency, and memory consumption during inference on five representative action recognition datasets.
We report results on all Kinetics datasets, including Kinetics 400 , Kinetics 600 , and Kinetics 700 , which contain 10-second, 250-frame video sequences at 25 fps labeled with 400, 600, and 700 action classes, respectively. We use examples that are available at the time of writing, which is 87.5%, 92.8%, and 96.2% of the training examples respectively (see Appendix C). Additionally, we experiment with Moments in Time , containing 3-second, 75-frame sequences at 25fps in 339 action classes, and Charades , which has variable-length videos with 157 action classes where a video can contain multiple class annotations. We include Something-Something V2 and Epic Kitchens 100 results in Appendix C.
For each dataset, all models are trained with RGB frames from scratch, i.e., we do not apply any pretraining. For all datasets, we train with 64 frames (except when the inference frames are fewer) at various frame-rates, and run inference with the same frame-rate.
We run TuNAS using Kinetics 600 and keep 7 MoViNets each having a FLOPs target used in . As our models get larger, our scaling coefficients increase the input resolution, number of frames, depth, and feature width of the networks. We also experiment with AutoAugment augmentation used in image classification, i.e., we sample a random image augmentation for each video and apply the same augmentation for each frame. For the architectures of the 7 models as well as training hyperparameters, see Appendix B.
We evaluate all our models with a single clip sampled from input video with a fixed temporal stride, covering the entire video duration. When the single-clip and multi-clip evaluations use the same number of frames in total so that FLOPs are equivalent, we find that single-clip evaluation yields higher accuracy (see Appendix C). This can be due in part to 3D CNNs being able to model longer-range dependencies, even when evaluating on many more frames than it was trained on. Since existing models commonly use multi-clip evaluation, we report the total FLOPs per video, not per clip, for a fair comparison.
However, single-clip evaluation can greatly inflate a network’s peak memory usage (as seen in Figure 1), which is likely why multi-clip evaluation is commonly used in previous work. The stream buffer eliminates this problem, allowing MoViNets to predict like they are embedding the full video, and incurs less peak memory than multi-clip evaluation.
We also reproduce X3D , arguably the most related work to ours, to test its performance under single-clip and 10-clip evaluation to provide more insights. We denote 30-clip to be the evaluation strategy with 10 clips times three spatial crops for each video, while 10-clip just uses one spatial crop. We avoid any spatial augmentation or temporal sampling in MoViNets during inference to improve efficiency.
1 Comparison Results on Kinetics 600
Table 2 presents the main results of seven MoViNets on Kinetics 600 before applying the stream buffer, mainly compared with various X3D models , which are recently developed for efficient video recognition. The columns of the table correspond to the Top-1 classification accuracy; GFLOPs per video a model incurs; resolution of the input video frame (where we shorten to 224); input frames per video, where means the 30-clip evaluation with 4 frames as input in each run; frames per second (FPS), determined by the temporal stride in the search space for MoViNets; and a network’s number of parameters.
MoViNet-A0 has fewer GFLOPs and is 10% more accurate than the frame-based MobileNetV3-S (where we train MobileNetV3 using our training setup, averaging logits across frames). MoViNet-A0 also outperforms X3D-S in terms of both accuracy and GFLOPs. MoViNet-A1 matches the GFLOPs of X3D-S, but its accuracy is 2% higher than X3D-S.
Growing the target GFLOPs to the range between X3D-S and 30-clip X3D-XS, we arrive at MoViNet-A2. We can achieve a little higher accuracy than 30-clip X3D-XS or X3D-M by using almost half of their GFLOPs. Additionally, we include the frame-by-frame MobileNetV3-L and verify that it can benefit from TSM by about 3%.
There are more significant margins between larger MoViNets (A3–A6) and their counterparts in the X3D family. It is not surprising because NAS should intuitively be more advantageous over the handcrafting method for X3D when the design space is large. MoViNet-A5 and MoViNet-A6 outperform several state-of-the-art video networks, including recent Transformer models like ViViT and TimeSformer (see the last 6 rows of Table 2). MoViNet-A6 with AutoAugment achieves 84.8% accuracy (without pretraining) while still being substantially more efficient than comparable models (most often by an order of magnitude). Even when compared to fully Transformer models like TimeSformer-HR , MoViNet-A6 outperforms it by 1% accuracy and using 40% of the FLOPs.
Our base MoViNet architectures may consume lots of memory in the absence of modifications, especially as the model sizes and input frames grow. Using the stream buffer with causal operations, we can have an order of magnitude peak memory reduction for large networks (MoViNets A3-A6), as shown in the last column of Table 3.
Moreover, Figure 3 visualizes the streaming architectures’ effect on memory. From the left panel at the top, we see that our MoViNets are more accurate and more memory-efficient across all model sizes compared to X3D, which employs multi-clip evaluation. We also demonstrate constant memory as we scale the total number of frames in the input receptive field at the top’s right panel. The bottom panel indicates that the streaming MoViNets remain efficient in terms of the GFLOPs per input video.
We also apply our stream buffer to ResNet3D-50 (see the last two rows in Table 3). However, we do not see as much of a memory reduction, likely due to larger overhead when using full 3D convolution as opposed to the depthwise convolution in MoViNets.
We see from Table 3 only a small 1% accuracy drop across all models after applying the stream buffer. We can restore the accuracy using the temporal ensembling without any additional inference cost. Table 3 reports the effect of ensembling two models trained at half the frame rate of the original model (so that GFLOPs remain the same). We can see the accuracy improvements in all streaming architectures, showing that ensembling can bridge the gap between streaming and non-streaming architectures, especially as model sizes grow. It is worth noting that, unlike prior works, the ensembling balances accuracy and efficiency (GFLOPs) in the same spirit as , not just to boost the accuracy.
2 Comparison Results on Other Datasets
Figure 4 summarizes the main results of MoViNets on all the five datasets along with state-of-the-art models that have results reported on the respective datasets. We compare MoViNets with X3D , MSNet , TSM , ResNet3D , SlowFast , EfficientNet-L2 , TVN , SRTG , and AssembleNet . Appendix Ctabulates the results with more details.
Despite only searching for efficient architectures on Kinetics 600, NAS yields models that drastically improve over prior work on other datasets as well. On Moments in Time, our models are 5-8% more accurate than Tiny Video Networks (TVNs) at low GFLOPs, and MoViNet-A5 achieves 39.9% accuracy, outperforming AssembleNet (34.3%) which uses optical flow as additional input (while our models do not). On Charades, MoViNet-A5 achieves the accuracy of 63.2%, beating AssembleNet++ (59.8%) which uses optical flow and object segmentation as additional inputs. Results on Charades provide evidence that our models are also capable of sophisticated temporal understanding, as these videos can have longer duration clips than what is seen in Kinetics and Moments in Time.
3 Additional Analyses
We provide some ablation studies about some critical MoViNet operations in Table 4. For the base network without the stream buffer, SE is vital for achieving high accuracy; MoViNet-A1’s accuracy drops by 2.9% if we remove SE. We see a much larger accuracy drop when using CausalConv without SE than CausalConv with a global SE, which indicates that the global SE can take some of the role of standard Conv to extract information from future frames. However, when we switch to a fully streaming architecture with CausalConv and CausalSE, this information from future frames is no longer available, and we see a large drop in accuracy, but still significantly improved from CausalConv without SE. Using PosEnc, we can gain back some accuracy in the causal model.
We provide the architecture description of MoViNet-A2 in Table 5 — Appendix Bhas the detailed architectures of other MoViNets. Most notably, the network prefers large bottleneck width multipliers in the range [2.5, 3.5], often expanding or shrinking them after each layer. In contrast, X3D-M with similar compute requirements has a wider base feature width with a smaller constant bottleneck multiplier of 2.25. The searched network prefers balanced 3x3x3 kernels, except at the first downsampling layers in the later blocks, which have 5x3x3 kernels. The final stage almost exclusively uses spatial kernels of size 1x5x5, indicating that high-level features for classification benefit from mostly spatial features. This comes at a contrast to S3D , which reports improved efficiency when using 2D convolutions at lower layers and 3D convolutions at higher layers.
3.1 Hardware Benchmark
MoViNets A0, A1, and A2 represent the fastest models that would most realistically be used on mobile devices. We compare them with MobileNetV3 in Figure 5 with respect to both FLOPs and real-time latency on an x86 Intel Xeon W-2135 CPU at 3.70GHz. These models are comparable in per-frame computation cost, as we evaluate on 50 frames for all models. From these results we can conclude that streaming MoViNets can run faster on CPU while being more accurate at the same time, even with temporal modifications like TSM. While there is a discrepancy between FLOPs and latency, searching over a latency target explicitly in NAS can reduce this effect. However, we still see that FLOPs is a reasonable proxy metric for CPU latency, which would translate well for mobile devices.
We also show benchmarks for MoViNets running on an Nvidia V100 GPU in Table 6. Similar to mobile CPU, our streaming model latency is comparable to single-clip X3D models. However, we do note that MobileNetV3 can run faster than our networks on GPU, showing that the FLOPs metric for NAS has its limitations. MoViNets can be made more efficient by targeting real hardware instead of FLOPs, which we leave for future work.
Conclusion
MoViNets provide a highly efficient set of models that transfer well across different video recognition datasets. Coupled with stream buffers, MoViNets significantly reduce training and inference memory cost while also supporting online inference on streaming video. We hope our approach to designing MoViNets can provide improvements to future and existing models, reducing memory and computation costs in the process.
References
Appendices
We supplement the main text by the following materials.
Appendix A provides more details of the search space, the technique to scale the search space, and the search algorithm.
Appendix B is about the neural architectures of MoViNets A0-A7.
Appendix C reports additional results on the datasets studied in the main text along with ablation studies.
To produce models that scale well, we progressively expand the search space across width, depth, input resolution, and frame rate, like EfficientNet . Specifically, we use a single scaling parameter to define the size of our search space. Then we define the following coefficients:
such that . This will ensure that an increase in by 1 will multiply the average model size in the search space by 4. Here we use a multiplier of 4 (instead of 2) to spread out our search spaces so that we can run the same search space with multiple efficiency targets and sample our desired target model size from it.
As a result, our parameters for a given search space is the following:
We round each of the above parameters to the nearest multiple of 8. If , this forms the base search space for MoViNet-A2. Note that is defined relative to , so we do not need coefficients for it.
We found coefficients using a random search over these parameters. More specifically, we select values in the range at increments of 0.05 to represent possible values of the coefficients. We ensure that the choice of coefficients is such that , where the initial computation target for each frame is 300 MFLOPs. For each combination, we scale the search space by the coefficients where , and randomly sample three architectures from each search space. We train models for a selected search space for 10 epochs, averaging the results of the accuracy for that search space. Then we select the coefficients that maximize the resulting accuracy. Instead of selecting a single set of coefficients, we average the top 5 candidates to produce the final coefficients. While the sample size of models is small and would be prone to noise, we find that the small averages work well in practice.
A.2 Search Algorithm
During search, we train a one-shot model using TuNAS that overlaps all possible architectures into a hypernetwork. At every step during optimization, we alternate between learning the network weights and learning a policy which we use to randomly sample a path through the hypernetwork to produce an initially random network architecture. is learned using REINFORCE , optimized on the quality of sampled architectures, defined as the absolute reward consisting of the sampled network’s accuracy and cost. At each stage, the RL controller must choose a single categorical decision to select an architectural component. The network architecture is a result of binding a value to each decision. For example, the decision might choose between a spatial 1x3x3 convolution and a temporal 5x1x1 convolution. We use FLOPs as the cost metric for architecture search, and use Kinetics 600 as the dataset to optimize for efficient video networks. During search, we obtain validation set accuracies on a held-out subset of the Kinetics 600 training set, training for a total of 90 epochs.
The addition of SE to our search space increases FLOPs by such a small amount (%) that the search enables it for all layers. SE plays a similar role as the feature gating in S3D-G , except with a nonlinear squeeze inside the projection operation.
A.3 Stream Buffers for NAS
We apply the stream buffers to MoViNets as a separate step after NAS in the main text. We can also leverage them for NAS to reduce memory usage during search. Memory poses one of the biggest challenges for NAS, as models are forced to use a limited number of frames and small batch sizes to be able to keep the models in memory during optimization. While this does not prevent us from performing search outright, it requires the use of accelerators with very high memory requirements, requiring a high cost of entry. To circumvent this, we can use stream buffers with a small clip size to reduce memory. As a result, we can increase the total embedded frames and increase the batch size to provide better model accuracy estimation while running NAS. Table 7 provides an example of an experiment where the use of a stream buffer can reduce memory requirements in this manner. Using a stream buffer, we can reduce the input size from a single clip of 16 frames to 2 clips of 8 frames each, and double the batch size. This results in a relatively modest increase in memory, compared to not using the buffer where we can run into out-of-memory (OOM) issues.
We note that the values of in each layer influences the memory consumption of the model. This is dependent entirely on the temporal kernel width of the 3D convolution. If , then we only need to cache the last 4 frames. could be larger, but but it will result in extra frames we will discard, so we set it to the smallest value to conserve memory. Therefore, it is not necessary to specify it directly with NAS, as NAS is only concerned with the kernel sizes. However, we can add an objective to NAS to minimize memory usage, which will apply pressure to reduce the temporal kernel widths and therefore will indirectly affect the value of in each layer. Reducing memory consumption even further by keeping kernel sizes can be explored in future work.
B Architectures of MoViNets
See Tables 17, 18, 19, 20, 21, and 22 for the architecture definitions of MoViNet A0-A5 (we move the tables to the final pages of the Appendices to reduce clutter). For MoViNet-A6, we ensemble architectures A4 and A5 using the strategy described in the main text, i.e., we train both models independently and apply an arithmetic mean on the logits during inference. All layers of all models have SE layers enabled, so we remove this search hyperparameter from all tables for brevity.
We apply additional changes to our architectures and model training to improve performance even further. To improve convergence speed in searching and training we use ReZero by applying zero-initialized learnable scalar weights that are multiplied with features before the final sum in a residual block. We also apply skip connections that are traditionally used in ResNets, adding a 1x1x1 convolution in the first layer of each block which may change the base channels or downsample the input. However, we modify this to be similar to ResNet-D where we apply 1x3x3 spatial average pooling before the convolution to improve feature representations.
We apply Polyak averaging to the weights after every optimization step, using an Exponential Moving Average (EMA) with decay 0.99. We adopt the Hard Swish activation function, which is a variant of SiLU/Swish proposed by MobileNetV3 that is friendlier to quantization and CPU inference. We use the RMSProp optimizer with momentum 0.9 and a base learning rate of 1.8. We train for 240 epochs with a batch size of 1024 with synchronized batch normalization on all datasets and decay the learning rate using a cosine learning rate schedule with a linear warm-up of 5 epochs.
We use a softmax cross-entropy loss with label smoothing 0.1 during training, except for Charades where we apply sigmoid cross-entropy to handle the multiple-class labels per video. For Charades, we aggregate predictions across frames similar to AssembleNet , where we apply a softmax across frames before applying temporal global average pooling to find multiple action classes that may occur in different frames.
Some works also expand the resolution for inference. For instance, X3D-M trains with a resolution while evaluating when using spatial crops. We evaluate all of our models on the same resolution as training to make sure the FLOPs per frame during inference is unchanged from training.
Our choice of frame-rates can vary from model to model, providing different optimality depending on the architecture. We plot the accuracy of training various MoViNets on Kinetics 600 with different frame-rates in Figure 6. Most models have good efficiency at 50 frames (5fps) or 80 frames (8fps) per video. However, we can see MoViNet-A4 benefits from a higher frame-rate of 12fps. For Charades, we use 64 frames at 6fps for both training and inference.
C More Implementation Details and Experiments
To make a temporal convolution operation causal, we can apply a simple padding trick which shifts the receptive field forward such that the convolutional kernel is centered at the frame furthest into the future. Figure 7 illustrates this effect. With a normal 3D convolution operation with kernel size and stride , the padding with respect to dimension is given as:
where are the left and right padding amounts respectively. For causal convolutions, we transform as:
such that the effective temporal receptive field of a voxel at time position only spans .
C.2 Additional Details of Datasets
We note that for all the Kinetics datasets are gradually shrinking over time due to videos being taken offline, making it difficult to compare against less recent works. We report on the most recently available videos. While putting our work at a disadvantage compared to previous work, we wish to make comparisons more fair for future work. Nevertheless, we report the numbers as-is and report the reduction of examples in the datasets in Table 8.
We report full results of all models in the following tables: Table 9 for Kinetics 400, Table 10 for Kinetics 600 (with top-5 accuracy), Table 11 for Kinetics 700, Table 12 for Moments in Time, Table 13 for Charades, Table 14 for Something-Something V2, and Table 15 for Epic Kitchens 100. For a table of results on Kinetics 600, see Table 10 and also the main text.
C.3 Single-Clip vs. Multi-Clip Evaluation
We report all of our results on a single view without multi-clip evaluation. Additionally, we report the total number of frames used for evaluation and the frame rate (note that the evaluation frames can exceed the total number of frames in the reference video when subclips overlap).
As seen in Figure 1 and Table 2 (in the main text), switching from a multi-clip to single-clip X3D model on Kinetics 600 (where we cover the entire 10-second clip) results in much higher computational efficiency per video. Existing work typically factors out FLOPs in terms of FLOPs per subclip, but it can hide the true cost of computation, since we can keep adding more clips to boost accuracy higher.
We also evaluate the differences between training the same MoViNet-A2 model on smaller clips vs. longer clips and evaluating the models with multi-clip vs. single-clip, as seen in Figure 8. For multi-clip evaluation, we can see that accuracy improves when the number of clips fill the whole duration of the video (this can be seen at 5 clips for 8 training frames and at 3 clips for 16 training frames), and only very slightly improves as we add more clips. However, if we train MoViNet-A2 on 16 frames and evaluate on 80 frames (so that we cover all 10 seconds of the video), this results in higher accuracy than the same number of frames using multi-clip eval. Furthermore, we can boost this accuracy even higher if we use 48 frames to train our model. Using stream buffers, we can reduce memory usage of training so that we can train using 48 frames while only using the memory of embedding 16 frames at a time.
C.4 Streaming vs. Non-Streaming Evaluation
One question we have wondered is if the distribution of features learned is different from streaming and non-streaming architectures. In Figure 9, we plot the average accuracy across Kinetics 600 of a model evaluated on a single frame by embedding an entire video, pooling across spatial dimensions, and applying the classification layers independently on each frame.
We first notice that the accuracy MobileNetV3 and MoViNet-A2 exhibit a Laplace distribution, on average peaking at the center frame of each video. Since MobileNetV3 is evaluated on each frame independently, we can observe that the most salient part of the actions is on average in the video’s midpoint. This is a good indicator that the videos in Kinetics are trimmed very well to center around the most salient part of each action. Likewise, MoViNet-A2, with balanced 3D convolutions, has the same characteristics as MobileNetV3, just with higher accuracy.
However, the dynamics of streaming MoViNet-A2 with causal convolutions is entirely different. The distribution of accuracy fluctuates and varies more than non-streaming architectures. By removing the ability for the network to see all frames as a whole with causal convolutions, the aggregation of features is not the same as when using balanced convolutions. Despite this difference, overall, the accuracy difference across all videos is only about 1%. And by looking at top-5 accuracy in Table 10, we can notice that streaming architectures nearly perform the same, despite the apparent information loss when transitioning to a model with a time-unidirectional receptive field.
C.5 Long Video Sequences
Figure 10 shows how training clip duration affects the accuracy of a model evaluated at different durations. We can see that MoViNet can generalize well beyond the original clip duration it was trained with, always improving in accuracy with more frames. However, the model does notably worse if evaluated on clips with shorter durations than it was trained on. Longer clip duration for training translates to better accuracy for evaluation on longer clips overall. And with a stream buffer, we can train on even longer sequences to boost evaluation performance even higher.
However, we also see we can operate frame-by-frame with stream buffers, substantially saving memory, showing better memory efficiency than multi-clip approaches and requiring constant memory as the number of input frames increase (and therefore temporal receptive field). Despite the accuracy reduction, we can see MoViNet-Stream models perform very well on long video sequences and are still more efficient than X3D which requires splitting videos into smaller subclips. We encourage future work using multi-clip evaluation to report results without overlapping subclips, which not only provides a much more representative accuracy measurement, tends to be more efficient as well.
C.6 Stream Buffers with Other Operations
WaveNet introduces causal convolution, where the receptive field on a stack of 1D convolutions is forced to only see activations up to the current time step, as opposed to balanced convolutions which expand their receptive fields in both directions. We take inspiration from causal convolutions to design stream buffers. However, WaveNet only proposes 1D convolutions for generative modeling, using them for their autoregressive property. We generalize the idea of causal convolution to any local operation, and introduce stream buffers to be able to use causal operations for online inference, allowing frame-by-frame predictions. In addition, Transformer-XL caches activations in a temporal buffer much like our work, for use in long-range sequence modeling. However, the model is only causal across fixed sequences while our work can be causal across individual frames, and can even vary the number of frames in each clip (so long as frames are consecutive with no gaps or overlaps between clips). We can apply the same principle to other operations as well to generalize causal operations. Note that this approach is not inherently tied to any data type or modality. Stream buffers can also be used to model many kinds of temporal data, e.g., audio, text.
Additionally, support for efficient 3D convolutions on mobile devices is currently fragmented, while 2D convolutions are well supported. We include the option to search for (2+1)D architectures, splitting up any 3D depthwise convolutions into a 2D spatial convolution followed by a 1D temporal convolution. We show that trivially changing a 3D architecture to (2+1)D decreases FLOPs while also keeping similar accuracy, as seen in table 16. Here we define MoViNet-A2b as a searched model similar to MoViNet-A2.