Timeception for Complex Action Recognition

Noureldien Hussein, Efstratios Gavves, Arnold W. M. Smeulders

Introduction

In ordinary life, activities of daily living pop up frequently. Our conversations include actions like “cooking a meal” or “cleaning the house” much more frequently than actions like “jumping” or “cutting a cucumber”. The latter, which we call one-actions, exhibit one visual pattern, possibly repetitive. They are usually short in time, homogeneous in motion and coherent in form. In contrast, cooking a meal or cleaning the house are very different actions. We refer to them as complex actions, characterized by: i. They are typically composed of several one-actions, see figure 1. ii. These one-actions, contained in a complex action, exhibit large variations in their temporal duration and temporal order. iii. As a consequence of the composition, a complex action takes much longer to unfold. And, by the in-homogeneity in composition, the complex action needs to be sampled in full, not to miss crucial parts.

In the recent literature, the main focus is the recognition of short-range actions like in HMDB, UCF and Kinetics . Few attention has been paid to the recognition of long-range and complex actions, as in Charades and EventNet , which we study here. The first challenge is minute-long temporal modeling while maintaining attention to seconds-long details. Statistical temporal pooling, as applied in falls short of learning temporal order. Neural temporal modeling and spatio-temporal convolutions of various types successfully learns temporal order of 8 or 128 timesteps . But the computational cost is far beyond scaling up to 1000 timesteps needed for complex actions. The second challenge is tolerating variations in temporal extent and temporal order of one-actions. Related methods learn spatio-temporal convolutions with fixed-size kernels, which would be too rigid for complex actions. To address these challenges, we present Timeception, a novel convolutional layer dedicated only for temporal modeling. It learns long-range temporal dependencies with attention to short-range details. Plus, it tolerates the differences in temporal extent of one-actions comprising the complex action. As a result, we demonstrate success in recognizing the long and complex actions, and achieving state-of-the-art-results in Charades , Breakfast Actions and MultiTHUMOS .

The novelties of of this paper are: i. We introduce a convolutional temporal layer effectively and efficiently learn minute-long action ranges of 1024 timesteps, a factor of 8 longer than best related work. ii. We introduce multi-scale temporal kernels to account for large variations in duration of action components. iii. We use temporal-only convolutions, which are better suited for complex actions than spatiotemporal counterparts.

Related Work

Temporal Modeling. The stark difference between video and image classification is the temporal dimension, which necessitates temporal modeling. A widely used approach is statistical pooling: max and average pooling , attention pooling , rank pooling , dynamic images and context gating , to name a few. Beyond statistical pooling, vector aggregation is also used. uses Fisher Vector to aggregate spatio-temporal features over time, while extend VLAD to use local convolution features extracted from video frames. The downside of statistical pooling and vector aggregation is completely neglecting temporal patterns – an important visual cue.

Other strands of work use neural methods for temporal modeling. LSTMs are used to model the sequence in action videos . While TA-DenseNet extents DenseNet to exploit the temporal dimension. To our knownedge, no substantial improvements have been reported recently.

Short-range Action Recognition. Few works learn deep appearance features by frame-level classification of actions, using 2D CNNs. Others complement deep appearance features with shallow motion features, as IDT . Also, auxiliary image representations is fused with RGB signals: uses OpticalFlow channels, while use Dynamic Images. 3D CNNs are the natural evolution of their 2D counterparts. C3D proposes 3D CNNs to capture spatio-temporal patterns of 8 frames in a sequence. In the same vein, I3D inflates the kernels of ImageNet-pretrained 2D CNN to jump-start the training of 3D CNNs. While effective in short-range video sequences of few seconds, 3D convolutions are too computationally expensive to address minute-long videos, which is our focus.

Long-range Action Recognition. To learn long-range temporal patterns, uses CRF on top of CNN feature maps to model human activities. To learn video-wide representations, TRN learns relations between several video segments. TSN learns temporal structure in long videos. LTC considers different temporal resolutions as a substitute to bigger temporal windows. Inspired by self-attention , non-local networks proposes a 3D CNN with a long temporal footprint of 128 timesteps.

All aforementioned methods succeed in modeling temporal footprint of 128 timesteps (∼\sim4-5 sec) at max. In this work, we address complex actions with long-range temporal dependencies of up to 1024 timesteps, jointly.

Convolution Decomposition. CNNs succeed in learning spatial and spatiotemporal action concepts, but existing convolutions grow heavy in computation, specially at the higher layers where the number of channels can grow as much as 2k . To control the computational complexity, several works propose the decomposition of 2D and 3D convolutions. Xception argues that separable 2D convolutions are as effective as typical 2D convolutions. Similarly, S3D considers separable 2+1D convolutions to reduce the complexity of typical 3D convolutions. ResNet reduces the channel dimension using 1×11{\mkern-2.0mu\times\mkern-2.0mu}1 2D convolution before applying the costly 3×33{\mkern-2.0mu\times\mkern-2.0mu}3 2D spatial convolution. ShuffleNet models cross-channel correlation by channel shuffling instead of 1×11{\mkern-2.0mu\times\mkern-2.0mu}1 2D convolution. ResNeXt proposes grouped convolutions, while Inception replaces the fixed-size 2D spatial kernels into multi-scale 2D spatial kernels of different sizes.

In this work, we propose the decomposition of spatiotemporal convolutions into depthwise-separable temporal convolutions, which we show to be better suited for long-range temporal modeling that 2+1D convolutions. Moreover, to account for the differences in temporal extents, we propose temporal convolutions with multi-scale kernels.

Method

i. Subspace Modularity. In the context of deep network cascades, a decomposition should be modular, such that between subspaces, it retains the nature of the respective subspaces across subsequent layers. Namely, after a cascade of spatial and a temporal convolutions, it must be possible that yet another cascade (of spatial and temporal convolutions) is possible and meaningful.

ii. Subspace Balance. A decomposition should make sure that a balance is retained between the subspaces and their parameterization in different layers. Namely, increasing the number of parameters for modeling a specific subspace should come at the expense of reducing the number of parameters of another subspace. A typical example is conventional 2D CNN, in which the spatial subspace (S\mathcal{S}) is reduced while the semantic channel subspace (C\mathcal{C}) is expanded.

iii. Subspace Efficiency. When designing the decomposition for a specific task, we should make sure that the bulk of the available parameter budget is dedicated to subspaces that are directly relevant to the task at hand. For instance, for long-range temporal modeling, a logical choice is a decomposition that increases the convolutional parameters for the temporal subspace (T\mathcal{T}).

Motivated by the aforementioned design principles, we propose a new temporal convolution layer for encoding long-range patterns in complex actions, named Timeception, see figure 2. First, we discuss the Timeception layer. Then we describe how to stack Timeception layers on top of existing 2D or 3D CNNs.

2 Timeception Layer

For modeling complex actions in long videos, our temporal modeling layer faces two objectives. First, we would like to learn the possible long-range temporal dependencies between one-actions throughout the entire video, and for a frame sequence of up to 1000 timesteps. Second, we would like to tolerate the variations in the temporal extents of one-actions throughout the video.

Next, we present the Timeception layer, designed with these two objectives in mind. Timeception is a layer that sits on top of either previous Timeception layers, or a CNN. The CNN can be either purely spatial; processing frames independently, like ResNet , or short-range spatiotemporal; processing nearby bursts of frames, like I3D .

Long-range Temporal Dependencies. There exist two design consequences for modeling long-range temporal dependencies between one-actions throughout the video. The first consequence is that our temporal network must be composed of deeper stacks of temporal layers. Via successive layers, thereafter, complex and abstract spatiotemporal patterns can emerge, even when they reside at temporally very distant locations in the video. Given that we need deeper temporal stacks and we have a specific parameter budget for the complete model, the second consequence is that the temporal layers must be as cost-effective as possible.

The simplified temporal-only kernel has some interesting properties. Each kernel acts on only one channel. As the kernels do not extend to the channel subspace, they are encouraged to learn generic and abstract, rather than semantically-specific, temporal combinations. For instance, the kernels learn to detect the temporal pattern of one latent concept represented by one channel. Last, as the parameter complexity of a single Timeception layer is approximately O(T+logL)\mathcal{O}(T+logL), it is computationally feasible to train a deep model to encode temporal patterns of up to 1024 timesteps. This amounts to about 40 seconds of video sequences.

Unfortunately, by stacking temporal-only convolutions one after the other, we violate the first design principle of subspace modularity. The reason is that the semantic subspace in long-range spatiotemporal patterns is ignored. To this end, we propose to use channel grouping operation before the temporal-only convolutions and channel shuffling operation after the temporal-only convolutions. The purpose of channel grouping is reducing the complexity of cross-channel correlations, by modeling it separately for each group. Clearly, as each group contains a random subset of channels, not all possible correlations are accounted for. This is mitigated by channel shuffling and channel concatenation, which makes sure that the channels are grouped altogether albeit in a different order. As such, the next Timeception layer will group a different subset of channels. Together, channel grouping and channel shuffling is more cost-effective operation to learn cross-channel correlations than 1×11{\mkern-2.0mu\times\mkern-2.0mu}1 2D convolutions .

Tolerating Variant Temporal Extents. The second objective for the Timeception layer is to tolerate the differences in temporal extents of complex actions. While in the previous description we assume a fixed length for the temporal-only kernels, one-actions in a complex video may vary in length. To this end, we propose to replace fixed-size temporal kernels with multi-scale temporal kernels. There are two possible ways to implement multi-scale kernels, see figure 3. The first way, inspired by Inception for images, is to adopt KK kernels, each of a different size kk. The second way, inspired by , is to employ dilated convolutions.

3 The Final Model

The final model consists of four Timeception layers stacked on top of the last convolution layer of a CNN, used as backbone. We explore two backbone choices: a spatial 2D CNN and a short-range spatiotemporal 3D CNN.

Implementation. When training the model on a specific dataset, first we pretrain the backbone CNN on this dataset. We use uniformly sampled frames for the 2D backbone and uniformly sampled video segments (each has 8 successive frames) for the 3D backbone. After pre-training, we plug-in Timeception and MLP layers on top of the last convolution layer of the backbone and fine-tune the model on the same dataset. At this stage, only Timeception layers are trained, while the backbone CNN is frozen. The model is trained with batch-size 32 for 100 epoch. It is optimized with SGD with 0.10.1, 0.90.9 and 1e1e-55 as learning rate, momentum and weight decay, respectively. Our public implementation uses TensorFlow and Keras .

Experiments

The scope of this paper is complex actions with their three properties: composition, temporal extent and temporal order –see figure 1. Thus, we choose to conduct our experiments on Charades , Breakfast Actions and MultiTHUMOS . Other infamous datasets for action recognition do not meet the properties of complex actions.

Charades is multi-label, action classification, video dataset with 157 classes. It contains 8k, 1.2k and 2k videos for training, validation and test splits, respectively (67 hrs for training split). On average, each complex action (i.e. each video) is 30 seconds and contains 6 one-actions. Thus, Charades meets the criteria of complex actions. We use mean Average Precision (mAP) for evaluation. As labels of test set are held out, we report results on the validation set, similar to all related works .

Breakfast Actions is a dataset for unscripted cooking-oriented human activities. It contains 1712 videos in total, 1357 for training and 335 for test. The average length of videos is 2.3 minutes. It is a video classification task of 12 categories of breakfast activities, where each video represents only one activity. Besides, each video has temporal annotation of one-actions composing its activity. In total, there are 48 classes of one-actions. In our experiments, we only use the activity annotation, and we do not use the temporal annotation of the one-actions.

MultiTHUMOS is a dataset for human activities in untrimmed videos, with the primary focus on temporal localization. It contains 65 action classes and 400 videos (30 hrs). Each video can be thought of a complex action, which comprises 11 one-actions on average. MultiTHUMOS extends the original THUMOS-14 by providing multi-label annotation for the videos in validation and test splits. Having multiple and dense labels for the video frames enable temporal models to benefit from the temporal relations between one-actions across the video. Similar to Charades, mAP is used for evaluation.

2 Tolerating Temporal Extents

In this experiment, we evaluate the capacity of the multi-scale kernels to tolerate the differences in temporal extents of actions. The experiment is carried out on Charades.

Original v.s. Altered Temporal Extents First, we train two baselines, one with multi-scale temporal kernels (as in Timeception) and the other with fixed-size kernels. The training is done on the original temporal extent of training videos. Then, at test time only, we alter the temporal extents of test videos. Specifically, we split each test video into segments. Then, we temporally expand or shrink these segments. Expansion is done by repeating frames, while shrinking is done by dropping frames. We use 4 types of alterations with varying granularity to test the model in different scenarios: (a) very-coarse, (b) coarse, (c) fine, and (d) very-fine, see in figure 4.

The results of this controlled experiment are shown in table 1. We observe that Timeception is more effective than fixed-size kernels in handling unexpected variations in temporal extents. The same observations is confirmed using either I3D or ResNet as backbone architecture.

Fixed-size vs. Multi-scale Temporal Kernels This experiment points out the merit of using multi-scale temporal kernels. For this, we compare fixed-size temporal convolutions against multi-scale temporal-only convolutions, either with different kernel sizes kk or dilation rates dd. And we train 3 baseline models with different configurations of k,dk,d: i. Fixed kernel size and fixed dilation rate d=1,k=3d=1,k=3. This is the typical configuration used in 3D CNNs . ii. Different kernel sizes k∈{1,3,5,7}k\in\{1,3,5,7\} and fixed dilation rate d=1d=1. iii. Fixed kernel size k=3k=3 and different dilation rates d∈{1,2,3}d\in\{1,2,3\}.

The result of this experiment are shown in table 2. We observe that using multi-scale kernels is better suited for modeling complex actions than fixed-size kernels. The same observation holds for both I3D and ResNet as backbones. Also, we observe little to no change in performance when using different dilation rates dd instead of different kernel sizes kk.

3 Long-range Temporal Dependencies

In this experiment, we demonstrate the capacity of multiple Timeception layers to learn long-range temporal dependencies for complex actions. We train several baseline models equipped with Timeception layers. These baselines use different number of input timesteps. We experiment on Charades, with both ResNet and I3D as backbone.

ResNet is used, with a different number of timesteps as inputs: T∈{32,64,128}T\in\{32,64,128\}, followed by Timeception layers. ResNet processes one frame at a time. Hence, in one feedforward pass, the number of timesteps consumed by Timeception layers is equal to that consumed by ResNet.

I3D is considered, with a different number of timesteps as inputs: T∈{256,512,1024}T\in\{256,512,1024\}, followed by Timeception layers. I3D processes 88 frames into one super-frame at a time. Thus, Timeception layers model T′∈{32,64,128}T^{\prime}\in\{32,64,128\} super-frames, Practically however, as each super-frame is related to a segment of 88 frames, both I3D+Timeception process in total T∈{256,512,1024}T\in\{256,512,1024\} frames.

We report results in table 3 and we make two observations. First, stacking Timeception layers leads to an improved accuracy when using both ResNet and I3D as backbone. As the only change between these models is the number of Timeception layers, we deduce that the Timeception layers have succeeded in learning temporal abstractions. Second, despite stacking more and more Timeception layers, the number of parameters is controlled. Interestingly, using 4 Timeception layers on I3D processing 1024 timesteps requires half the parameters needed for a ResNet processing 128 timesteps. The reason is the number of channels from ResNet is twice as much as from I3D (2048 v.s. 1024). We conclude that Timeception layers allow for deep and efficient models, able to learn long-range temporal abstractions, which is crucial for complex actions.

Learned Weights of Timeception. Figure 5 visualizes the learned weights by our model. Specifically, three Timeception layers trained on top of I3D backbone. The figure depicts the weights of multi-scale temporal convolutions with different kernel sizes k∈{3,5,7}k\in\{3,5,7\}. For simplicity, only the first 30 kernels from each kernel-size, are shown. We make two remarks for these learned weights. First, at layer 1, we notice that long kernels (k=7k=7) captures fine-grained temporal dependencies, because of the rapid transition of kernel weights. But at layer 3, these long kernels tend to focus on coarse-grained temporal correlations, because of the smooth transition between kernel weights. The same behavior prevails for the short (k=3k=3) and medium (k=5k=5) kernels. Second, at layer 3, we observe that long-range and short-range temporal patterns are learned by short kernels (k=3k=3) and long kernels (k=7k=7), respectively. The conclusion is that for complex actions, both video-wide and local temporal reasoning, even at the top layer, is crucial for recognition.

4 Effectiveness of Timeception

To demonstrate the effectiveness of Timeception, we compare it against related temporal convolution layers: i. separable temporal convolution , that models both T,C\mathcal{T},\mathcal{C} simultaneously. ii. grouped separable temporal convolution to model T\mathcal{T}, followed by 1×11{\mkern-2.0mu\times\mkern-2.0mu}1 2D convolution to model C\mathcal{C}. iii. grouped separable temporal convolution to model T\mathcal{T}, followed by channel shuffling to model C\mathcal{C}. Interestingly in figure 6 top, Timeception is very efficient in maintaining a reasonable increase in number of parameters as the network goes deeper. Also, figure 6 bottom shows how Timeception improves mAP on Charades, scales up temporal capacity of backbone CNNs while maintaining the overall model size.

5 Experiments on Benchmarks

Charades is used to evaluate our model, and to compare against related works. In this experiment, our baseline networks use 4 Timeception layers. The number of convolutional groups is 8 for I3D and 16 for ResNet, be it 2D or 3D. The results in table 4 shows that Timeception monotonically improves the performance of the backbone CNN. The absolute gain on top of ResNet and I3D is 8.8%8.8\% and 4.3%4.3\%, respectively.

Beyond the overall mAP, how beneficial is Timeception? And in what cases exactly does it help? To answer this question, we make two comparisons to assess the relative performance of Timeception. We experiment two scenarios: i. short-range (3232 timesteps) v.s. long-range (128128 timesteps), ii. fixed-scale v.s. multi-scale kernels. The results are shown in figures 7, 8, and we make two observations.

First, when comparing the relative performance of multi-scale vs. fixed-size Timeception, see figure 7, we observe that multi-scale Timeception excels in complex actions with dynamic temporal patterns. As an example, “take clothes + tidy clothes + put clothes”, one actor may take longer than others to tidy clothes. In contrast, fixed-size Timeception excels in the cases where the complex action is more rigorous in the temporal pattern, e.g. “open window + close window”. Second, when comparing the relative performance of short-range (3232 timesteps) v.s. long-range (10241024 timesteps) Timeception, see figure 8, the later excels in complex actions than requires the entire video to unfold, e.g. “fix door + close door”. However, short-range Timeception would do better in one-actions, like “open box + close box” or “turn on light + turn of light”.

Breakfast Actions is used as a second dataset to experiment our model. The average length of a video in this datataset is 2.3 sec. For this experiment, we use 3 layers of Timeception. And as for the backbone, we use I3D and 3D ResNet-50. None of the backbones is fine-tuned on this dataset, only Timeception layers are trained. To make one video consumable by our baseline, from each video we uniformly sample 64 video snippet, each of 8 sucessive frames. That makes the total timesteps modeled by the baseline is 512. Finally, we report result in table 5.

MultiTHUMOS is used as a third dataset to experiment our model. This helps in investigating the generality on different datasets. Related works use this dataset for temporal localization of one-actions in each video of complex action. Differently, we use this dataset to serve our objective: multi-label classification of complex actions, i.e. the entire video. As such, the evaluation method used is mAP . To assess the performance of our model, we compare against I3D as a baseline. As shown in the results in table 6, Timeception, equipped with multi-scale kernels outperforms that with fixed-size kernel.

Conclusion

Complex actions such as “cooking a meal” or “cleaning the house” can only be recognized when processed fully. This is in contrast to one-actions, that can be recognized from a small burst of frames. This paper presents Timeception, a novel temporal convolution layer for complex action recognition. Thanks to using efficient temporal-only convolutions, Timeception can scale up to minute-long temporal modeling. In addition, thanks to multi-scale temporal convolutions, Timeception can tolerate the changes in temporal extents of complex actions. Interestingly, when visualizing the temporal weights we observe that earlier timeception layers learn fast temporal changes, whereas later timeception layers focus on more global temporal transitions. Evaluating on popular benchmarks, the proposed Timeception improves the state-of-the-art notably.

We thank Xiaolong Wang for sharing the code.

References