Kernelized Covariance for Action Recognition

Jacopo Cavazza, Andrea Zunino, Marco San Biagio, Vittorio Murino

I Introduction

In the past three decades, motion capture systems – MoCap – have been engineered with the ultimate goal of tracking and recording human motion while guaranteeing high resolutions in both spatial and temporal domains. The acquired data consist of time series of joint/marker 3D positions and are broadly used for several different applications, e.g., studying human motions in sport sciences, inferring biometric patterns for person identification or generating realistic motion sequences in computer animation to name a few . Among these ones, action and activity recognition displays a crucial role in human-robot interaction, autonomous driving vehicles and video-surveillance . However, devising effective methods to analyze MoCap data is demanding due to the many yet unsolved problems related, for instance, to missing acquisitions of joints coordinates or to highly corrupted data.

Previous attempts to face these issues either rely on some distance learning techniques (e.g., subspace view invariant metric ) or applied stochastic techniques to model the degree of uncertainty in the data. For instance, a hidden Markov model is used in to produce weak classifiers which are enhanced by AdaBoost. Furthermore, proposed an action graph to model the dynamics for action recognition and exploited a bag of 3D points as feature representation.

Since the spatial and/or temporal dimensions of the recorded data can be heavy, dimensionality reduction or feature selection methods have been devised. However, in general, the classification is subsequent to a design phase of discriminative features such as actionlets , random occupancy patterns , pose-based sets , space-time trajectories , velocity and acceleration , normal vectors or Lie group geometry embedding .

As a different paradigm to a customized class of task-specific features, generalizable representations driven by covariance matrix were shown to be promising, either encoding spatio-temporal derivatives of joint positions or producing a hierarchical temporal pyramids of descriptors .

Recently, the new state of the art for action and activity recognition from MoCap data was set by , where several Gram matrices are computed to produce multiple representations of the joint positions of each trial and, once a fusion step is performed, a log-Euclidean kernel feeds the SVM classifier. Therein, the covariance is replaced by kernel matrices and this is motivated by the observation that the former can only understand linear relationships while the latter allows to model general ones. In this work, we pursue an opposite perspective, focusing on the covariance representation and rigorously devising a kernelized version to extend its discriminative power.

Indeed, by the direct usage of a kernel, we can avoid any preliminary explicit feature encoding (as, for instance, occurs in ) and, for a general class of kernel functions, we recover the kernel trick for covariance matrix estimation. As a result, its descriptiveness increases from linear to arbitrary relationships modelling, while the efficiency in the computation is preserved.

To the best of our knowledge, this problem was never faced before in this principled way in both machine learning and pattern recognition fields.

To sum up, we highlight the contributions of this paper.

We propose a new kernelized representation for covariance matrix, namely Kernelized-COV. By recovering the well-known kernel trick, we can capture more general interdependencies between variables in a way that the usual covariance descriptor becomes a particular case and the overall computational cost does not increase.

In order to prove the effectiveness of our approach for action and activity recognition of MoCap data, we compare our method against different ones on MSR-Action3D , MSR-Daily-Activity , MSRC-Kinect12 and HDM-05 benchmark datasets. With respect to the state-of-the-art methods , the registered performance shows comparable results in the first two datasets and better scores in the remaining ones. This properly certifies that our kernelization is able to bridge the gap between covariance and kernel-based representation.

The rest of the paper is outlined as follows. In Section II, we sketch some theoretical background about the covariance matrix. In Section III, we present our framework which is experimentally validated in Section IV. Finally, Section V draws some conclusions and profiles future work.

II Background

where X\mathbf{X} represents the 3n×T3n\times T data matrix which stacks by columns all the temporal acquisitions x(1),…,x(T),\mathbf{x}(1),\dots,\mathbf{x}(T), whose average is denoted by μ.\boldsymbol{\mu}. In matrix notation, (1) becomesFor a matter of space, the technical proof of deriving equation (2) from (1) was moved to the Supplementary Material.

once defined P\mathbf{P} as the T×TT\times T matrix whose (s,t)(s,t)-th entry is

This latter direction actually grounds on the mathematical properties of positive definite matrices, exploiting Riemannian metrics on manifold for image classification: once moved from a finite to an infinite dimensional space, the performance enhances and only recently deep learning approaches have shown to be superior. However, one of the main limitation related to covariance matrix is that it only enables to capture linear inter-relationships . For instance, principal component analysis actually exploits a covariance matrix to remove linear correlation of data points . Among the attempts for modeling more complicated relationships, additional statistics, such as entropy and mutual information , and kernels have been adopted. As a different paradigm, one can model non-linear behaviors by preliminary applying a preprocessing step and encode raw data by means of a transformation which increases the feature space. For instance, applied such idea for spatial and temporal derivatives for gesture recognition, considered both different color spaces and edge detectors for image classification, and used filter bank responses as features to estimate head orientation. In this latter approach, once defined the feature map Φ\Phi and the transformed data matrix Φ(X)\boldsymbol{\Phi}(\mathbf{X}) whose tt-th column is Φ(x(t))\Phi(\mathbf{x}(t)), the covariance (2) is now expressed by

III Method

In (6), once exploited the assumption that Φ(hj)=ej,\Phi(\mathbf{h}_{j})=\mathbf{e}_{j}, for some hj,\mathbf{h}_{j}, we can define the dim⁡(H)×T\dim(\mathcal{H})\times T matrix K[X,h]\mathbf{K}[\mathbf{X},\mathbf{h}] whose (i,s)(i,s)-th entry k(x(s),hi)k(\mathbf{x}(s),\mathbf{h}_{i}) is ⟨Φ(x(s)),ei⟩H=⟨Φ(x(s)),Φ(hi)⟩H\langle\Phi(\mathbf{x}(s)),\mathbf{e}_{i}\rangle_{\mathcal{H}}=\langle\Phi(\mathbf{x}(s)),\Phi(\mathbf{h}_{i})\rangle_{\mathcal{H}} and consequently we deduce

Lemma 1 certifies that we are able to compute the covariance in terms of the sole kernel k.k. However, some issues pertain to the practical feasibility of the assumption

for any jj, which is nevertheless fundamental for our purposes.

Let ω=[ω1,…,ω3n]\boldsymbol{\omega}=[\omega_{1},\dots,\omega_{3n}] a collection of 3n3n independent samples jointly distributed as a mixture of discrete Dirac’s deltas and define ψ(x)=⟨ω,x⟩.\psi(\mathbf{x})=\langle\boldsymbol{\omega},\mathbf{x}\rangle. Then, the expectation of ψ(x)ψ(z)\psi(\mathbf{x})\psi(\mathbf{z}) under the distribution of ω\boldsymbol{\omega} is

where δij\delta_{ij} denotes the Kronecker symbol. ∎

Since we proved that Ψ\boldsymbol{\Psi} approximates the kernel kk in the sense explained above, the final stage is solving the issue related to (8).

The map Ψ\boldsymbol{\Psi} satisfies the assumption (8), that is, for every i = 1,…,M1,\dots,M, it results

The relationship (12) displays a system of equations, stochastically dependent on the randomness of Ψ\boldsymbol{\Psi}. Actually, in our case, it is enough to solve the system (12) and prove the existence of h1,…,hM\mathbf{h}_{1},\dots,\mathbf{h}_{M} under a specific realization of NN and ω\boldsymbol{\omega}, the two sources of randomness in Ψ\boldsymbol{\Psi}. In other words, we can solve (12) in a maximum likelihood sense by considering the samples of NN and ω\boldsymbol{\omega} which verify (12) with probability 11. Thus, we use a prior on NN so that N=1N=1 and, once absorbed into hi\mathbf{h}_{i} all the multiplicative constant defining Ψ\boldsymbol{\Psi}, then (12) becomes

The theoretical discussion leads to derive Algorithm 1 and to apply the proposed kernelized covariance for the task of action and activity recognition. For a better understanding, we also visualize such pipeline in Figure 1.

Computational cost. The complexity of our trial-specific kernelized covariance is O(M2T2)O(M^{2}T^{2}). Thus, differently from previous approaches , the proposed framework is very efficient if compared to the cubic complexity of methods like which require eigen-decomposition. Under a mathematical point of view, our kernelized covariance is a natural generalization of the classical covariance matrix, which can be retrieved as a particular case in our paradigm once fixed the kernel function (9) to be a linear one. On the other hand, the computational cost still remains the same if compared with the classical covariance descriptor.

IV Experimental results

In this section, we present the experimental results obtained with our Kernelized-COV method on different publicly available MoCap datasets for action recognition. Precisely, the following algorithms were compared in our experiments: Region-COV (covariance region descriptor), temporal pyramid of covariance descriptors (Hierarchy of COVs) and, finally, an infinite covariance operator which exploits Bregman divergence, namely COV-JHJ_{\mathcal{H}}-SVM . Furthermore, we also report the comparison against the recent state-of-the-art methods, namely Ker-RP-POL and Ker-RP-RBF .

In all the experiments, we only used the 3D skeleton coordinates available in the following datasets:

MSR-Action3D , where there are 20 classes of mostly sport-related action (e.g., jogging or tennis-serve) involving 10 subjects. Since each subject performs each action 2 or 3 times, the overall number of trials is 567. For each of them, Kinect sensor is used to acquire depth maps, from which 20 joints are extracted to model the human pose of any of the human agents.

MSR-Daily-Activity , captured by using a Kinect device and it is composed by 16 different classes related to every-day actions such as read book or lie down on sofa. All of them are performed by 10 subjects. The main difficulty of this dataset originates from the fact that any activity class is performed in an either standing/sitting position, with a consequent misleading motion pattern to mess up the classification.

MSRC-Kinect12 , consisting of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system. 594 sequences of approximate total length of six hours and 40 minutes are collected from 30 people performing 12 gestures: in total, 6,244 gesture instances. The motion files contain Kinect estimated trajectories of 20 joints.

HDM-05 , containing more than tree hours of systematically recorded and well-documented MoCap data using a 240Hz VICON system to acquire the gestures of 5 non-professional actors via 31 markers. Motion clips have been manually cut out and annotated into roughly 100 different motion classes: on average, 10-50 realizations per class are available.

In all cases, we used the same splits adopted in : for MSR-Action3D, MSR-Daily-Activity and MSRC-Kinect12, training is performed on odd-index subject, while the even-index ones are left for testing (cross-subject pipeline of ), while, in HDM-05, the training split exploits all the data from the “bd” and “mm” subjects and testing is performed on “bk”, “dg” and “tr”.

Furthermore, for the HDM-05 dataset we removed some severely corrupted samples and, as performed by , selected only the following classes: clap above head, deposit floor, elbow to knee, grab high, hop both legs, jog, kick forward, lie down floor, rotate both arms backward, sit down chair, sneak, squat, stand up lie and throw basketball. All the data are pre-processed in a common way. In particular, in MSR-Action3D and MSR-Daily-Activity, we computed the velocity and acceleration from the raw positions of the joints adopting either first and second order finite different scheme respectively as in .

Table I shows the results of Kernelized-COV on the four different datasets in comparison with all the other methods. Therein, in the case of MSR-Action3D and MSR-Daily-Activity, our proposed method is able to achieve comparable results with a small deviation from the state-of-the-art , but it outperforms all the other competitors. More impressively, on MSRC-Kinect12, Kernelized-COV improves the-state-of-the-art by 2.7%2.7\%. Even in the last dataset, namely HDM-05, the accuracy of the proposed method is 1.3%1.3\% higher of the best score achieved by the other competitors. In this case, referring to , we did not report the accuracy on HDM-05 due to the different experimental settings: Hierarchy of COVs scored 95.41%95.41\% on a simplified 1111-class problem, while, in the same conditions, we scored 98.8%98.8\%. Furthermore, it is worth noting that, on all the considered datasets our Kernelized-COV works even better than a recent infinite covariance operator , more discriminatively encoding the data.

The improvements in classification accuracies demonstrate the effectiveness of Kernelized-COV. Moreover, our proposed principled way of encoding non-linearities conveyed by the data is always superior to classical covariance based methods such as and does not suffer the gap in performance showed by covariance representation in .

As a final remark, it is interesting to compare the performance of our Kernelized-COV with other not covariance-based methods. To this aim, we take into account the MSR-Action3D dataset and we compared with many previous approaches in the literature, already introduced in Section I. From this analysis, the results presented in Table II give a further evidence of the effectiveness of the proposed use of the kernelized covariance, which is able to overcome , the best score reported, by a margin of 3.1%.

V Conclusions & Future Perspectives

This paper presents a principled mathematical paradigm to recover the applicability of kernel trick for covariance matrix, in order to better model more general class of relationships other than the linear ones. This enhances the descriptiveness of the classical covariance matrix which is retrievable as a particular case of our general theoretical framework. Experimentally, Kernelized-COV closes the gap between covariance and kernel-based representations in many action recognition datasets, namely MSR-Action3D, MSR-Daily-Activity, MSRC-Kinect12 and HDM-05. The proposed method is able to improve the previous best accuracies, setting the new state-of-the-art performance on the last two datasets.

As a future work, we either tackle the applicability of this novel framework to other classification problems and we will also investigate how a similar pipeline can be extended to more general classes of kernel functions.

References