Skeleton based action recognition using translation-scale invariant image mapping and multi-scale deep cnn

Bo Li, Mingyi He, Xuelian Cheng, Yucheng Chen, Yuchao Dai

I Introduction

Action recognition is an important research area in computer vision, which has a wide range of applications, e.g., human computer interaction, video surveillance, robotics, and etc. Recent years, the cost-effective depth sensor combining with real-time skeleton estimation algorithms can provide reliable joint coordinates . As an intrinsic high level representation, 3D skeleton is valuable and comprehensive for summarizing a series of human dynamics in the video, and thus benefits general action analysis . Besides its succinctness and effectiveness, it has a significant advantage of great robustness to illumination, clustered background, and camera motion. Based on these advantages, 3D skeleton based activity analysis has drawn great attentions .

Very recently, human pose estimation from 2D RGB videos have also been studied with deep CNN method . Rather accurate 2D skeleton joints could be evaluated from RGB videos. However, it seems still lack effective method to deal with this kind of 2D skeleton video data for recognition purpose.

Previously, hand-crafted skeleton features have been devised . However, these hand-crafted features are always shallow and dataset-dependent, thus limiting their performance.

Recently, deep learning based methods have achieved great success in high-level computer vision tasks such as image recognition, classification, detection, and semantic segmentation etc. As for the 3D skeleton based action recognition problem, Recurrent Neural Networks (RNNs) have been widely adopted . RNNs could effectively extract the temporal information and learn the contextural information well. However, RNNs tend to overemphasize the temporal information especially when the training data is insufficient, thus leading to over-fitting .

Convolutional Neural Networks (CNNs) have also been applied to this problem . Different from the RNNs, how to effectively represent 3D skeleton data and feed into deep CNNs is still an open problem. Wang et al. proposed the Joint Trajectory Maps (JTM), which represents both spatial configuration and dynamics of joint trajectories into three texture images through color encoding, and then fed these texture images to CNNs for classification. However, this kind of JTM is a little complicated, and may lose some important information when projecting 3D skeleton into 2D image. Du et al. proposed to represent each skeleton sequence as an image, where the temporal dynamics of the sequence are encoded as changes in columns and the spatial structure of each frame is represented as column. Their encoding method is dataset dependent and translation-scale variant which means their encoding method need dataset information and human’s translation and action scale may influence the final mapping results.

To tackle the above shortcomings in skeleton-based video recognition problem, in this paper, we present a new framework consisting of translation-scale invariant image mapping and multi-scale deep CNN classifier. The overall flowchart of our method is illustrated in Fig. 1. We propose to map the 3D skeleton video to a color image, where the color image achieves translation and scale invariance and dataset independent. The proposed mapping could easily handle the translation and scale changes in 3D skeleton data, thus is more distinctive, and dataset independent.

Although the skeleton images are very different from natural images, the widely used pre-trained deep CNN model e.g., AlexNet, VGGNet, ResNet could still be transferred to it well. It is especially valuable when there is insufficient annotated skeleton videos. The fine-tune strategy could avoid training millions of parameters afresh and improve the performance significantly . More importantly, due to the special property of this kind of skeleton images, we propose a simple yet effective multi-scale deep CNN to enhance the frequency adjustment ability of our method.

In addition, we extend our method to deal with 2D skeleton-based video recognition problem. Surprisingly, our method could also work well. Experimental results on the popular benchmark dataset like NTU RGB-D, UTD-MHAD, MSRC-12, and G3D demonstrate the effectiveness our proposed framework. We also give extensive analysis experiments to show the propoerties of our method.

In conclusion, our main contributions are summarized as following:

We propose a translation-scale invariant image mapping method for 3D skeleton video data. This mapping could avoid the translation and scale influence of the skeleton data, thus is more distinctive and dataset independent.

A multi-scale deep CNN is proposed to enhance the frequency adjustment ability of our method.

We test our method to 2D skeleton data and achieve excellent results which shows our method could also work on 2D skeleton data well.

We achieve the state-of-the-art results on the widely used benchmarks like NTU RGB-D, UTD-MHAD, MSRC-12, and G3D dataset. In addition, extensive component analysis experiments are conducted.

This paper are organized as following. Related works are summarized and presented in Sec. II. We present our method in Sec. III, including our Translation-scale invariant image mapping method, multi-scale deep CNN and data augmentation method. Experiments on the popular benchmarks are presented in Sec. V. At last, more analysis experiments and results on 2D skeleton data are presented in Sec. VI.

II Related Work

Tranditionaly, hand-crafted skeleton features have been devised to capture the spatial-tempory information. There have been great amount of hand-craft features proposed for 3D skeleton-based video recognition . Generally speaking, spatial descriptor, geometric descriptor, or key poses are extensively studied. We would like suggest the readers refer to for more summary. However, these hand-crafted features are always shallow and dataset-dependent, thus limiting their performance.

Recently, deep learning methods have been adopted on this field. Du et al. divided the human skeleton into 5 parts and a hierarchical recurrent neural network is proposed for this problem. Veeriah et al. proposed a kind of differential recurrent neural network which emphasizes on the change in information gain caused by the salient motions between the successive frames. Zhu et al. proposed a kind of regularized deep LSTM network which take the skeleton as the input at each time slot and introduce a novel regularization scheme to learn the co-occurrence features of skeleton joints. Liu et al. propose a more powerful tree-structure based traversal method. To handle the noise and occlusion in 3D skeleton data, they introduce new gating mechanism within LSTM to learn the reliability of the sequential input data.

Convolutional Neural Networks (CNNs) have also been applied to this problem . Wang et al. proposed the Joint Trajectory Maps (JTM), which represents both spatial configuration and dynamics of joint trajectories into three texture images through color encoding, and then fed these texture images to CNNs for classification. Du et al. proposed to represent each skeleton sequence as an image, where the temporal dynamics of the sequence are encoded as changes in columns and the spatial structure of each frame is represented as column. How to effectively represent 3D skeleton data and feed into deep CNNs is still an open problem.

In recent years, human pose estimation from 2D RGB videos have been studied with deep CNN method . Rather accurate 2D skeleton joints could be evaluated from RGB videos.

Our work is also related to recent works on transfer learning and deep learning. In , Krizhevsky et al. trained a large deep CNN on the ImageNet dataset and achieved a performance leap on the classification problem. Recently, more and more work show that pre-trained deep CNN features can be transferred to new classification or recognition problems and boost performance . In , a very deep CNN is proposed, which achieved state-of-the-art classification results in ImageNet challenge 2014. In , a kind of residual structure is proposed which makes the hundreds layers CNN trainable.

This paper is an extended version of our conference paper published in ICMEW , In which, we achieve the 3rd place in the “Large Scale 3D Human Activity Analysis Challenge in Depth Videos”. However, this journal version substantively improves the performance and gives more insights by more experiments and analysis.

III Method

Our framework consist of 3 parts (1) Translation-scale invariant image mapping. (2) A multi-scale deep CNN for classification. (3) Data augmentation methods we utilized for 3D skeleton data. A conceptual illustration of our framework is presented in Fig.1.

Following the work of , we divide all human skeleton joints in each frame into five main parts according to human physical structure, i.e. two arms, two legs and a trunk. To preserve the local motion characteristics, joints in each part are concatenated as a vector by their physical connections. Then the five parts are concatenated as the representation of each frame. To map the 3D skeleton video to an image, a natural and direct way is to represent the three coordinate components (x,y,z)(x,y,z) of each joint as the corresponding three components (R,G,B)(R,G,B) of each pixel in an image. Specially, each row of the action image is defined as Ri=[xi1,xi2,...,xiN]R_{i}=[x_{i1},x_{i2},...,x_{iN}], Gi=[yi1,yi2,...,yiN]G_{i}=[y_{i1},y_{i2},...,y_{iN}], Bi=[zi1,zi2,...,ziN]B_{i}=[z_{i1},z_{i2},...,z_{iN}], where ii denotes the joint index and NN indicates the number of frames in a sequence. By concatenating all joints together, we obtain the resultant action image representation of the original skeleton video.

Due to the coordinate difference in 3D skeleton and image, proper normalization is needed. Du et al. proposed to quantify the float matrix to discrete image representation with respect to the entire training dataset. Specifically, given the joint coordinate cjkc^{jk} (x,y,zx,y,z), the corresponding pixel value pjkp^{jk} is defined as

where cmaxc_{max} and cminc_{min} are the maximum and minimum of all joint coordinates in the training set respectively. jkjk represent the kk-th channel (x,y,zx,y,z) of the jj-th skeleton video sequence. floorfloor is the rounding down function. However, as the normalization is conducted with respect to the entire training dataset, the resultant action image is dataset dependent and translation-scale variant. 1. This encoding method need cmin,cmaxc_{min},c_{max} which need statistic on the dataset. Different dataset may have different cmin,cmaxc_{min},c_{max}, thus this kind of encoding method is dataset dependent. 2. The same action conducted on the different position(caused by the translation of the person or camera) will result in a different action image, thus is translation variant. 3. The same action conducted by different subjects may have some scale variant, while this normalization also could not guarantee scale invariant in encoding the skeleton video.

To tackle the above problems in encoding the skeleton video, we propose a simple and yet effective translation-scale invariant image mapping method, which is presented in Equ. 2

where cmaxjkc^{jk}_{max} and cminjkc^{jk}_{min} are the maximum and minimum coordinate value of the kk-th channel (x,y,zx,y,z) of the jj-th skeleton video sequence.

Compared with Du et al. , our image mapping method owns the following properties:

Translation invariant: Our image mapping transforms the 3D coordinates with respect to the minimum coordinate in each sequence rather than the entire training dataset, thus the translation in 3D does not affect the action video as illustrated in Fig. 9.

Scale invariant: Similarly, by normalizing the 3D skeleton coordinates with respect to the 3D variation along the three axis in each sequence, the scale change has also been eliminated. In addition, our normalization is isometric to each coordinate x,y,zx,y,z, thus the relative scale in different axis has been well preserved.

Dataset independent: As the minimum and maximum coordinates are extracted from each sequence independently, the normalization is thus independent to each specific skeleton video dataset.

III-B Action recognition through multi-scale CNN

Our overall multi-scale CNN architecture is presented in Fig. 4. Our CNN architecture could be built on the fully-CNN based pre-trained model like AlexNet, VGGNet, and ResNet .

The multi-scale structure is motivated by the frequency variant of this kind of skeleton-images. Under our skeleton mapping framework, if we fix the size of the convolution kernel, different input size will bring different frequency variance. The multi-scale (multi-frequency) input often includes rich cues to the activity recognition problem.

In order to reduce the amount of the parameters in our model, the weights of the Fully-CNN parts are shared by all the different resolution inputs. Then global pooling is performed on the correspondent feature maps which is critical to result in the same size feature vectors. In order to further regularize the training, we put on the softmax loss on the all output of different resolution as well as the average of them.

We train our network with the multinomial logistic loss:

Here, NN is the number of training samples, kk the correspondent label of sample nn, mm is the number of classes. xx is the input of the softmax layer.

III-C Data augmentation

Data augmentation has been proved as an effective way in deep CNN based image classification. In this paper, we have encoded the 3D skeleton data to RGB images. In order to augment the dataset and leap the classification performance, we have specially designed different data augmentation strategies, such as 3D coordinate random rotation , Gaussian noise, video crop etc.

The augmentation methods we utilized including:

3D coordinate rotation: 3D coordinate are randomly rotated in the range of [−30o,30o][-30^{o},30^{o}], along the x,y,zx,y,z axis. Some examples are presented in Fig. 5.

Gaussian noise: We randomly add gaussian noise in the 3D coordinate with the θ=0,σ=0.01\theta=0,\sigma=0.01.

Video crop: We randomly crop the videos by the range of [0.7,1][0.7,1] in the random locations of the video .

IV Implementation details

Before proceeding to the experimental results, we give implementation details for our method. The proposed network is trained by using stochastic gradient decent momentum of 0.9, and weight decay of 0.0004. Weights are initialized by the pre-trained model from AlexNet, VGGNet, and ResNet . The network is trained with fixed learning rate 0.001 in the first 8 epoches, then divided by 10 every 5 epoches.

Our implementation is based on the efficient CNN toolbox: caffe with an NVIDIA Tesla Titian X GPU.

V Experiments

In this section, we evaluate the proposed method on public benchmark datasets: the large NTU RGB+D Dataset, UTD-MHAD, MSRC-12 Kinect Gesture Dataset, and G3D. The final recognition results were compared with the state-of theart reporteds on the same datasets.

To the best of our knowledge, NTU RGB-D dataset is the largest action recognition dataset. We adopt the same train-test protocol as in .

The dataset has more than 56 thousands sequences and 4 million frames, containing 60 actions performed by 40 subjects aging between 10 and 35. It consists of front view, two side views and one left, right 45 degree views. This dataset is very challenging due to the large intra-class and viewpoint variations.

For a fair comparison and evaluation, the same protocol as that in was used. It has both cross-subject and cross view evaluation. In the cross-subject evaluation, samples of subjects 1,2,4,5,8,9,13,14,15,16,17,18,19,25,27,28,31,34,35 and 38 were used as training samples and the remaining subjects were reserved for testing. In the cross-view evaluation, samples taken by camera 2 and 3 were used as training, while testing set includes the samples of camera 1. We further augment the training dataset by 2 times.

In the dataset of NTU, there are some samples consisted of more than 1 person. We choose the simplest strategy to deal with this kind of situation, that, we just concat these two person’s coordinate and present them in one image. An sample is presented in Fig. 6.

In Table I, we report the performance comparison between our method and the state-of-the-art methods. Clearly, our proposed method achieves the best performance in both cross-subject and cross-view evaluation. Our method outperforms the current state-of-the-art methods with a margin of 12% in cross-subject evaluation and 11% in cross-view evaluation.

V-B UTD-MHAD

UTD-MHAD is a multimodal action dataset, captured by one Microsoft Kinect camera and one wearable inertial sensor. This dataset contains 27 actions performed by 8 subjects (4 females and 4 males) with each subject performing each action 4 times. After removing three corrupted sequences, the dataset has 861 sequences. The actions are: “right arm swipe to the left”, “right arm swipe to the right”, “right hand wave”, “two hand front clap”, “right arm throw”, “cross arms in the chest”, “basketball shoot”, “right hand draw x”, “right hand draw circle (clockwise)”, “right hand draw circle (counter clockwise)”, “draw triangle”, “bowling (right hand)”, “front boxing”, “baseball swing from right”, “tennis right hand forehand swing”, “arm curl (two arms)”, “tennis serve”, “two hand push”, “right hand know on door”, “right hand catch an object”, “right hand pick up and throw”, “jogging in place”,“walking in place”, “sit to stand”, “stand to sit”, “forward lunge (left foot forward)” and “squat (two arms stretch out)”.It covers sport actions (e.g. “bowling”, “tennis serve” and “baseball swing”), hand gestures (e.g. “draw X”, “draw triangle”, and “draw circle”), daily activities (e.g. “knock on door”, “sit to stand” and “stand to sit”) and training exercises(e.g. “arm curl”, “lung” and “squat”). For this dataset, cross-subjects protocol was adopted as in , namely, the data from the subjects numbered 1, 3, 5, 7 were used for training while subjects 2, 4, 6, 8 were used for testing. Table VIII compares the performance of the proposed method and those reported in .

V-C MSRC-12 Kinect Gesture Dataset

MSRC-12 is a relatively large dataset for gesture/action recognition from 3D skeleton data captured by a Kinect sensor. The dataset has 594 sequences, containing 12 gestures by 30 subjects, 6244 gesture instances in total. The 12 gestures are: “lift outstretched arms”, “duck”, “push right”, “goggles”,“wind it up”, “shoot”, “bow”, “throw”, “had enough”, “beat both”, ”change weapon” and “kick”. For this dataset, cross-subjects protocol was adopted, that is, odd subjects were used for training and even subjects were for testing.

V-D G3D Dataset

Gaming 3D Dataset (G3D) focuses on real-time action recognition in a gaming scenario. It contains 10 subjects performing 20 gaming actions: “punch right”, “punch left”, “kick right”, “kick left”, “defend”, “golf swing”, “tennis swing forehand”, “tennis swing backhand”, “tennis serve”, “throw bowling ball”, “aim and fire gun”, “walk”, “run”, “jump”, “climb”, “crouch”, “steer a car”, “wave”, “flap” and “clap”. For this dataset, the first 4 subjects were used for training, the fifth for validation and the remaining 5 subjects were for testing as configured in . Table VII compared the performance of the proposed method and those reported in .

VI More analysis

To demonstrate the effectiveness of our translation-scale invariant image mapping and our data augmentation method, we compared with other skeleton data image mapping method and the results are presented in Table.V, which clearly demonstrates the effectiveness of our image mapping method. For a fair comparison, all the encoded action images are fine-tuned on the Alexnet . and the result of Wang et al. is quoted from directly, in order to avoid the hyper-parameter setting influence. It is worth noting that the results of Wang etal utilized data augmentation. For fair comparison, we also give the results of our method with data augmentation.

It is clear from the Tab V, our mapping method outperform in all the dataset with a clearly margin, which prove our translation-scale property is important. In addtion, our method outperform in most dataset except G3D. More importantly, our advantage is obvious.

VI-B Effect of our multi-scale architecture

We compared the performance of different pre-trained CNN net in Table VI. It is obvious that our multi-scale network structure also improves the performance significately.

VI-C Adapted to 2D skeleton

As for 2D skeleton, we set the missing coordinate to 0. The correspondent results are presented in Tab. VII. Surprisingly, the 2D skeleton is just a little worse than the 3D skeleton. We would like to argu that, to our best knowledge, we are the first to conduct this kind of experiment. This promissing results show that our method could also work well on the 2D skeleton data.

In addition, we also give the result of 1D skeleton data in Tab. VII. The performance of 1D skeleton data decrease sharply.

VII Conclusion and Future Work

In this paper, we present a skeleton based action recognition method by using both translation-scale invariant image mapping and multi-scale deep CNNs. Experiments on the large scale challenging NTU RGB-D, UTD-MHAD MSRC-12 dataset show that our method outperforms the state-of-the-art methods by a large margin. In addition, we extend our method to 2D skeleton-based video recogntion problem and it performs well.

Acknowledgements

This work was supported in part by Natural Science Foundation of China grants (61420106007, 61671387) and Australian Research Council grants (DE140100180).

References