Actional-Structural Graph Convolutional Networks for Skeleton-based Action Recognition

Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, Qi Tian

Introduction

Human action recognition, broadly applicable to video surveillance, human-machine interaction, and virtual reality 6126548; 1032808; Sudha2017Approaches, has recently attracted much attention in computer vision. Skeleton data, representing dynamic 3D joint positions, have been shown to be effective in action representation, robust against sensor noise, and efficient in computation and storage ijcai_ChaoLi; Yan2017. The skeleton data are usually obtained by either locating 2D or 3D spatial coordinates of joints with depth sensors or using pose estimation algorithms based on videos Cao_2017_CVPR.

The earliest attempts of skeleton action recognition often encode all the body joint positions in each frame to a feature vector for pattern learning Vemulapalli_2014_CVPR; Fernando_2015_CVPR; Du_2015_CVPR; vis_cnn. These models rarely explore the internal dependencies between body joints, resulting to miss abundant actional information. To capture joint dependencies, recent methods construct a skeleton graph whose vertices are joints and edges are bones, and apply graph convolutional networks (GCN) to extract correlated features kipf_iclr2017. The spatio-temporal GCN (ST-GCN) is further developed to simultaneously learn spatial and temporal features AAAI1817135. ST-GCN though extracts the features of joints directly connected via bones, structurally distant joints, which may cover key patterns of actions, are largely ignored. For example, while walking, hands and feet are strongly correlated. While ST-GCN tries to aggregate wider-range features with hierarchical GCNs, node features might be weaken during long diffusion AAAI1816098.

We here attempt to capture richer dependencies among joints by constructing generalized skeleton graphs. In particular, we data-driven infer the actional links (A-links) to capture the latent dependencies between any joints. Similar to pmlr-v80-kipf18a, an A-link inference module (AIM) with an encoder-decoder structure is proposed. We also extend the skeleton graphs to represent higher order relationships as the structural links (S-links). Based on the generalized graphs with the A-links and S-links, we propose an actional-structural graph convolution to capture spatial features. We further propose the actional-structural graph convolution network (AS-GCN), which stacks multiple of actional-structural graph convolutions and temporal convolutions. As a backbone network, AS-GCN adapts various tasks. Here we consider action recognition as the main task and future pose prediction as the side one. The prediction head promotes self-supervision and improve recognition by preserving detailed features. Figure 1 presents the characteristics of the AS-GCN model, where we learn the actional links and extend the structural links for action recognition. The feature responses present that we could capture more global joint information than ST-GCN, which only uses the skeleton graph to model the local relations.

To verify the effectiveness of the proposed AS-GCN, we conduct extensive experiments on two distinct large-scale data sets: NTU-RGB+D Shahroudy_2016_CVPR and Kinetics DBLP:journals/corr/KayCSZHVVGBNSZ17. The experiments have demonstrated that AS-GCN outperforms the state-of-the-art approaches in action recognition. Besides, AS-GCN accurately predicts future frames, showing that sufficient detailed information is captured. The main contributions in this paper are summarized as follows:

We propose the A-link inference module (AIM) to infer actional links which capture action-specific latent dependencies. The actional links are combined with structural links as generalized skeleton graphs; see Figure 1;

We propose the actional-structural graph convolution network (AS-GCN) to extract useful spatial and temporal information based on the multiple graphs; see Figure 2;

We introduce an additional future pose prediction head to predict future poses, which also improves the recognition performance by capturing more detailed action patterns;

The AS-GCN outperforms several state-of-the-art methods on two large-scale data sets; As a side product, AS-GCN is also able to precisely predict the future poses.

Related Works

Skeleton data is widely used in action recognition. Numerous algorithms are developed based on two approaches: the hand-crafted-based and the deep-learning-based. The first approach designs algorithms to capture action patterns based on the physical intuitions, such as local occupancy features cvpr_wang, temporal joint covariances ijcai_Hussein and Lie group curves Vemulapalli_2014_CVPR. On the other hand, the deep-learning-based approach automatically learns the action faetures from data. Some recurrent-neural-network (RNN)-based models capture the temporal dependencies between consecutive frames, such as bi-RNNs Du_2015_CVPR, deep LSTMs Shahroudy_2016_CVPR; 10.1007/978-3-319-46487-9_50, and attention-based model AAAI1714437. Convolutional neural networks (CNN) also achieve remarkable results, such as residual temporal CNN a8014941, information enhancement model vis_cnn and CNN on action representations Ke_2017_CVPR. Recently, with the flexibility to exploit the body joint relations, the graph-based approach draws much attention AAAI1817135; Si_2018_ECCV. In this work, we adopt the graph-based approach for action recognition. Different from any previous method, we learn the graphs adaptively from data, which captures useful non-local information about actions.

Background

In this section, we cover the background material necessary for the rest of the paper.

2 Spatio-Temporal GCN

Actional-Structural GCN

The generalized graph, named actional-structural graph, is defined as Gg(V,Eg){\cal G}_{g}(V,E_{g}), where VV is the original set of joints and EgE_{g} is the set of generalized links. There are two types of links in EgE_{g}: structural links (S-links), explicitly derived from the body structure, and actional links (A-links), directly inferred from skeleton data. See the illustration of both types in Figure 3.

Many human actions need far-apart joints to move collaboratively, leading to non-physical dependencies among joints. To capture corresponding dependencies for various actions, we introduce actional links (A-links), which are activated by actions and might exist between arbitrary pair of joints. To automatically infer the A-links from actions, we develop a trainable A-link inference module (AIM), which consists of an encoder and a decoder. The encoder produces A-links by propagating information between joints and links iteratively to learn link features; and the decoder predict future joint positions based on the inferred A-links; see Figure 4.

We use AIM to warm-up the A-links, which are further adjusted during the training process.

Encoder. The functionality of an encoder is to estimate the states of the A-links given the 3D joint positions across time; that is,

where fv(⋅)f_{v}(\cdot) and fe(⋅)f_{e}(\cdot) are both multi-layer perceptrons, ⊕\oplus is vector concatenation, and F(⋅)\mathcal{F}(\cdot) is an operation to aggregate link features and obtain the joint feature; such as averaging and elementwise maximization. After propagating for KK times, the encoder outputs the linking probabilities as

where r\mathbf{r} is a random vector, whose elements are i.i.d. sampled from Gumbel(0,1)\rm{Gumbel(0,1)} distribution and τ\tau controls the discretization of Ai,j,:\mathcal{A}_{i,j,:}. Here we set τ=0.5\tau=0.5. We obtain the linking probabilities Ai,j,:\mathcal{A}_{i,j,:} in the approximately categorical form by Gumbel softmax jang_iclr2017.

Decoder. The functionality of the decoder to predict the future 3D joint positions conditioned on the A-links inferred by the encoder and previous poses; that is,

We pretrain AIM for a few epoches to warm-up A-links. Mathematically, the cost function of AIM is

where A:,:,c(0)\mathcal{A}^{(0)}_{:,:,c} is the prior of A\mathcal{A}. In experiments, we find the performance boosts when p(A)p(\mathcal{A}) promotes the sparsity. The intuition behind is that too many links would capture useless dependencies to confuse action pattern learning; however, in (3), we ensure that ∑c=1CAi,j,c=1\sum_{c=1}^{C}\mathcal{A}_{i,j,c}=1. Since the probability one is allocated to CC link types, it is hard to promote sparsity when CC is small. To control the sparsity level, we introduce a ghost link with a large probability, indicating that two joints are not connected through any A-link. The ghost link still ensures that the probabilities sum up to one; that is, for ∀i,j\forall i,j, Ai,j,0+∑c=1CAi,j,c=1\mathcal{A}_{i,j,0}+\sum_{c=1}^{C}\mathcal{A}_{i,j,c}=1, where Ai,j,0\mathcal{A}_{i,j,0} is the probability of isolation. Here we set the prior A:,:,0(0)=P0\mathcal{A}^{(0)}_{:,:,0}=P_{0} and A:,:,c(0)=P0/C\mathcal{A}^{(0)}_{:,:,c}=P_{0}/C for c=1,2,⋯ ,Cc=1,2,\cdots,C. In the training of AIM, we only update the probabilities of A-links Ai,j,c\mathcal{A}_{i,j,c}, where c=1,⋯ ,Cc=1,\cdots,C.

We accumulate LAIM\mathcal{L}_{\rm AIM} for multiple samples and minimize it to obtain a warmed-up A\mathcal{A}. Let Aact(c)=A:,:,c∈n×n\mathbf{A}_{\rm{act}}^{(c)}=\mathcal{A}_{:,:,c}\in^{n\times n} be the cc-th type of linking probability, which represents the topology of the cc-th actional graph. We define the actional graph convolution (AGC), which uses the A-links to capture the actional dependencies among joints. In the AGC, we use A^act(c)\hat{\mathbf{A}}^{(c)}_{\rm{act}} as the graph convolution kernel, where A^act(c)=Dact(c)−1Aact(c)\hat{\mathbf{A}}_{\rm{act}}^{(c)}={\mathbf{D}_{\rm{act}}^{(c)}}^{-1}\mathbf{A}_{\rm{act}}^{(c)}. Given the input Xin\mathbf{X}_{\rm{in}}, the AGC is

where Wact(c)\mathbf{W}_{\rm{act}}^{(c)} is the trainable weight to capture feature importance. Note that we use the AIM to warm-up A-links in the pretraining process; during the training of action recognition and pose prediction, the A-links are further optimized by forward-passing the encoder of AIM only.

2 Structural Links (S-links)

As shown in (1), A(p)~Xin\widetilde{\mathbf{A}^{(p)}}\mathbf{X}_{\rm{in}} aggregates the 11-hop neighbors’ information in skeleton graph; that is, each layer in ST-GCN only diffuse information in a local range. To obtain long-range links, we use the high-order polynomial of A{\bf A}, indicating the S-links. Here we use A^L\hat{\mathbf{A}}^{L} as the graph convolution kernel, where A^=D−1A\hat{\mathbf{A}}=\mathbf{D}^{-1}\mathbf{A} is the graph transition matrix and LL is the polynomial order. A^\hat{\mathbf{A}} introduces the degree normalization to avoid the magnitude explosion and has probabilistic intuition ilprints361; DBLP:journals/tsp/ChenTFVK18. With the LL-order polynomial, we define the structural graph convolution (SGC), which can directly reach the LL-hop neighbors to increase the receptive field. The SGC is formulated as

3 Actional-Structural Graph Convolution Block

To integrally capture the actional and structural features among arbitrary joints, we combine the AGC and SGC and develop the actional-structural graph convolution (ASGC). In (4) and (5), we obtain the joint features from AGC and SGC in each time stamp, respectively. We use a convex combination of both as the response of the ASGC. Mathematically, the ASGC operation is formulated as

where λ\lambda is a hyper-parameter, which trades off the importance between structural features and actional features. A non-linear activation function, such as ReLU(⋅)\rm{ReLU(\cdot)}, can be further introduced after ASGC.

The linearity ensures that ASGC effectively preserves information from both structural and actional aspects; for example, when the response from the action aspect is stronger, it can be effectively reflected through ASGC.

To capture the inter-frame action features, we use one layer of temporal convolution (T-CN) along the time axis, which extracts the temporal feature of each joint independently but shares the weights on each joint. Since ASGC and T-CN learns spatial and temporal features, respectively, we concatenate both layers as an actional-structural graph convolution block (AS-GCN block) to extract temporal features from various actions; see Figure 5. Note that ASGC is a single operation to extract only spatial information and the AS-GCN block includes a series of operations to extract both spatial and temporal information.

4 Multitasking of AS-GCN

Backbone network. We stack a series of AS-GCN blocks to be the backbone network, called AS-GCN; see Figure 6. After the multiple spatial and temporal feature aggregations, AS-GCN extracts high-level semantic information across time.

Action recognition head. To classify actions, we construct a recognition head following the backbone network. We apply the global averaging pooling on the joint and temporal dimensions of the feature maps output by the backbone network, and obtain the feature vector, which is finally fed into a softmax classifier to obtain the predicted class-label y^\hat{\mathbf{y}}. The loss function for action recognition is the standard cross entropy loss

where y\mathbf{y} is the ground-truth label of the action.

Future pose prediction head. Most previous works on the analysis of skeleton data focused on the classification task. Here we also consider pose prediction; that is, using AS-GCN to predict future 3D joint positions given by historical skeleton-based actions.

Joint model. In practice, when we train the recognition head and future prediction head together, recognition performance gets improved. The intuition behind is that the future prediction module promotes self-supervision and avoids overfitting in recognition.

Experiments

NTU-RGB+D. NTU-RGB+D, containing 56,88056,880 skeleton action sequences completed by one or two performers and categorized into 6060 classes, is one of the largest data sets for skeleton-based action recognition Shahroudy_2016_CVPR. It provides the 3D spatial coordinates of 25 joints for each human in an action. For evaluating the models, two protocols are recommended: Cross-Subject and Cross-View. In Cross-Subject, 40,32040,320 samples performed by 2020 subjects are separated into training set, and the rest belong to test set. Cross-View assigns data according to camera views, where training and test set have 37,92037,920 and 18,96018,960 samples, respectively.

Model Setting. We construct the backbone of AS-GCN with 9 AS-GCN blocks, where the features dimensions are 64, 128, 256 in each three blocks. The structure and operations of future pose prediction module are symmetric to the recognition module and we use the residual connection. In the AIM, we set the hidden features dimensions to be 128. The number of A-link types C=3C=3 and the prior of the ghost link P0=0.95P_{0}=0.95. λ=0.5\lambda=0.5. We use PyTorch 0.4.1 and train the model for 100 epochs on 8 GTX-1080Ti GPUs. The batch size is 32. We use the SGD algorithm to train both recognition and prediction heads of AS-GCN, whose learning rate is initially 0.1, decaying by 0.1 every 20 epochs. We use Adam optimizer Kingma_iclr2015 to train the A-link inference module with the initial learning rate 0.00050.0005. All hyper-parameters are selected using a validation set.

2 Ablation Study

To analyze each individual component of AS-GCN, we conduct extensive experiments on Cross-Subject benchmark of the NTU-RGB+D data set Shahroudy_2016_CVPR.

Effect of link types. Here we focus on validating the proposed A-links and S-links. In the experiments, we consider three link-type combinations, including S-links, A-links and AS-links (A-links ++ S-links), with the original skeleton links. While involving S-links, we respectively set the polynomial order L=1,2,3,4L=1,2,3,4 in the model. Note that when L=1L=1, the corresponding S-link is exactly the skeleton itself.

Table 1 shows the classification accuracy of action recognition. We see that (1) either S-links with higher polynomial order or A-links can improve the recognition performance; (2) when using both S-links and A-links together, we achieve the best performance; (3) with only A-links and skeleton graphs, the classification accuracy result reaches 83.2%83.2\%, which is higher than S-links with polynomial order 1 (81.5%81.5\%). These results validate the limitation of the original skeleton graph and the effectiveness of the proposed S-link and A-link.

Visualizing A-links. Various actions may activate different actional dependencies among joints. Figure 8 shows the inferred A-links for three actions. The A-links with probabilities larger than 0.9 are illustrated as orange lines, where wider lines represent larger linking probabilities.

We see that (1) in Plots (a) and (c), the actions of hand waving and taking a selfie are mainly upper limb actions, where arms have large movements and interact with the whole bodies, so that many A-links are built between arms and other body parts. (2) In Plot (b), the action of kicking something shows that the kicked leg is highly correlated to the other joints, indicating the body balancing during this action. These results validate that richer information of action patterns is captured by A-links.

The number and priors of A-links. To select the appropriate CC: the number of A-link types; and P0P_{0}: the prior of the ghost links for training the AIM. We test the models with different CC and P0P_{0} to obtain the corresponding recognition accuracies, which are presented in Table 2.

We see that when C=3C=3 and P0=0.95P_{0}=0.95, we could obtain the highest recognition accuracy. The intuition is that too few A-link types cannot capture significant actional relationships and too many causes overfitting. And the sparse A-links would promote the recognition performance.

Effect of prediction head. To analyze the effect of the prediction head on improving recognition performance, we perform two groups of contrast tests. For the first group, AS-GCNs only employs S-links for action recognition but one has prediction head and the other does not have. In the other group, AS-GCN with/without prediction head additionally employ A-links. The polynomial order of S-links is from 1 to 4.

Table 3 shows the classification results with/without prediction heads. We obtain better recognition performance consistently by around 1%1\% when we introduce the prediction head. The intuition behind is that the prediction modules promotes to preserve more detailed information and introduce self-supervision to help recognition module avoid overfitting and achieve higher action recognition performance. The sparse skeleton actions may sometimes rely on the detailed motions rather than coarse profiles which are easily-confused in some actions classes.

Feature visualization. To validate how the features of each joint effect on the final performance, we visualize the feature maps of actions in Figure 9, where the circle around each joint indicates magnitude of feature responses of this joint in the last AS-GCN block of the recognition module of AS-GCN.

Plot (a) shows the feature responses of the action ’hand waving’ at different time. At the initial phase of action, namely Frame 15, many joints of upper limb and trunk have approximately comparative responses; however, in the subsequent frames, large responses are distributed on the upper body especially waving arm. Note that other non-functional joints are not much neglected, because abundant hidden relationships are built. Plot (b) shows the other two actions, where we are able to capture many long-range dependencies. Plot (c) compares the features between AS-GCN and ST-GCN. ST-GCN does apply multi-layer GCNs to cover the entire spatial domain; however, the feature are weakened during the propagation and distant joints cannot interact effectively, leading to localized feature responses. On the other hand, the proposed AS-GCN could capture useful long-range dependencies to recognize the actions.

3 Comparisons to the State-of-the-Art

We compare AS-GCN on skeleton-based action recognition tasks with the state-of-the-art methods on the data sets of NTU-RGB+D and Kinetics. On NTU-RGB+D, we train AS-GCN on two recommended benchmarks: Cross-Subject and Cross-View, then we respectively obtain the top-1 classification accuracies in the test phase. We compare with covering hand-crafted methods Vemulapalli_2014_CVPR, RNN/CNN-based deep learning models Du_2015_CVPR; Shahroudy_2016_CVPR; 10.1007/978-3-319-46487-9_50; a8014941; vis_cnn; Ke_2017_CVPR; ijcai_ChaoLi and recent graph-based methods AAAI1817135; Tang_2018_CVPR; Si_2018_ECCV. Specifically, ST-GCN AAAI1817135 combines GCN with temporal CNN to capture spatio-temporal features, and SR-TSL Si_2018_ECCV use gated recurrent unit (GRU) to propagate messages on graphs and use LSTM to learn the temporal features. Table 4 shows the comparison. We see that the proposed AS-GCN outperforms the other methods.

In the Kinetics dataset, we compare AS-GCN with four state-of-the-art approaches. A hand-crafted based method named Feature Encoding Fernando_2015_CVPR is presented at first. Then Deep LTSM and Temporal ConvNet Shahroudy_2016_CVPR; a8014941 are implemented as two deep learning models on Kinetics skeletons. Additionally, ST-GCN is also evaluated for Kinetics action recognition. Table 5 shows the top-1 and top-5 classification performances. We see that AS-GCN outperforms the other competitive methods in both top-1 and top-5 accuracies.

4 Future Pose Prediction

We evaluate the performance of AS-GCN for future pose prediction. For each action, we take all frames except for the last ten as the input. We attempt to predict the last ten frames.

Figure 10 visualizes the original and predicted action. We sample five frames at regular intervals in ten. The predicted frame provides the future joint position with a low error, especially the characteristic actional body parts, e.g. shoulders and arms. As for the peripheral parts such as legs and feet, the predicted positions have larger error, which is the secondary information of the action pattern. These results show that AS-GCN preserves more detailed features especially for the action-functional joints.

Conclusions

We propose the actional-structural graph convolution networks (AS-GCN) for skeleton-based action recognition. The A-link inference module captures actional dependencies. We also extend the skeleton graphs to represent higher order relationships. The generalized graphs are fed to AS-GCN block for a better representation of actions. An additional future pose prediction head captures more detailed patterns through self-supervision. We validate AS-GCN in action recognition using two data sets, NTU-RGB+D and Kinetics. The AS-GCN achieves large improvement compared with the previous methods. Moreover, AS-GCN also shows promising results for future pose prediction.

Acknowledgement

We are supported by The High Technology Research and Development Program of China (2015AA015801), NSFC (61521062), and STCSM (18DZ2270700).

References

Appendix A: Theorem Proof

The actional-structural graph convolution is a valid linear operation; that is, when Y1=ASGC(X1)\mathbf{Y}_{1}={\rm ASGC}\left(\mathbf{X}_{1}\right) and Y2=ASGC(X2)\mathbf{Y}_{2}={\rm ASGC}\left(\mathbf{X}_{2}\right). Then, aY1+bY2=ASGC(aX1+bX2)a\mathbf{Y}_{1}+b\mathbf{Y}_{2}={\rm ASGC}\left(a\mathbf{X}_{1}+b\mathbf{X}_{2}\right), ∀a,b\forall a,b.

The operations in actional graph convolution (AGC) are all linear, as well as the structural graph convolution (SGC). The AGC satisfies

With both AGC and SGC operations, the actional-structural convolution (ASGC) is formulated as

which is a linear summation of AGC and SGC. Therefore, we have

The ASGC is a linear operation for the input data.

Appendix B: Model Architectures

In this section, we show the detailed architectures of the proposed AS-GCN model.

The activation functions of MLPs in the encoder are exponential linear unit (elu) functions, and ’bn’ denotes the batch normalization to the features. ⊕\oplus is the concatenation operation.

Decoder

We present the detailed configuration of the decoder of AIM. Given the position of joint viv_{i} at time tt, xit\mathbf{x}_{i}^{t}, the decoder aims to predict the future joint position xit+1\mathbf{x}_{i}^{t+1} conditioned on the sourrounding A-links, Ai,j,:\mathcal{A}_{i,j,:}. The architectures are presented in Table 7.

GRU(⋅)\rm{GRU}(\cdot) denotes a GRU unit, whose hidden feature dimension is 6464. It predicts the future position of all joints conditioned on A-links and previous frames.

Backbone

The backbone network of the AS-GCN extracts the rich spatial and temporal feature of actions with the proposed ASGC and temporal CNN (T-CN). For example, we build AS-GCN on NTU-RGB+D dataset and Cross-Subject benchmark Shahroudy_2016_CVPR. There are 25 joints, 3D spatial positions and 300 padded frames for each action. The architecture of the backbone is presented in Table 8.

There are nine ASGC blocks consisting of the backbone of AS-GCN model. The input/output feature maps are 3D tensors, where the three axes represent the joint number, feature dimension and frame number, respectively. The shapes of operations have the consistent dimensions with input and output feature maps, where the first axis is the filter number or output feature dimension, and the other three correspond to the input shape. AA and BB are the types of A-links and S-links. For action recognition, we obtain the last feature map whose shape is and apply a global average pooling operation on the time and joint axis, i.e. the 1st and 3rd axis. Thus we obtain a semantic feature vector of the action, whose dimension is 256.

Future Action Prediction Head

The architecture of the future action prediction head of AS-GCN model are presented in Table 9.

We input the output feature map from the backbone network in to the prediction head. The input tensor are calculated by nine ASGC blocks. The first five blocks reduce the frame number to aggregate higher-level action features. The last four blocks work on action regeneration. For the last four blocks, we concatenate the last input frame to each feature map. Finally, with a residual connection, we obtain a tensor with shape from a fully connected layer, which contains the joint position of the predicted 10 frames.

Appendix C: More Future Pose Predictions

More future action prediction results of different actions are illustrated in Figure 11,

which contains the action of ’wipe face’, ’throw’ and ’nausea or vomiting condition’. As we see, the actions are predicted with very low error.