Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action Recognition

Jungho Lee, Minhyeok Lee, Dogyoon Lee, Sangyoun Lee

Introduction

Human action recognition (HAR) is a task that categorizes action classes by receiving video data as input. HAR is used in many applications, such as human–computer interaction and virtual reality. Recently, several RGB-based and skeleton-based HAR methods have been proposed with the development of deep learning technology. However, RGB-based methods cannot robustly recognize human actions because they are strongly influenced by environmental noises such as background color, brightness of light, and clothing. Therefore, methods using skeleton modality have attracted attention because they are not affected by these noises. These methods recognize action by receiving 2D or 3D coordinates of major human joints as time-series inputs.

Recent approaches have adopted graph convolutional networks (GCNs) to apply human-skeleton graphs to convolutional layers. However, existing GCN-based methods have the following limitations. (1) With the widely used handcrafted graph, the relationships between distant joint nodes are not identified since they use only the relationships of PC edges in the human skeleton. Although the graph with PC edges has a semantic significance, the graph with only PC edges suffers from long-range dependency problem as they are heuristically fixed. However, for humans to recognize actions, relationships between structurally distant joints as well as between adjacent joints are strongly correlated. Although several methods have attempted to solve such limitation by training attention-guided learnable graphs, they still use ’s handcrafted graph with their learnable graphs. Moreover, as the element values of ’s graph are more dominant than those of the learnable graphs, they do not adequately highlight the relationships between distant nodes. (2) Some recent methods risk falling into suboptimality by simply aggregating the edge features and ignoring the contribution of each edge, thus incompletely recognizing which edges are significant for each skeleton sample. For example, in the case of a ‘squat down’ action, the interactions between the legs and arms should be highlighted.

Motivated by these limitations, we propose a hierarchically decomposed graph convolutional network (HD-GCN) with a hierarchically decomposed graph (HD-Graph) and attention-guided hierarchy aggregation (A-HA) module. In addition, we present a six-way ensemble method to effectively utilize our HD-Graph. The framework of our proposed methods is shown in Fig. 1 for ‘squat down’ action.

The HD-GCN incorporates GCNs with our HD-Graph, which identifies the relationships between distant joint nodes in the same semantic spaces (e.g.\textit{e}.\textit{g}. right and left hands, right and left feet). The same semantic spaces are formed by moving out step by step from the Center of Mass (CoM) node of the graph. For example, if belly is a CoM node, the first semantic space includes the belly node, the next space includes the chest and hip nodes, and the subsequent space includes the left and right shoulder and the left and right hip nodes. The nodes in the same semantic space are defined as hierarchy node set. To detect the relationships between distant joint nodes, network should have large receptive field. The proposed HD-Graph contains both meaningful adjacent and distant joint nodes by connecting all the nodes in neighboring hierarchy node sets and identifies the connectivity between those nodes for large receptive field. We adopt rooted tree-like structure to effectively represent every edges. We apply a spatial edge convolution (S-EdgeConv) layer to consider semantically close edges which cannot be captured by the HD-Graph for each sample. To create the S-EdgeConv layer, we borrow the structure of , which is widely used in 3D point clouds.

To consider the contribution of each edge set, the process of selecting the dominant hierarchical information should depend on the action data sample to give proper attention to the most dominant edge sets. For example, in order to recognize the “clapping” action, a hierarchy edge set that includes both hands must be emphasized. To tackle this issue, we propose an attention-guided hierarchy aggregation (A-HA) module, which consists of two submodules: representative spatial average pooling (RSAP) and hierarchical edge convolution (H-EdgeConv). A scaling bias problem occurs if we use the spatial average pooling module without any node extraction process because each node has a different number of adjacent nodes. To prevent this, we apply RSAP, which includes a representative node extraction process that triggers features after the pooling layer to represent each node. To effectively manage hierarchical features obtained by RSAP, we apply a hierarchical edge convolution (H-EdgeConv) layer. The H-EdgeConv treats each hierarchical feature as a graph node and identifies which hierarchical features should be highlighted via the Euclidean distance in feature space. With the RSAP and the H-EdgeConv, our model successfully determines which hierarchy edge sets and joints should be emphasized among the input features.

The existing ensemble method uses four-stream data composed of the joint, bone, joint motion, and bone motion streams, which are the original skeletal coordinates, spatial differential between joint coordinates, and temporal differential of joint, and temporal differential of the bone, respectively. Most existing ensemble methods use additional motion data, but models that solely utilize motion data exhibit relatively inferior performance. To address this problem, we present a new method, a six-way ensemble. We apply this ensemble method by setting three HD-Graphs with joint and bone stream data. Each graph has different CoM nodes to extract features of different semantic spaces (see Appendix).

We conduct extensive experiments on four benchmark action recognition datasets: NTU-RGB+D 60 , NTU-RGB+D 120 , Kinetics-Skeleton , and Northwestern-UCLA .

Our main contributions are summarized as follows:

We propose a hierarchically decomposed graph (HD-Graph) to thoroughly identify the significant distant edges between the same hierarchy node sets.

We propose an attention-guided hierarchy aggregation (A-HA) module to highlight the key edge sets with representative spatial average pooling (RSAP) and hierarchical edge convolution (H-EdgeConv).

We use a new six-way ensemble method for skeleton-based action recognition with HD-Graphs that have different center of mass (CoM), which outperforms regular ensemble without any motion data.

Our HD-GCN outperforms the state-of-the-arts on four benchmarks for skeleton-based action recognition.

Related Work

In skeleton-based action recognition, human skeletal data are represented by a graph with joint nodes. Most recent approaches use GCN-based methods with ’s graph structure, which identifies physical connections in the human skeleton. Those GCN-based methods perform remarkably better than methods using handcrafted features . They extract the spatial features representing the relationships between physically connected edges among human skeleton, and they outperform other methods by using them to construct the major relationships between joint nodes in the human skeleton. In particular, and propose adaptive attention-based graph structures to learn the sample-wise topological features. However, they might fall into suboptimality because they do not consider the physical prior of the human skeletal structure and allow too much flexibility in network training. To address this issue, we introduce a novel HD-Graph, referencing the known tendencies of human perception

2 Attention Modules for Action Recognition

The attention mechanism is an essential element for constructing a deep neural network. Using recent attention modules , networks emphasize important information along a specific dimension. For example, Hu et al. applies channel-wise attention, and Woo et al. applies both channel-wise and spatial-wise attentions. These techniques are divided into two categories for GCNs: (1) attention-based graph construction which is a method of forming topologies using a non-local block or customized correlation matrices, and (2) spatial-wise, temporal-wise, channel-wise attention, which are commonly used attentions in , and several other networks.

Methodology

In Sec. 3.2, we detail the HD-Graph convolution to solve the problems of the conventional human-skeleton graph , which includes only PC edges. We also explain the A-HA module in Sec. 3.3 to highlight dominant hierarchical features. In Sec. 3.4, we replace the widely used four-stream ensemble method with a six-way ensemble without motion data streams. Finally, we introduce the HD-GCN, which uses these proposed methods.

The spatio-temporal graph for human skeleton is represented by G(V,E)\mathcal{G(V,E)}, where V\mathcal{V} and E\mathcal{E} denote the joint and edge groups, respectively. Physically connected edges and fully-connected edges used in Sec. 3.2 are denoted as PC-edges and FC-edges, respectively.

Graph Convolutional Networks.

2 Hierarchically Decomposed Graph

Most recent methods have adopted the handcrafted graph proposed by Yan et al. , but the HD-Graph is derived through a newly presented method. Fig. 2 shows the framework of the HD-Graph.

where Ek\mathcal{E}_{k} denotes the concatenation of the three edge subsets of S={sid,scp,scf}\mathit{S}=\{s_{id},s_{cp},s_{cf}\} and sid,scp,scfs_{id},s_{cp},s_{cf} indicate the identity, centripetal, and centrifugal edge subsets, respectively. Through this construction policy, we create a skeletal graph with bidirectional and identity edges.

Fully Connected Inter-Hierarchy Edges.

Decomposed graph A↔HD\overleftrightarrow{\mathbf{A}}_{\textrm{HD}} has a different number of edge sets from the conventional graph, but the edges are all the same. To identify the relationships between major distant joint nodes, especially those in the same semantic space, we connect all nodes between neighboring hierarchy node sets. In addition, since ’s graph contains the connectivity of only PC edges, not distant relationships, the receptive field is very small with this sparse graph. Applying our fully connected (FC) edges to the rooted tree, the graph becomes denser and makes the receptive field larger than before with more meaningful distant connectivity as shown in Fig. 2 (b). Then, the adjacency matrices are normalized with degree matrices for training stability and we leave all elements of the matrices as learnable parameters for training adaptability.

HD-Graph Convolution.

Our HD-Graph convolution includes four parallel branch operations: three graph convolution through HD-Graph and an additional EdgeConv operation. To reduce the computational complexity, a linear transformation is applied to all four operations. For three of these operations, our method performs a subset-wise GCN operation in the same way as for each hierarchy edge set with three edge subsets. However, rather than summing the output values for each subset as in Eq. 1, we concatenate these output values to the channel dimension:

where zVz_{\mathcal{V}} denotes the S-EdgeConv operation.

All four branch outputs are concatenated to the channel dimension, with all four computed in the same way for NLN_{L} hierarchy edge sets. Due to the inherent characteristics of skeletal data, the number of joint nodes included in each dataset is different, and, consequently, the number of hierarchy sets is different. Therefore, we adopt an addition policy for NLN_{L} hierarchy-wise outputs and a concatenation policy for NSN_{S} subset-wise outputs. In this way, the dimensionality is maintained, and the common hierarchy-wise aggregation policy is followed for every skeletal dataset by adding all the outputs for different numbers of hierarchical sets.

3 Attention-Guided Hierarchy Aggregation

The HD-Graph convolution uses an aggregation policy of adding all the hierarchy-wise outputs. However, because each data sample has relationship between specific major edges, we propose an attention-guided hierarchy aggregation (A-HA) module, which applies a weighted-sum policy to the hierarchy dimension with proper attention to the hierarchy-wise outputs. The framework of the HD-Graph convolution with the A-HA module is shown in Fig. 3.

where NkN_{k} denotes the number of vertices in the HkH_{k} set.

Hierarchical Edge Convolution.

After the RSAP layer, NLN_{L} hierarchy-wise features in attention feature map M\mathbf{M} have not yet shared their information with each other. We treat all NLN_{L} features as nodes on a graph to learn and reflect similarities in the hierarchical feature space. To apply this process, representative features of these nodes are fed into EdgeConv , and the similarities of those nodes are learned based on the Euclidean distance. We also include the self-loop shown in the bottom section of Fig. 3 so that the node’s own features can be reflected. Our attention map M\mathbf{M} operates as follows:

where zLz_{L} and σ\sigma denote H-EdgeConv and the sigmoid function, respectively.

The attention map M\mathbf{M} obtained is multiplied by the HD-Graph convolution output feature map FHD\mathbf{F}_{\textrm{HD}}, and the output feature map Fout\mathbf{F}_{out} is obtained through a weighted sum to the hierarchy axis as shown in Fig. 3. Similar to S-EdgeConv in Sec. 3.2, The H-EdgeConv method incorporates the concept of hierarchical edge sets in a physically proximal manner in earlier layers, while in the deeper layers, it emphasizes the presence of semantically similar edge sets. This approach highlights different hierarchical edge sets for each sample and enables the model to learn meaningful representations that capture both physical and semantic properties of the input data.

4 Six-Way Ensemble

Shi et al. have applied a four-stream ensemble method using streams for joints, bones, joint motion, and bone motion. However, as the performances of motion streams are relatively poorer than the performances of joint and bone streams, we adopt an ensemble method with the joint and bone streams without any motion streams. We use three different HD-Graphs, and each graph is used for training with joint and bone streams. The three HD-Graphs have different CoM nodes, which are chest, belly, hip nodes, respectively. In other words, we train joint and bone streams with HD-Graph with the CoM node of chest, and we train the same when the CoM node is belly or hip node. As models with the three different graphs should be trained in different aspects, each of the graphs is composed of different edge sets. For example, if the CoM node is belly, both thigh edges and both upper arm edges are included in the same edge set, whereas when the CoM node is chest, both thigh edges and both forearm edges are included in the same edge set. The details of our six-way ensemble are specified on our Appendix.

5 Network Architecture

As shown in Fig. 4, we adopt as our baseline network architecture with a total of nine stacked GCN blocks. The numbers of output channels for the blocks are 64, 64, 64, 128, 128, 128, 256, 256, and 256. Each block contains a residual connection and is divided into a spatial module, in which the GCN operation proceeds, and a temporal module, which includes the temporal convolutions. Our method use the temporal module of , whose baseline module is . This module consists of four branch operations. Two are dilated temporal convolutions with kernel size five and dilation one and two, respectively. The remaining branch operations are pointwise convolution and max pooling with kernel size three. Our spatial module consists of an HD-Graph convolution operation and an A-HA module, as introduced in Sec. 3.2 and Sec. 3.3. After passing through all GCN layers with attention to the hierarchy-wise features, the network compresses the feature map through the global average pooling layer and classifies the action sample through the softmax function.

Experiments

NTU-RGB+D 60 is a large dataset used in skeletal action recognition. It contains 56,880 skeleton action samples, performed by 40 different participants and classified into 60 classes. The authors of this dataset recommend two benchmarks. (1) Cross-Subject (X-Sub): 20 of the 40 subjects’ actions are used for training, and the remaining 20 are for validation. (2) Cross-View (X-View): Two of the three camera-views are used for training, and the other one is used for validation.

NTU-RGB+D 120.

NTU-RGB+D 120 is a dataset in which 57,367 new action samples are added to the NTU-RGB+D 60 dataset. It contains a total of 114,480 skeleton action samples over 120 classes, performed by 106 different subjects. The authors of this dataset recommend two benchmarks: (1) Cross-Subject (X-Sub): 53 of the 106 subjects’ actions are used for training, and the remaining 53 are used for validation. (2) Cross-Setup (X-Set): Of the 32 setups, data with even setup IDs are used for training, and the remaining data with odd IDs are used for validation.

Kinetics-Skeleton.

The Kinetics-Skeleton dataset is derived from the Kinetics 400 video dataset , utilizing the OpenPose pose estimation to extract 240,436 training and 19,796 testing skeleton sequences across 400 classes. The dataset restricts the number of skeletons per time step to two and eliminates skeletons with lower confidence scores, ensuring high-quality sequences for human action recognition and pose estimation research.

Northwestern-UCLA.

The Northwestern-UCLA skeleton dataset contains 1494 video clips over 10 classes. Each action is captured through three Kinect cameras with different camera views and is performed by 10 subjects. We adopt the same protocol as NW-UCLA: Two of the three camera-views are used for training, and the other one is used for validation.

Experimental Settings.

In our experiments, we adopt as the backbone. The SGD optimizer is employed with a Nesterov momentum of 0.9 and a weight decay of 0.0004. The number of learning epochs is set to 90, with a warm-up strategy applied to the first five epochs for more stable learning. We set the learning rate to decay with cosine annealing , with a maximum learning rate of 0.1 and a minimum learning rate of 0.0001. For the NTU-RGB+D datasets, we set the batch size to 64 and use the data preprocessing method from . For Kinetics-Skeleton, the batch size is set to 128. In addition, to overcome the absence of belly and hip nodes in the Kinetics-Skeleton, we define the center of both hip joints as CoM hip node, and the center of chest and the hip node as CoM belly node, resulting in a total of 20 nodes. For the Northwestern-UCLA dataset, we set the batch size to 16 and use the data preprocessing method from . All our experiments are conducted on a single RTX 3090 GPU.

2 Comparison with State-of-the-Arts Methods

Most recent state-of-the-art networks adopt a four-way ensemble method, but we adopt the six-way ensemble method described in Sec. 3.4.

We compare ours with state-of-the-art networks on three datasets: NTU-RGB+D 60 , NTU-RGB+D 120 , Northwestern-UCLA , and Kinetics-Skeleton . Comparisons for each dataset are shown in Tab. 1. The recognition performance of our HD-GCN has exceeded the state-of-the-arts on every dataset without any motion streams, as shown in Tab. 1. With our proposed ensemble method, HD-GCN outperforms the state-of-the-art and shows comparable performance to the 6-way ensemble state-of-the-art using only 4-way ensemble method.

3 Ablation Study

In this section, we demonstrate the effectiveness of the proposed HD-GCN. Performance is specified as the cross-subject and cross-setup classification accuracy on the NTU-RGB+D 120 joint stream data.

To proceed with the ablation study for HD-Graph, we set Yan et al. ’s graph as the conventional graph. Here, we use the temporal convolution module of , as mentioned in Sec. 3.5, to compare the performance of networks fairly with various graphs. The experimental results are shown in 2.

We set the edges of the HD-Graph in different ways to show a gradual performance increase according to the type of graph. There are four main versions of HD-Graph, the first of which is graph A containing only the PC edges. Unlike the conventional graph with one edge set including three fixed subsets, HD-Graph has a flexible number of edge sets, each divided by hierarchy layers with three subsets. Graph B is an extension of A, with the additional operation S-EdgeConv. Graph C contains FC edges for NHN_{H} hierarchy node sets, and graph D is similar to C but includes S-EdgeConv. The HD-Graph with only PC edges performs better than the conventional graph by a large margin, even though they share the same edges. This proves that it is meaningful to divide the joint nodes by hierarchy edge sets. In addition, the HD-Graph with FC edges and S-EdgeConv performs better for every datasets.

Attention-Guided Hierarchy Aggregation.

To prove the effectiveness of the A-HA module, we use a method to change or remove specific parts of our attention module, with the results shown in Tab. 3. Spatial average pooling (SAP) simply averages along the spatial axis without the representative node extraction process, which performs worse than our RSAP. The poorer performance is due to two factors: (1) scaling bias occurs because the number of nodes in each hierarchy node set is different, and (2) attention through SAP does not represent the corresponding hierarchy node set because it brings the average of the feature vectors of all nodes, not a specific node set. Furthermore, it performs better with H-EdgeConv, which recognizes each hierarchy edge set as a graph node. This proves that because the major edge sets are different for each data sample, it is important to find and highlight edge sets with high similarity based on the Euclidean distance through H-EdgeConv.

The results of the attention score M\mathbf{M} of our A-HA module are shown in Fig. 5. These results show that our module scores edge sets 4 and 5 higher for the “running” class, which includes knees and feet, elbows and hands. For the “Kicking” class, A-HA gives the highest score to edge set 3, which includes shoulders and hips, followed by edge set 4 and 5. It is reasonable for human visual recognition that the dynamically moving edge set 4, 5 are more important than the stationary and barely moving edge set 3 when running rather than when kicking something.

Six-Way Ensemble.

We use the ensemble method to which three graphs with different CoM nodes are applied, excluding motion streams. Tab. 1 shows that the HD-GCN with 4-way ensemble outperforms the state-of-the-art 4-way methods with motion data and shows comparable performance with the state-of-the-art 6-way method. In addition, when the 6-way ensemble with three different graphs is applied to HD-GCN, it outperforms the state-of-the-art methods. This proves that the features extracted with different CoM nodes are learned in different learning aspects.

4 Comparison of Complexity with Other Models

Although our model has multiple branch layers for multiple edge sets, it does not cause high complexity because it precedes channel reduction layers. Comparisons of computational complexity with other models are shown in Tab. 4. We conduct experiments in the same environment by fixing window size to 64. Our model shows the best performance on NTU-RGB+D 120 joint stream by a large margin even though the computational complexity of our model is the lowest. For multi-stream ensemble, our 4-stream HD-GCN shows almost similar performance to 6-stream InfoGCN while having 3.68G fewer FLOPs and 2.70M fewer parameters as shown in Tab. 5.

Conclusions

In this work, we propose a novel hierarchically decomposed graph convolutional network (HD-GCN) for skeleton-based action recognition. We also propose a new framework (HD-Graph) that replaces the existing framework, decomposes all the joint nodes by hierarchy edge sets and considers the connectivity between major distant nodes, which is difficult to identify naturally. We also present an effective attention module (A-HA) composed of representative spatial average pooling (RSAP) layer and hierarchical edge convolution (H-EdgeConv), which applies hierarchy-wise attention for the HD-Graph. In addition, our HD-GCN learns graph-wise features with different patterns through a six-way ensemble method. We derive an effective feature extractor by combining these three methods and empirically verify its effectiveness. Our approach outperforms current state-of-the-art methods on four benchmark datasets.

Appendix A Additional Details of HD-Graph

It is easier to construct our HD-Graph than existing handcrafted graph even if ours is composed of more edges than the existing one. ’s graph requires every physically adjacent edges for human joints as shown in Algorithm 1. On the other hand, our HD-Graph requires only the hierarchy-wise node sets as shown in Algorithm 2. It verifies that our HD-Graph is more universal than the existing graph in that the requirements of the HD-Graph are fewer than those of the existing one.

Tree Structures for HD-Graph.

Tree structures for skeletal modality has been already proposed in , which applies depth first search (DFS) algorithm to identify the kinematic dependency relations between the joints. It traverses every joint nodes from the root node to the leaf nodes to model the spatial dependency of the joints. Nevertheless, as ’s tree identifies only the adjacent connections of human joints, it cannot discover direct relationships between structurally distant nodes. Moreover, it is dependent to much on the fixed joint visiting order, which makes the model reflect only the topologically fixed edge features. On the other hand, our HD-Graph is free from those drawbacks. Although we also uses the tree structure to construct the HD-Graph, direct relationships of the structurally distant edges are identified by connecting every nodes for adjacent hierarchy node sets. In addition, because there is no fixed node visiting order in the process of constructing HD-Graph, HD-GCN leverages various edge features via FC-edges and adaptively highlights significant edge sets by A-HA module.

Appendix B Effectiveness of Six-way Ensemble

As we mentioned in our main paper, we propose the ensemble method with joint and bone streams without motion streams. Model with each stream is trained with three different HD-Graphs, which have different CoM nodes; chest, belly, and hip. In other words, training ways for our ensemble methods are as follows: (1) joint stream with CoM of chest node, (2) bone stream with CoM of chest node, (3) joint stream with CoM of belly node, (4) bone stream with CoM of belly node, (5) joint stream with CoM of hip node, (6) bone stream with CoM of hip node. As shown in Fig. 6, corresponding edge sets for all graphs are different from each others. We compare the performance of those three graphs for several labels on NTU-RGB+D 120 joint dataset as shown in Fig. 7. The differences between the maximum and minimum accuracy for those three graphs range from 4% to 13%. It heuristically proves that the models trained on each of the three graphs have different learning patterns.

For most recent skeleton-based action recognition models , optimal coefficients for their ensemble methods should be chosen, which have different values depending on their model. For example, suggests ensemble coefficients of [1.0, 1.0, 0.6, 0.6], which represent joint, bone, joint motion, bone motion streams. Moreover, presents [0.7, 0.7, 0.3, 0.3] and suggests [0.6, 0.6, 0.4, 0.4]. It reduces the universality of the model in that the coefficients should be manually determined. However, there is no need to set those coefficients for six-way ensemble because our method does not require motion streams. Instead of using low-performance motion streams, we use only joint and bone streams and apply the ensemble to all six models with equal contribution. In other words, our ensemble method does not require any ensemble coefficients that determine how much each stream contributes to the model. Applying our ensemble method, our HD-GCN outperforms state-of-the-art methods without the motion streams and manually fixed ensemble coefficients.

Additional Experimental Results.

Tab. 6 shows every single experimental result for our six-way ensemble method. Comparing the results of ours in Tab. 6 and other models shown in Tab. 7, it shows that our model outperforms the others even on single-stream experiments by a large margin.

Appendix C Architectures for Kinetics-Skeleton

We modify the original graph of Kinetics-Skeleton to apply our HD-Graph. The original architecture of the dataset contains 18 nodes, which does not have hip and belly nodes for CoM, so we manually set those nodes by using existing nodes. Firstly, we set the CoM hip node, which is the middle point of left and right hip nodes. In addition, the belly CoM node is the middle point of chest and hip nodes. The modified skeleton architecture contains 20 nodes due to the generated CoM hip and belly nodes. The original and modified versions of the skeleton are shown in Fig. 8.

References