IGFormer: Interaction Graph Transformer for Skeleton-based Human Interaction Recognition

Yunsheng Pang, Qiuhong Ke, Hossein Rahmani, James Bailey, Jun Liu

Introduction

Human interaction recognition plays a significant role in a wide range of applications . For example, it can be used in visual surveillance to detect dangerous events such as “kicking” and “punching”. It can also be used for robot controlling for human-robot interaction. This paper addresses human interaction recognition from skeleton sequences . Compared with RGB videos, skeleton sequences provide only 3D coordinates of human joints, which are more robust to unconventional and variable conditions, such as unusual viewpoints and cluttered backgrounds.

Compared with single-person action recognition, one additional crucial cue in recognizing a human interaction is the interactive body parts of the interactive persons. For example, the interactive hands of two persons are critical in understanding a “shaking hands” interaction. Generally, the interactive body parts in interactions demonstrate semantic correlations and correspondence. For example, in the interaction of “Taking a photo” shown in Fig. 1 (a), the hands holding the camera of one person and the hands with “yeah” of the other person demonstrate a strong correlation. Similarly, in “Shaking hands” shown in Fig. 1 (b), the interactive hands of the two persons correspond to each other. In these cases, exploring the semantic correlation between the interactive body parts is crucial for interaction understanding. In addition, for some interactions, the interactive body parts demonstrate distance evolution. For example, the hands of the two persons gradually approach each other when the two persons are “shaking hands”. Measuring the distance between body parts of the interactive persons can provide additional useful information to the semantic correlation for better interaction recognition.

Inspired by the above observation and the successful application of Transformer in many fields , we propose a novel Transformer-based model named Interaction Graph Transformer (IGFormer) for interaction recognition from skeleton sequences. In particular, the proposed IGFormer consists of a Graph Interaction Multi-head Self-Attention (GI-MSA) module, which aims at modeling the relationship of interactive persons from both semantic and distance levels to recognize actions. More specifically, the GI-MSA module learns a semantic-based graph and a distance-based interaction graph to represent the mutual relationship between body parts of the interactive persons. The semantic-based graph is learned by the attention mechanism in a data-driven manner to capture the semantic correlations of the interactive body parts. The distance-based graph is constructed by measuring the distance between pairs of body parts to excavate the distance information between interactive body parts. The two interaction graphs are combined to complement each other in a refinement way, making the model suitable for modeling different interactions.

To feed skeleton sequences to the IGFormer, one straightforward solution is to transform each skeleton sequence to a pseudo-image and divide the image into a sequence of patches, similar to the manner of ViT . However, this may destroy the spatial relationship among the skeleton joints in each body part, which could hinder effective modeling of the interactive body parts for interaction recognition. To tackle this problem, we propose a Semantic Partition Module (SPM) to transform the skeleton sequence of each subject into a new format, i.e., a Body-Part-Time (BPT) sequence, each of which is the representation of one body part during a short period. The BPT sequence encodes semantic information and temporal dynamics of the body parts, enhancing the capability of the network for modeling interactive body parts for interaction recognition.

We summarize the contributions of this paper as follows:

We introduce a Transformer-based model named IGFormer, which contains a novel GI-MSA module to learn the relationships of the interactive persons from both semantic and distance levels for skeleton-based human interaction recognition.

We introduce a Semantic Partition Module (SPM) transforming each skeleton sequence into a BPT sequence to enhance the modeling of interactive body parts.

We conduct extensive experiments on three challenging datasets and achieve state-of-the-art performance.

Related Work

Conventional deep learning-based methods model the human skeleton as a sequence of joint-coordinate vectors or a pseudo-image , which is then fed into RNNs or CNNs to predict the actions. However, representing the skeleton data as a vector sequence or a 2D grid cannot fully express the dependency between correlated joints since the human skeleton is naturally structured as a graph. Recently, GCN-based methods consider the human skeleton as a graph whose vertices are joints and edges are bones and apply graph convolutional networks (GCN) on the human graph to extract correlated features. These methods achieve better performance than RNN- and CNN-based methods, and become the mainstream methods in skeleton-based action recognition. However, these methods consider each person as an independent entity and cannot effectively capture human interaction. In this work, we focus on skeleton-based human interaction recognition and propose to model the interactive relationship of persons from both semantic and distance levels.

2 Human Interaction Recognition

Human interaction recognition is a sub-field of action recognition. Compared with single-person action recognition, human interaction methods should not only be able to model the behavior of each individual but also capture the interaction between them. Yun et al. evaluated several geometric relational body-pose features including joint features, plane features and velocity features for interaction modeling, and found out that joint features outperform others, whereas velocity features are sensitive to noise. Ji et al. built poselets by grouping joints that belong to the same body part of each individual to describe the interaction of each body part. Recently, Perez et al. proposed a two-stream LSTM-based interaction relation network called LSTM-IRN to model the intra relations of body joints from the same person and the inter relations of the joints from different persons. However, LSTM-IRN ignores the distance evolution of body parts, which is considered as an important prior knowledge for human interaction recognition. Different from the above-mentioned methods, we model the interaction relationship of interactive humans as two interaction graphs, which are constructed from the semantic and distance levels respectively to capture the semantic correlation and distance evolution between body parts.

3 Visual Transformer

Transformer was first proposed in for machine translation task and since then has been widely adopted in various natural language processing (NLP) tasks. Inspired by the successful application in NLP, Transformer has been applied to the computer vision and demonstrated its scalability and effectiveness in many vision tasks. Vision Transformer (ViT) was the first pure Transformer architecture for image recognition and obtained better performance and generalization than traditional convolutional neural networks (CNNs). After that, Transformer-based models with carefully designed and complicated architectures have been applied to various downstream vision tasks, such as object detection , semantic segmentation and video classification . In skeleton-based action recognition, Plizzari et al. proposed ST-TR to model the dependencies between joints by substituting the graph convolution operator with the self-attention operator. Different from ST-TR, we focus on human interaction modeling and propose a novel self-attention-based GI-MSA module to model the correlations between body parts of interactive persons.

Interaction Graph Transformer

One important cue in recognizing human interaction is the interactive body parts. In this section, we introduce an Interaction Graph Transformer (IGFormer), which contains a Graph Interaction Multi-head Self-Attention (GI-MSA) module to model the interactive body parts at both semantic and distance levels for skeleton-based interaction recognition. The proposed IGFormer is also equipped with a Semantic Partition Module (SPM), which aims at retaining the semantic and temporal information of each body part within the input skeleton sequences for better learning of the interactive body parts.

More specifically, each ITB contains three components including two shared-weight self-encoding (SE) modules, the Graph Interaction Multi-head Self-Attention (GI-MSA) module, and two Feed-Forward Networks (FFN). Each SE module is a standard one-layer Transformer , which aims at modeling the interaction among the body parts within each individual skeleton. The two outputs of the SE are fed into the GI-MSA to model the interactive body parts and generate an enhanced representation for each interactive person. Finally, each output of the GI-MSA is fed to a Layer Normalization (LN) followed by a FFN. We add an addition operation between the output of GI-MSA and FFNs to improve the representation capability of the model. The ITB can be formulated as follows:

where Hme\textbf{H}_{me} and Hne\textbf{H}_{ne} denote the outputs of the SE, H^me\hat{\textbf{H}}_{me} and H^ne\hat{\textbf{H}}_{ne} denote the outputs of the GI-MSA module, and H^mo\hat{\textbf{H}}_{mo} and H^no\hat{\textbf{H}}_{no} are the outputs of the ITB.

The two SE modules in the first ITB take the Body-Part-Time (BPT) representations of two interactive subjects, i.e, Hm\textbf{H}_{m} and Hn\textbf{H}_{n}, as input. The inputs of the SE in the following ITB are the outputs of the previous ITB. In the following subsections, we introduce the proposed SPM and GI-MSA in detail.

Different from natural 2D images that can be directly divided into a sequence of patches to feed to the Transformer , human skeleton sequences are represented as a set of 3D joints. Transforming the 3D skeleton sequences to 2D pseudo-images and passing them through a vision Transformer such as ViT may result in losing the temporal dependency between frames as well as the correlation between joints. To better retain both spatial and temporal information of the skeleton sequences, we propose SPM to transform the skeleton sequence of each subject into a sequence of BPT. Each element in the BPT is the representation of one body part during a short temporal period. The overall architecture of the proposed SPM is shown in Fig. 3. There are three main steps in the SPM, i.e., partitioning, resizing, and projection, which are explained below.

Projection. The projection operation aims to transform the resized body parts of each person into a BPT sequence to feed to the Transformer. Specifically, we apply a 2D convolution with kernel size of P×PP\times P on Sm,p\textbf{S}_{m,p} and Sn,p\textbf{S}_{n,p} to generate 2D feature maps, respectively. The size of each output feature map is L×DL\times D, where L=⌈(T+2×padding−P+1)/stride⌉L=\left\lceil(T+2\times padding-P+1)/stride\right\rceil and DD denotes the number of output channels. “paddingpadding” and “stridestride” denote the padding size and the stride of the convolutional filter. Each 2D feature map can then be split into a sequence of LL steps, where each step is a feature vector of dimension DD. The projection can be formulated as follows:

2 Graph Interaction Multi-head Self-Attention

To accurately recognize human interaction, one critical cue is the interactive body parts. Considering the semantic correspondence and the distance characteristics that may exist in the interactive body parts, we propose a Graph Interaction Multi-head Self-attention (GI-MSA) module to model the interactive body parts as two interaction graphs as shown in Fig. 2 (b). Specifically, GI-MSA contains a Semantic-based Dense Interaction Graph (SDIG) and a Distance-based Sparse Interaction Graph (DSIG). The SDIG is learned by exploring the semantic correlations of the interactive body parts in a data-driven manner while the DSIG is constructed based on the prior knowledge that the physically close body parts of the interactive persons are generally interactive body parts and should be connected. With the SDIG and DSIG, the proposed GI-MSA models the interaction relationships of humans from both semantic and distance spaces to capture critical interactive information. Finally, the representation of each individual is enhanced by aggregating interactive features from the other person.

2.2 Distance-based Sparse Interaction Graph

2.3 Interaction-based Feature Generation

Given the semantic- and distance-based interaction graphs, we aggregate the interactive information of the graphs with the individual features of the interactive persons to generate an enhanced representation for better interaction recognition as shown in Fig. 2 (b). Specifically, we first transform the input individual representation Hne\textbf{H}_{ne}, which is the output of the SE module for person nn, into the value features HneV\textbf{H}_{ne}^{V}:

where α\alpha is a trainable scalar to adjust the intensity of each graph enabling the network to be adaptively adjustable between distance evolution and semantic correlation of body parts. Similarly, H^ne\hat{\textbf{H}}_{ne} can be obtained in the same way.

We define the above steps of generating H^me\hat{\textbf{H}}_{me} and H^ne\hat{\textbf{H}}_{ne} from Hme\textbf{H}_{me} and Hne\textbf{H}_{ne} as Graph Interaction Self-Attention (GI-SA), which is formulated as:

Finally, GI-MSA is defined by considering hh attention “heads”, i.e., hh self-attention functions are applied to the input in parallel. Each head provides a sequence of size M×dM\times d, where d=D/hd=D/h. The outputs of the hh self-attention functions are concatenated to form an M×DM\times D sequence to be fed to the a Layer Normalization (LN) followed by a FFN. The GI-MSA can be formulated as:

Experiments

The proposed IGFormer is evaluated on three benchmark datasets, i.e., SBU , NTU-RGB+D and NTU-RGB+D120 , and is compared with state-of-the-art RNN-, CNN- and GCN-based human action and interaction recognition methods, including Co-LSTM , ST-LSTM , GCA-LSTM , 2s-GCA , FSNET , VA-LSTM , LSTM-IRN , ST-GCN , AS-GCN and CTR-GCN . Furthermore, to demonstrate the improvement of the proposed IGFormer over the standard Transformer model, we design a Transformer-based baseline named ViT-baseline, which is a ViT-base model taking the pseudo-image representation of the skeleton sequence as input.

SBU is a two-person interaction dataset, which contains eight classes of human interactions including approaching, departing, pushing, kicking, punching, exchanging objects, hugging, and shaking hands. Seven participants (pairing up to 21 different permutations) performed all eight interactions. In total, the dataset contains 282 short videos. Each video contains 3D coordinates of 15 joints per person at each frame. Following , we use the 5-fold cross validation protocol to evaluate our method.

NTU-RGB+D is a large-scale action dataset containing 56,578 skeleton sequences from 60 action classes. Each action is captured by 3 cameras at the same height but from different horizontal angles. Each human skeleton contains 3D coordinates of 25 body joints. There are two standard evaluation protocols for this dataset including 1) Cross-Subject, where half of the subjects are used for training and the remaining ones are used for testing, and 2) Cross-View, where two cameras are used for training, and the third one is used for testing. This dataset contains 11 human interaction classes including punch/slap, pat on the back, giving something, walking towards, kicking, point finger, touch pocket, walking apart, pushing, hugging and handshaking. The maximum number of frames in each sample is 256.

NTU-RGB+D120 extends NTU-RGB+D with an additional 57,367 samples from 60 extra action classes. In total, it contains 113,945 skeleton sequences from 120 action classes. There are two standard evaluation protocols for this dataset including 1) Cross-Subject, where half of the subjects are employed for training and the rest are left for testing, 2) Cross-Setup, where half of the setups are used for training, and the remaining ones are used for testing. In addition to the 11 interaction classes in the NTU-RGB+D, This dataset contains 15 additional human interaction classes including hit with object, wield knife, knock over, grab stuff, shoot with gun, step on foot, high-five, cheers and drink, carry object, take a photo, follow, whisper, exchange things, support somebody and rock-paperscissors, resulting a total of 26 interaction classes. In both NTU-RGB+D120 and NTU-RGB+D datasets, for samples with less than 256 frames, we repeat the sample until it reaches 256 frames.

2 Implementation Details

Transformer Architecture. We use a variant of ViT-Base as the backbone of our proposed IGFormer model. The original ViT-base model contains 12 Transformer layers with the hidden size of 768 (D=768). The dimension of each MLP layer is four times the hidden size. However, due to the small number of samples in the human interaction recognition datasets, a lighter model is more suitable to avoid overfitting. Therefore, we reduce the number of Transformer layers to 3 (N=3) and initialize them with the pre-trained weights of the first three layers of the ViT-base model. We also remove the classification token (CLS) and adopt the average pooling operation to obtain the final representation from each sequence of patches. We set the patch size PP in the Resizing step of SPM to 16 and the stride of convolution in the Projection step to 10, which results in BPT sequences with M=125 for each person in all datasets. In each body part, L equals to 25. kk in Eq. (9) is set to 15.

Training Details. The experiments are conducted on NVIDIA P100 GPU. We adopt SGD algorithm with Nesterov momentum of 0.9 as the optimizer. The initial learning rate is set to 0.01 and is divided by 10 at the 30th30^{th} and 40th40^{th} epochs. The training process is terminated at the 60th60^{th} epoch, batch size is 32.

3 Ablation Study

In this section, we conduct extensive ablation studies on both NTU RGB+D and NTU RGB+D 120 datasets to validate the effectiveness of the proposed SPM (Section 3.1) and GI-MSA (Section 3.2) modules.

Impacts of SPM. We compare two different representations of the skeleton sequences as the input of the proposed IGFormer to validate the effectiveness of the proposed SPM. The first one is Pseudo-Image representation, which have been widely used in CNN-based models by transforming each 3D skeleton sequence to a 2D pseudo-image. We define the numbers of frames TT and joints JJ of a skeleton sequence as the width and height of the image and then perform a linear projection on the image as ViT . The second representation is the BPT sequence, which is generated by the proposed SPM. Moreover, skeletons are transformed into the different lengths by changing the stride of convolution projection in ViT and SPM to validate the robustness of the proposed SPM under different input configurations. The experimental results are shown in Table 1. We observe that the BPT representation outperforms Pseudo-Image representation at all three configurations, which validates the effectiveness of the proposed SPM. We also evaluate a baseline that models each skeleton joint as a token of the Transformer sequence and fuses features of two persons, but the performance drops by 2.2% compared with our SPM on X-Sub of NTU-RGB+D.

GI-MSA versus Input/Late fusion. We design two interaction learning baselines, i.e., Input Fusion and Late Fusion, to compare with our proposed GI-MSA module. The Input Fusion baseline merges the BPT sequences of two subjects to form a single sequence and passes it through a standard Transformer to learn the interactions between two subjects. The Late Fusion baseline feeds the BPT sequences of two subjects individually through a Transformer model to extract their representations, which are then fused to model the interaction. As shown in Table 2, we observe that the performance of both input fusion and late fusion methods are worse than our proposed IGFormer on both datasets, demonstrating the efficacy of the proposed GI-MSA module for interactive learning.

Impacts of SDIG and DSIG. We evaluate the impacts of different components of the proposed GI-MSA, including SDIG, DSIG, the spatial and temporal context for learning SDIG. Here, we employ IGFormer without GI-MSA module as our baseline. Based on the results in Table 3, we draw three conclusions: (1) Both spatial and temporal context in Eq. (5) are important for learning key contextual features, i.e., the performance drops significantly by removing any of them. (2) The GI-MSA containing only SDIG can improve the performance of human interaction recognition, which validates the effectiveness of the proposed SDIG. (3) The DSIG, which serves as the prior knowledge of human interaction, does not perform well individually but provides extra information for interaction learning, leading to improved performance after being combined with SDIG.

Impacts of Number of ITB layers. Our IGFormer is built by stacking several Interaction Transformer Blocks (ITBs) to enhance the capability of interaction modeling. Here, we evaluate the influence of different number of ITBs on the performance of IGFormer. As shown in Table 4, stacking 3 layers of ITB achieves the best results on both NTU-RGB+D and NTU-RGB+D 120. Increasing the number of ITBs degrades the accuracy due to over-fitting problem.

Impacts of the joint noise on human interaction. The skeletons in NTU-RGB+D are usually noisy, e.g., some joints are missing. We evaluate the performance of our IGFormer on X-Sub of NTU-RGB+D by adding zero-mean noise to the skeleton sequences. IGFormer achieves 93.6%, 93.1%, 92.0%, 90.4% accuracy when the standard deviation (σ\sigma) is set to 0cm, 1 cm, 2cm, 4cm, respectively, which demonstrates that IGFormer is robust against the input noise.

4 Comparison with State-of-the-arts

The experimental results on the interaction classes of SBU, NTU-RGB+D and NTU-RGB+D 120 datasets are shown in Table 5. The proposed IGFormer achieves state-of-the-art performance compared with other skeleton-based human interaction recognition methods. Benefiting from the proposed SPM and GI-MSA modules, IGFormer outperforms the CNN- and RNN-based methods by a large margin. IGFormer also outperforms state-of-the-art GCN-based method, CTR-GCN , by 2.0%2.0\% and 2.2%2.2\% on X-Sub and X-View of NTU-RGB+D, and 2.2%2.2\% and 2.1%2.1\% on X-Sub and X-set of NTU-RGB+D 120. Compared with the baseline Transformer-based method, ViT-baseline, our IGFormer achieves 3.4%3.4\% and 3.2%3.2\% gains on X-Sub and X-View of NTU-RGB+D, and 3.9%3.9\% and 4.0%4.0\% gains on X-Sub and X-Set of NTU-RGB+D 120.

Conclusion

In this work, we presented IGFomer, which consists of a GI-MSA module to model the interaction of persons as graphs. The GI-MSA learns an SDIG and DSIG to capture the semantic and distance correlations between body parts of interactive persons. We also presented a SPM to transform each human skeleton into a BPT sequence for retaining interactive information of body parts. The proposed IGFormer outperformed state-of-the-art methods on three datasets.

Acknowledgement

The research is partially supported by University of Melbourne Early Career Researcher Grant (No: 2022ECR008). This research is also partially supported by TAILOR, a project funded by EU Horizon 2020 research and innovation programme under GA No 952215. This work is also partially supported by National Research Foundation, Singapore under its AI Singapore Programme (AISG Award No: AISG-100E-2020-065), SUTD Startup Research Grant and MOE Tier 1 Grant.

References