Space-Time Representation of People Based on 3D Skeletal Data: A Review
Fei Han, Brian Reily, William Hoff, Hao Zhang
Introduction
Human representation in spatiotemporal space is a fundamental research problem extensively investigated in computer vision and machine intelligence over the past few decades. The objective of building human representations is to extract compact, descriptive information (i.e., features) to encode and characterize a human’s attributes from perception data (e.g., human shape, pose, and motion), when developing recognition or other human-centered reasoning systems. As an integral component of reasoning systems, approaches to construct human representations have been widely used in a variety of real-world applications, including video analysis , surveillance , robotics , human-machine interaction , augmented and virtual reality , assistive living , smart homes , education , and many others .
During recent years, human representations based on 3D perception data have been attracting an increasing amount of attention . Comparing with 2D visual data, additional depth information provides several advantages. Depth images provide geometric information of pixels that encode the external surface of the scene in 3D space. Features extracted from depth images and 3D point clouds are robust to variations of illumination, scale, and rotation . Thanks to the emergence of affordable structured-light color-depth sensing technology, such as the Microsoft Kinect and Asus Xtion PRO LIVE RGB-D cameras, it is much easier and cheaper to obtain depth data. In addition, structured-light cameras enable us to retrieve the 3D human skeletal information in real time , which used to be only possible when using expensive and complex vision systems (e.g., motion capture systems ), thereby significantly popularizing skeleton-based human representations. Moreover, the vast increase in computational power allows researchers to develop advanced computational algorithms (e.g., deep learning ) to process visual data at an acceptable speed. The advancements contribute to the boom of utilizing 3D perception data to construct reasoning systems in computer vision and machine learning communities.
Since the performance of machine learning and reasoning methods heavily relies on the design of data representation , human representations are intensively investigated to address human-centered research problems (e.g., human detection, tracking, pose estimation, and action recognition). Among a large number of human representation approaches , most of the existing 3D based methods can be broadly grouped into two categories: representations based on local features and skeleton-based representations . Methods based on local features detect points of interest in space-time dimensions, describe the patches centered at the points as features, and encode them (e.g., using bag-of-word models) into representations, which can locate salient regions and are relatively robust to partial occlusion. However, methods based on local features ignore spatial relationships among the features. These approaches are often incapable of identifying feature affiliations, and thus the methods are generally incapable to represent multiple individuals in the same scene. These methods are also computationally expensive because of the complexity of the procedures including keypoint detection, feature description, dictionary construction, etc.
On the other hand, human representations based on 3D skeleton information provide a very promising alternative. The concept of skeleton-based representation can be traced back to the early seminal research of Johansson , which demonstrated that a small number of joint positions can effectively represent human behaviors. 3D skeleton-based representations also demonstrate promising performance in real-world applications including Kinect-based gaming, as well as in computer vision research . 3D skeleton-based representations are able to model the relationship of human joints and encode the whole body configuration. They are also robust to scale and illumination changes, and can be invariant to camera view as well as human body rotation and motion speed. In addition, many skeleton-based representations can be computed at a high frame rate, which can significantly facilitate online, real-time applications. Given the advantages and previous success of 3D skeleton-based representations, we have witnessed a significant increase of new techniques to construct such representations in recent years, as demonstrated in Fig. 1, which underscores the need of this survey paper focusing on the review of 3D skeleton-based human representations.
Several survey papers were published in related research areas such as motion and activity recognition. For example, Han et al. described the Kinect sensor and its general application in computer vision and machine intelligence. Aggarwal and Xia recently published a review paper on human activity recognition from 3D visual data, which summarized five categories of representations based on 3D silhouettes, skeletal joints or body part locations, local spatio-temporal features, scene flow features, and local occupancy features. Several earlier surveys were also published to review methods to recognize human poses, motions, gestures, and activities , as well as their applications . However, none of the survey papers specifically focused on 3D human representation based on skeletal data, which was the subject of numerous research papers in the literature and continues to gain popularity in recent years.
The objective of this survey is to provide a comprehensive overview of 3D skeleton-based human representations mainly published in the computer vision and machine intelligence communities, which are built upon 3D human skeleton data that is assumed as the raw measurements directly from sensing hardware. We categorize and compare the reviewed approaches from multiple perspectives, including information modality, representation coding, structure and transition, and feature engineering methodology, and analyze the pros and cons of each category. Compared with the existing surveys, the main contributions of this review include:
To the best of our knowledge, this is the first survey dedicated to human representations based on 3D skeleton data, which fills the current void in the literature.
The survey is comprehensive and covers the most recent and advanced approaches. We review 171 3D skeleton-based human representations, including 150 papers that were published in the recent five years, thereby providing readers with the complete, state-of-the-art methods.
This paper provides an insightful categorization and analysis of the 3D skeleton-based representation construction approaches from multiple perspectives, and summarizes and compares attributes of all reviewed representations.
In addition, we provide a complete list of available benchmark datasets. Although we also provide a brief overview of human modeling methods to generate skeleton data through pose recognition and joint estimation , the purpose is to provide related background information. Skeleton construction, which is widely studied in the research fields (such as computer vision, computer graphics, human-computer interaction, and animation) is not the focus of this paper. In addition, the main application domains of interest in this survey paper is human gesture, action, and activity recognition, as most of the reviewed papers focus on these applications. Although several skeleton-based representations are also used for human re-identification , however, skeleton-based features are usually used along with other shape or texture based features (e.g., 3D point cloud) in this application, as skeleton-based features are generally incapable to represent human appearance that is critical for human re-identification.
The remainder of this review is structured as follows. Background information including 3D skeleton acquisition and construction as well as public benchmark datasets is presented in Section 2. Sections 3 to 6 discuss the categorization of 3D skeleton-based human representations from four perspectives, including information modality in Section 3, encoding in Section 4, hierarchy and transition in Section 5, and feature construction methodology in Section 6. After discussing the advantages of skeleton-based representations and pointing out future research directions in Section 7, the review paper is concluded in Section 8.
Background
The objective of building 3D skeleton-based human representations is to extract compact, discriminative descriptions to characterize a human’s attributes from 3D human skeletal information. The 3D skeleton data encodes human body as an articulated system of rigid segments connected by joints. This section discusses how 3D skeletal data can be acquired, including devices that directly provide the skeletal data and computational methods to construct the skeleton. Available benchmark datasets including 3D skeleton information are also summarized in this section.
Several commercial devices, including motion capture systems, time-of-flight sensors, and structured-light cameras, allow for direct retrieval of 3D skeleton data. The 3D skeletal kinematic human body models provided by the devices are shown in Fig. 2.
Motion capture systems identify and track markers that are attached to a human subject’s joints or body parts to obtain 3D skeleton information. There are two main categories of MoCap systems, based on either visual cameras or inertia sensors. Optical-based systems employ multiple cameras positioned around a subject to track, in 3D space, reflective markers attached to the human body. In MoCap systems based on inertial sensors, each 3-axis inertial sensor estimates the rotation of a body part with respect to a fixed point. This information is collected to obtain the skeleton data without any optical devices around a subject. Software to collect skeleton data is provided with commercial MoCap systems, such as Nexus for Vicon MoCapVicon: http://www.vicon.com/products/software/nexus., NatNet SDK for OptiTrackOptiTrack: http://www.optitrack.com/products/natnet-sdk., etc. MoCap systems, especially based on multiple cameras, can provide very accurate 3D skeleton information at a very high speed. On the other hand, such systems are typically expensive and can only be used in well controlled indoor environments.
1.2 Structured-Light Cameras
Structured-light color-depth sensors are a type of camera that uses infrared light to capture depth information about a scene, such as Microsoft Kinect v1 , ASUS Xtion PRO LIVE , and PrimeSense , among others. A structured-light sensor consists of an infrared-light source and a receiver that can detect infrared light. The light projector emits a known pattern, and the way that this pattern distorts on the scene allows the camera to decide the depth. A color camera is also available on the sensor to acquire color frames that can be registered to depth frames, thereby providing color-depth information at each pixel of a frame or 3D color point clouds. Several drivers are available to provide the access to the color-depth data acquired by the sensor, including the Microsoft Kinect SDK , the OpenNI library , and the OpenKinect library . The Kinect SDK also provides 3D human skeletal data using the method described by Shotton et.al . OpenNI uses NITE – a skeleton generation framework developed as proprietary software by PrimeSense, to generate a similar 3D human skeleton model. Markers are not necessary for structured-light sensors. They are also inexpensive and can provide 3D skeleton information in real time. On the other hand, since structured-light cameras are based on infrared light, they can only work in an indoor environment. The frame rate (30 Hz) and resolution of depth images () are also relatively low.
1.3 Time-of-Flight (ToF) Sensors
ToF sensors are able to acquire accurate depth data at a high frame rate, by emitting light and measuring the amount of time it takes for that light to return – similar in principle to established depth sensing technologies, such as radar and LiDAR. Compared to other ToF sensors, the Microsoft Kinect v2 camera offers an affordable alternative to acquire depth data using this technology. In addition, a color camera is integrated into the sensor to provide registered color data. The color-depth data can be accessed by the Kinect SDK 2.0 or the OpenKinect library (using the libfreenect2 driver) . The Kinect v2 camera provides a higher resolution of depth images () at 30 Hz. Moreover, the camera is able to provide 3D skeleton data by estimating positions of 25 human joints, with better tracking accuracy than the Kinect v1 sensor. Similar to the first version, the Kinect v2 has a working range of approximately 0.5 to 5 meters.
2 3D Pose Estimation and Skeleton Construction
Besides manual human skeletal joint annotation , a number of approaches have been designed to automatically construct a skeleton model from perception data through pose recognition and joint estimation. Some of these are based on methods used in RGB imagery, while others take advantage of the extra information available in a depth or RGB-D image. The majority of the current methods are based on body part recognition, and then fit a flexible model to the now ‘known’ body part locations. An alternate main methodology is starting with a ‘known’ prior, and fitting the silhouette or point cloud to this prior after the humans are localized . This section provides a brief review of autonomous skeleton construction methods based on visual data according to the information that is used. A summary of the reviewed skeleton construction techniques is presented in Table 1.
Due to the additional 3D geometric information that depth imagery can provide, many methods are developed to build a 3D human skeleton model based on a single depth image or a sequence of depth frames.
Human joint estimation via body part recognition is one popular approach to construct the skeleton model . A seminal paper by Shotton et al. in 2011 provided an extremely effective skeleton construction algorithm based on body part recognition, that was able to work in real time. A single depth image (independent of previous frames) is classified on a per-pixel basis, using a randomized decision forest classifier. Each branch in the forest is determined by a simple relation between the target pixel and various others. The pixels that are classified into the same category form the body part, and the joint is inferred by the mean-shift method from a certain body part, using the depth data to ‘push’ them into the silhouette. While training the decision forests takes a large number of images (around 1 million) as well as a considerable amount of computing power, the fact that the branches in the forest are very simple allows this algorithm to generate 3D human skeleton models within about 5 ms. An extended work was published in , with both accuracy and speed improved. Plagemann et al. introduced an approach to recognize body parts using Haar features and construct a skeleton model on these parts. Using data over time, they construct a Bayesian network, which produces the estimated pose using body part locations and starts with the previous pose as a prior . Holt et al. proposed Connected Poselets to estimate 3D human pose from depth data. The approach utilizes the idea of poselets , which is widely applied for pose estimation from RGB images. For each depth image, a multi-scale sliding window is applied, and a decision forest is applied to detect poselets and estimate human joint locations. Using a skeleton prior inspired by pictorial structures , the method begins with a torso point and connects outwards to body parts. By applying kinematic inference to eliminate impossible poses, they are able to reject incorrect body part classifications and improve their accuracy.
Another widely investigated methodology to construct 3D human skeleton models from depth imagery is based on nearest-neighbor matching . Several approaches for whole-skeleton matching are based on the Iterative Closest Point (ICP) method , which can iteratively decide a rigid transformation such that the input query points fit to the points in the given model under this transformation. Using point clouds of a person with known poses as a model, several approaches apply ICP to fit the unknown poses by estimating the translation and rotation to fit the unknown body parts to the known model. While these approaches are relatively accurate, they suffer from several drawbacks. ICP is computationally expensive for a model with as many degrees of freedom as a human body. Additionally, it can be difficult to recover from tracking loss. Typically the previous pose is used as the known pose to fit to; if tracking loss occurs and this pose becomes inaccurate, then further fitting can be difficult or impossible. Finally, skeleton construction methods based on the ICP algorithm generally require an initial T-pose to start the iterative process.
2.2 Construction from RGB Imagery
Early approaches and several recent methods based on deep learning focused on 2D or 3D human skeleton construction from traditional RGB or intensity images, typically by identifying human body parts using visual features (e.g., image gradients, deeply learned features, etc.), or matching known poses to a segmented silhouette.
Methods based on a single image: Many algorithms were proposed to construct human skeletal model using a single color or intensity image acquired from a monocular camera . Wang et al. constructs a 3D human skeleton from a single image using a linear combination of known skeletons with physical constraints on limb lengths. Using a 2D pose estimator , the algorithm begins with a known 2D pose and a mean 3D pose, and calculates camera parameters from this estimation. The 3D joint positions are recalculated using the estimated parameters, and the camera parameters are updated. The steps continue iteratively until convergence. This approach was demonstrated to be robust to partial occlusions and errors in the 2D estimation. Dong et al. considered the human parsing and pose estimation problems simultaneously. The authors introduced a unified framework based on semantic parts using a tailored And-Or graph. The authors also employed parselets and Mixture of Joint-Group Templates as the representation.
Recently, deep neural networks have proven their ability in human skeleton construction . Toshev and Szegedy employed Deep Neural Networks (DNNs) for human pose estimation. The proposed cascade of DNN regressors obtains pose estimation results with high precision. Fan et al. uses Dual-Source Deep Convolutional Neural Networks (DS-CNNs) for estimating 2D human poses from a single image. This method takes a set of image patches as the input and learns the appearance of each local body part by considering their previous views in the full body, which successfully addresses the joint recognition and localization issue. Tompson et al. proposed a unified learning framework based on deep Convolutional Networks (ConvNets) and Markov Random Fields, which can generate a heat-map to encode a per-pixel likelihood for human joint localization from a single RGB image.
Methods based on multiple images: When multiple images of a human are acquired from different perspectives by a multi-camera system, traditional stereo vision techniques can be employed to estimate depth maps of the human. After obtaining the depth image, a human skeleton model can be constructed using methods based on depth information (Section 2.2.1). Although there exists a commercial solution that uses marker-less multi-camera systems to obtain highly precise skeleton data at 120 frames per second (FPS) and approximately 25-50ms latency , computing depth maps is usually slow and often suffers from problems such as failures of correspondence search and noisy depth information. To address these problems, algorithms were also studied to construct human skeleton models directly from the multi-images without calculating the depth image . For example, Gall et al. introduced an approach to fully-automatically estimate the 3D skeleton model from a multi-perspective video sequence, where an articulated template model and silhouettes are obtained from the sequence. Another method was also proposed by Liu et al. , which uses a modified global optimization method to handle occlusions.
3 Benchmark Datasets With Skeletal Data
In the past five years, a large number of benchmark datasets containing 3D human skeleton data were collected in different scenarios and made available to the public. This section provides a complete review of the datasets as listed in Table 2. We categorize and discuss these datasets according to the type of devices used to acquire the skeleton information.
Early 3D human skeleton datasets were usually collected by a MoCap system, which can provide accurate locations of a various number of skeleton joints by tracking the markers attached on human body, typically in indoor environments. The CMU MoCap dataset is one of the earliest resources that consists of a wide variety of human actions, including interaction between two subjects, human locomotion, interaction with uneven terrain, sports, and other human actions. It is capable of recording 120 Hz with images of 4 megapixel resolution. The recent Human3.6M dataset is one of the largest MoCap datasets, which consists of 3.6 million human poses and corresponding images captured by a high-speed MoCap system. There are 4 basler high-resolution progressive scan cameras to acquire video data at 50 Hz. It contains activities by 11 professional actors in 17 scenarios: discussion, smoking, taking photo, talking on the phone, etc., as well as provides accurate 3D joint positions and high-resolution videos. The PosePrior dataset is the newest MoCap dataset that includes an extensive variety of human stretching poses performed by trained athletes and gymnasts. Many other MoCap datasets were also released, including the Pictorial Human Spaces , CMU Multi-Modal Activity (CMU-MMAC) Berkeley MHAD , Standford ToFMCD , HumanEva-I , and HDM05 MoCap datasets.
3.2 Datasets Collected by Structured-Light Cameras
Affordable structured-light cameras are widely used for 3D human skeleton data acquisition. Numerous datasets were collected using the Kinect v1 camera in different scenarios. The MSR Action3D dataset was captured using the Kinect camera at Microsoft Research, which consists of subjects performing American Sign Language gestures and a variety of typical human actions, such as making a phone call or reading a book. The dataset provides RGB, depth, and skeleton information generated by the Kinect v1 camera for each data instance. A large number of approaches used this dataset for evaluation and validation . The MSRC-12 Kinect gesture dataset is one of the largest gesture databases available. Consisting of nearly seven hours of data and over 700,000 frames of a variety of subjects performing different gestures, it provides the pose estimation and other data that was recorded with a Kinect v1 camera. The Cornell Activity Dataset (CAD) includes CAD-60 and CAD-120 , which contains 60 and 120 RGB-D videos of human daily activities, respectively. The dataset was recorded by a Kinect v1 in different environments, such as an office, bedroom, kitchen, etc. The SBU-Kinect-Interaction dataset contains skeleton data of a pair of subjects performing different interaction activities - one person acting and the other reacting. Many other datasets captured using a Kinect v1 camera were also released to the public, including the MSR Daily Activity 3D , MSR Action Pairs , Online RGBD Action (ORGBD) , UTKinect-Action , Florence 3D-Action , CMU-MAD , UTD-MHAD , G3D/G3Di , SPHERE , ChaLearn , RGB-D Person Re-identification , Northwestern-UCLA Multiview Action 3D , Multiview 3D Event , CDC4CV pose , SBU-Kinect-Interaction , UCF-Kinect , SYSU 3D Human-Object Interaction , Multi-View TJU , , and 3D Iconic Gesture datasets. The complete list of human-skeleton datasets collected using structured-light cameras are presented in Table 2.
3.3 Datasets Collected by Other Techniques
Besides the datasets collected by MoCap or structured-light cameras, additional technologies were also applied to collect datasets containing 3D human skeleton information, such as multiple camera systems, ToF cameras such as the Kinect v2 camera, or even manual annotation.
Due to the low price and improved performance of the Kinect v2 camera, it has become increasingly widely adopted to collect 3D skeleton data. The Telecommunication Systems Team (TST) created a collection of datasets using Kinect v2 ToF cameras, which include three datasets for different purposes. The TST fall detection dataset contains eleven different subjects performing falling activities and activities of daily living in a variety of scenarios; The TST TUG dataset contains twenty different individuals standing up and walking around; and the TST intake monitoring dataset contains food intake actions performed by 35 subjects .
Manual annotation approaches are also widely used to provide skeleton data. The KTH Multiview Football dataset contains images of professional football players during real matches, which are obtained using color sensors from 3 views. There are 14 annotated joints for each frame. Several other skeleton datasets are collected based on manual annotation, including the LSP dataset , and the TUM Kitchen dataset , etc.
Information Modality
Skeleton-based human representations are constructed from various features computed from raw 3D skeletal data that can be possibly acquired from various sensing technologies. We define each type of skeleton-based features extracted from each individual sensing technique as a modality. From the perspective of information modality, 3D skeleton-based human representations can be classified into four categories based on joint displacement, orientation, raw position, and combined information. Existing approaches falling in each categories are summarized in detail in Tables 3–6, respectively.
Features extracted from displacements of skeletal joints are widely applied in many skeleton-based representations due to the simple structure and easy implementation. They use information from the displacement of skeletal joints, which can either be the displacement between different human joints within the same frame or the displacement of the same joint across different time periods.
Representations based on relative joint displacements compute spatial displacements of coordinates of human skeletal joints in 3D space, which are acquired from the same frame at a time point.
The pairwise relative position of human skeleton joints is the most widely studied displacement feature for human representation . Within the same skeleton model obtained at a time point, for each joint in 3D space, the difference between the location of joint and joint is calculated by . The joint locations are often normalized, so that the feature is invariant to the absolute body position, initial body orientation and body size . Chen and Koskela implemented a similar feature extraction method based on pairwise relative position of skeleton joints with normalization calculated by , which is illustrated in Fig. 3(a).
Another group of joint displacement features extracted from the same frame for skeleton-based representation construction is based on the difference to a reference joint. In these features, the displacements are obtained by calculating the coordinate difference of all joints with respect to a reference joint, usually manually selected. Given the location of a joint and a given reference joint in the world coordinate system, Rahmani et al. defined the spatial joint displacement as , where the reference joint can be the skeleton centroid or a manually selected, fixed joint. For each sequence of human skeletons representing an activity, the computed displacements along each dimension (e.g., , or ) are used as features to represent humans. Luo et al. applied similar position information for feature extraction. Since the joint hip center has relatively small motions for most actions, they used that joint as the reference.
1.2 Temporal Joint Displacement
3D human representations based on temporal joint displacements compute the location difference across a sequence of frames acquired at different time points. Usually, they employ both spatial and temporal information to represent people in space and time.
A widely used temporal displacement feature is implemented by comparing the joint coordinates at different time steps. Yang and Tian introduced a novel feature based on the position difference of joints, called EigenJoints, which combines three categories of features including static posture, motion, and offset features. In particular, the joint displacement of the current frame with respect to the previous frame and initial frame is calculated. Ellis et al. introduced an algorithm to reduce latency for action recognition using a 3D skeleton-based representation that depends on spatio-temporal features computed from the information in three frames: the current frame, the frame collected 10 time steps ago, and the frame collected 30 frames ago. Then, the features are computed as the temporal displacement among those three frames. Another approach to construct temporal displacement representations incorporates the object being interacted with in each pose . This approach constructs a hierarchical graph to represent positions in 3D space and motion through 1D time. The differences of joint coordinates in two successive frames are defined as the features. Hu et al. introduced the joint heterogeneous features learning (JOULE) model through extracting the pose dynamics using skeleton data from a sequence of depth images. A real-time skeleton tracker is used to extract the trajectories of human joints. Then relative positions of each trajectory pair is used to construct features to distinguish different human actions.
The joint movement volume is another feature construction approach for human representation that also uses joint displacement information for feature extraction, especially when a joint exhibits a large movement . For a given joint, extreme positions during the full joint motion are computed along , , and axes. The maximum moving range of each joint along each dimension is then computed by , where ; and the joint volume is defined as , as demonstrated in Fig. 3(b). For each joint, and are flattened into a feature vector. The approach also incorporates relative joint displacements with respect to the torso joint into the feature.
2 Orientation-Based Representations
Another widely used information modality for human representation construction is based on joint orientations, since in general orientation-based features are invariant to human position, body size, and orientation to the camera.
Approaches based on spatial orientations of pairwise joints compute the orientation of displacement vectors of a pair of human skeletal joints acquired at the same time step.
A popular orientation-based human representation computes the orientation of each joint to the human centroid in 3D space. For example, Gu et al. collected the skeleton data with fifteen joints and extracted features representing joint angles with respect to the person’s torso. Sung et al. computed the orientation matrix of each human joint with respect to the camera, and then transformed the joint rotation matrix to obtain the joint orientation with respect to the person’s torso. A similar approach was also introduced in based on the orientation matrix. Xia et al. introduced Histograms of 3D Joint Locations (HOJ3D) features by assigning 3D joint positions into cone bins in 3D space. Twelve key joints are selected and their orientation are computed with respect to the center torso point. Using linear discriminant analysis (LDA), the features are reprojected to extract the dominant ones. Since the spherical coordinate system used in is oriented with the axis aligned with the direction a person is facing, their approach is view invariant.
Another approach is to calculate the orientation of two joints, called relative joint orientations. Jin and Choi utilized vector orientations from one joint to another joint, named the first order orientation vector, to construct 3D human representations. The approach also proposed a second order neighborhood that connects adjacent vectors. The authors used a uniform quantization method to convert the continuous orientations into eight discrete symbols to guarantee robustness to noise. Zhang and Tian used a two mode 3D skeleton representation, combining structural data with motion data. The structural data is represented by pairwise features, relating the positions of each pair of joints relative to each other. The orientation between two joints and was also used, which is given by where denotes the geometry distance between two joints and in 3D space.
2.2 Temporal Joint Orientation
Human representations based on temporal joint orientations usually compute the difference between orientations of the same joint across a temporal sequence of frames. Campbell and Bobick introduced a mapping from the Cartesian space to the “phase space”. By modeling the joint trajectory in the new space, the approach is able to represent a curve that can be easily visualized and quantifiably compared to other motion curves. Boubou and Suzuki described a representation based on the so-called Histogram of Oriented Velocity Vectors (HOVV), which is a histogram of the velocity orientations computed from 19 human joints in a skeleton kinematic model acquired from the Kinect v1 camera. Each temporal displacement vector is described by its orientation in 3D space as the joint moves from the previous position to the current location. By using a static skeleton prior to deal with static poses with little or no movement, this method is able to effectively represent humans with still poses in 3D space in human action recognition applications.
3 Representations Based on Raw Joint Positions
Besides joint displacements and orientations, raw joint positions directly obtained from sensors are also used by many methods to construct space-time 3D human representations.
A category of approaches flatten joint positions acquired in the same frame into a column vector. Given a sequence of skeleton frames, a matrix can be formed to naively encode the sequence with each column containing the flattened joint coordinates obtained at a specific time point. Following this direction, Hussein et al. computed the statistical Covariance of 3D Joints (Cov3DJ) as their features, as illustrated in Fig. 4. Specifically, given human joints with each joint denoted by , a feature vector is formed to encode the skeleton acquired at time : . Given a temporal sequence of skeleton frames, the Cov3DJ feature is computed by , where is the mean of all . Since not all the joints are equally informative, several methods were proposed to select key joints that are more descriptive . Chaaraoui et al. introduced an evolutionary algorithm to select a subset of skeleton joints to form features. Then a normalizing process was used to achieve position, scale and rotation invariance. Similarly, Reyes et al. selected 14 joints in 3D human skeleton models without normalization for feature extraction in gesture recognition applications.
Another group of representation construction techniques utilize the raw joint position information to form a trajectory, and then extract features from this trajectory, which are often called the trajectory-based representation. For example, Wei et al. used a sequence of 3D human skeletal joints to construct joint trajectories, and applied wavelets to encode each temporal joint sequence into features, which is demonstrated in Fig. 5. Gupta et al. proposed a cross-view human representation, which matches trajectory features of videos to MoCap joint trajectories and uses these matches to generate multiple motion projections as features. Junejo et al. used trajectory-based self-similarity matrices (SSMs) to encode humans observed from different views. This method showed great cross-view stability to represent humans in 3D space using MoCap data.
Similar to the application of deep learning techniques to extract features from images where raw pixels are typically used as input, skeleton-based human representations built by deep learning methods generally rely on raw joint position information. For example, Du et al. proposed an end-to-end hierarchical recurrent neural network (RNN) for the skeleton-based representation construction, in which the raw positions of human joints are directly used as the input to the RNN. Zhu et al. used raw 3D joint coordinates as the input to a RNN with Long Short-Term Memory (LSTM) to automatically learn human representations.
4 Multi-Modal Representations
Since multiple information modalities are available, an intuitive way to improve the descriptive power of a human representation is to integrate multiple information sources and build a multi-modal representation to encode humans in 3D space. For example, the spatial joint displacement and orientation can be integrated together to build human representations. Guerra-Filho and Aloimonos proposed a method that maps 3D skeletal joints to 2D points in the projection plane of the camera and computes joint displacements and orientations of the 2D joints in the projected plane. Gowayyed et al. developed the histogram of oriented displacements (HOD) representation that computes the orientation of temporal joint displacement vectors and uses their magnitude as the weight to update the histogram in order to make the representation speed-invariant.
Multi-modal space-time human representations were also actively studied, which are able to integrate both spatial and temporal information and represent human motions in 3D space. Yu et al. integrated three types of features to construct a spatio-temporal representation, including pairwise joint distances, spatial joint coordinates, and temporal variations of joint locations. Masood et al. implemented a similar representation by incorporating both pairwise joint distances and temporal joint location variations. Zanfir et al. introduced the so-called moving pose feature that integrates raw 3D joint positions as well as first and second derivatives of the joint trajectories, based on the assumption that the speed and acceleration of human joint motions can be described accurately by quadratic functions.
5 Summary
Through computing the difference of skeletal joint positions in 3D real-world space, displacement-based representations are invariant to absolute locations and orientations of people with respect to the camera, which can provide the benefit of forming view-invariant spatio-temporal human representations. Similarly, orientation-based human representations can provide the same view-invariance because they are also based on the relative information between human joints. In addition, since orientation-based representations do not rely on the displacement magnitude, they are usually invariant to human scale variations. Representations based directly on raw joint positions are widely used due to the simple acquisition from sensors. Although normalization procedures can make human representations partially invariant to view and scale variations, more sophisticated construction techniques (e.g., deep learning) are typically needed to develop robust human representations.
Representations without involving temporal information are suitable to address problems such as pose and gesture recognition. However, if we want the representations to be capable of encoding dynamic human motions, temporal information needs to be integrated. Activity recognition can benefit from spatio-temporal representations that incorporate time and space information simultaneously. Among space-time human representations, approaches based on joint trajectories can be designed to be insensitive to motion speed invariance. In addition, fusion of multiple feature modalities typically results in improved performance (further analysis is provided in Section 7.1).
Representation Encoding
Feature encoding is a necessary and important component in representation construction , which aims at integrating all extracted features together into a final feature vector that can be used as the input to classifiers or other reasoning systems. In the scenario of 3D skeleton-based representation construction, the encoding methods can be broadly grouped into three classes: concatenation-based encoding, statistics-based encoding, and bag-of-words encoding. The encoding technique used by each reviewed human representation is summarized in the Feature Encoding column in Tables 3–6.
We loosely define feature concatenation as a representation encoding approach, which is a popular method to integrate multiple features into a single feature vector during human representation construction. Many methods directly use extracted skeleton-based features, such as displacements and orientations of 3D human joints, and concatenate them into a 1D feature vector to build a human representation . For example, Fothergill et al. encoded the feature vector by concatenating 35 skeletal joint angles, 35 joint angle velocities, and 60 joint velocities into a 130-dimensional vector at each frame. Then, feature vectors from a sequence of frames are further concatenated into a big final feature vector that is fed into a classifier for reasoning. Similarly, Gong et al. directly concatenated 3D joint positions into a 1D vector as a representation at each frame to address the time series segmentation problem.
2 Statistics-Based Encoding
Statistics-based encoding is a common but effective method to incorporate all features into a final feature vector, without applying any feature quantization procedure. This encoding methodology processes and organizes features through simple statistics. For example, the Cov3DJ representation , as illustrated in Fig. 4, computes the covariance of a set of 3D joint position vectors collected across a sequence of skeleton frames. Since a covariance matrix is symmetric, only upper triangle values are utilized to form the final feature in . An advantage of this statistics-based encoding approach is that the size of the final feature vector is independent of the number of frames. Moreover, Wang et al. proposed an open framework by using the kernel matrix over feature dimensions as a generic representation and elevated the covariance representation to the unlimited opportunities.
The most widely used statistics-based encoding methodology is histogram encoding, which uses a 1D histogram to estimate the distribution of extracted skeleton-based features. For example, Xia et al. partitioned the 3D space into a number of bins using a modified spherical coordinate system and counted the number of joints falling in each bin to form a 1D histogram, which is called the Histogram of 3D Joint Positions (HOJ3D). A large number of skeleton-based human representations using similar histogram encoding methods were also introduced, including Histogram of Joint Position Differences (HJPD), Histogram of Oriented Velocity Vectors (HOVV), and Histogram of Oriented Displacements (HOD), among others . When multi-modal skeleton-based features are involved, concatenation-based encoding is usually employed to incorporate multiple histograms into a single final feature vector .
3 Bag-of-Words Encoding
Unlike concatenation and statistics-based encoding methodologies, bag-of-words encoding applies a coding operator to project each high-dimensional feature vector into a single code (or word) using a learned codebook (or dictionary) that contains all possible codes. This procedure is also referred to as feature quantization. Given a new instance, this encoding methodology uses the normalized frequency vector of code occurrence as the final feature vector. Bag-of-words encoding is widely employed by a large number of skeleton-based human representations . According to how the dictionary is learned, the encoding methods can be broadly categorized into two groups, based on clustering or sparse coding.
The k-means algorithm is a popular unsupervised learning method that is commonly used to construct a dictionary. Wang et al. grouped human joints into five body parts, and used the k-means algorithm to cluster the training data. The indices of the cluster centroids are utilized as codes to form a dictionary. During testing, query body part poses are quantized using the learned dictionary. Similarly, Kapsouras and Nikolaidis used the k-means clustering method on skeleton-based features consisting of joint orientations and orientation differences in multiple temporal scales, in order to select representative patterns to build a dictionary.
Sparse coding is another common approach to construct efficient representations of data as a (often linear) combination of a set of distinctive patterns (i.e., codes) learned from the data itself. Zhao et al. introduced a sparse coding approach regularized by the norm to construct a dictionary of templates from the so-called Structured Streaming Skeletons (SSS) features in a gesture recognition application. Luo et al. proposed another sparse coding method to learn a dictionary based on pairwise joint displacement features. This approach uses a combination of group sparsity and geometric constraints to select sparse and more representative patterns as codes. An illustration of the dictionary learning method to encode skeleton-based human representations is presented in Fig. 6.
4 Summary
Due to its simplicity and high efficiency, the concatenation-based feature vector construction method is widely applied in real-time online applications to reduce processing latency. The method is also used to integrate features from multiple sources into a single vector for further encoding/processing. By not requiring a feature quantization process, statistics-based encoding, especially based on histograms, is efficient and relatively robust to noise. However, the statistics-based encoding method is incapable of identifying the representative patterns and modeling the structure of the data, thus making it lacking in discriminative power. Bag-of-words encoding can automatically find a good over-complete basis and encode a feature vector using a sparse solution to minimize approximation error. Bag-of-words encoding is also validated to be robust to data noise. However, dictionary construction and feature quantization require additional computation. According to the performance reported by the papers (as further analyzed in Section 7.1), the bag-of-words encoding can generally obtain superior performance.
Structure and Topological Transition
While most skeleton-based 3D human representations are based on pure low-level features extracted from the skeleton data in 3D Euclidean space, several works studied mid-level features or feature transition to other topological space. This section categorizes the reviewed approaches from the structure and transition perspective into three groups: representations using low-level features in Euclidean space, representations using mid-level features based on human body parts, and manifold-based representations. The major class of each representation categorized from this perspective is listed in the Structure and Transition column in Tables 3–6.
A simple, straightforward framework to construct skeleton-based representations is to use low-level features computed from 3D skeleton data in Euclidian space, without considering human body structures or applying feature transition. Most of the existing representations fall in this category. The representations can be constructed by single-layer methods, or by approaches with multiple layers.
An example of the single-layer representation construction method is the EigenJoints approach introduced by Yang and Tian . This approach extracts low-level features from skeletal data, such as pairwise joint displacements, and uses Principal Component Analysis (PCA) to perform dimension reduction. Many other existing human representations are also based on low-level skeleton-based features without modeling the hierarchy of the data.
Several multi-layer techniques were also implemented to create skeleton-based human representations from low-level features. In particular, deep learning approaches inherently consist of multiple layers with the intermediate and output layers encoding different levels of features . The multi-layer deep learning approaches have attracted an increasing attention in recent several years to learn human representations directly from human joint positions . Inspired by the spatial pyramid method to incorporate multi-layer image information, temporal pyramid methods were introduced and used by several skeleton-based human representations to capture the multi-layer information in the time dimension . For example, a temporal pyramid method was proposed by Zhang et al. to capture long-term dependencies, as illustrated in Fig. 7. In this example, a temporal sequence of eleven frames is used to represent a tennis-serve motion, and the joint of interest is the right wrist, as denoted by the red dots in Fig 7. When three levels are used in the temporal pyramid, level 1 uses human skeleton data at all time points (); level 2 selects the joints at odd time points (); and level 3 continues this selection process and keeps half of the temporal data points () to compute long-term orientation changes.
2 Representations Based on Body Part Models
Mid-level features based on body part models are also used to construct skeleton-based human representations. Since these mid-level features partially take into account the physical structure of human body, they can usually result in improved discrimination power to represent humans .
Wang et al. decomposed a kinematic human body model into five parts, including the left/right arms/legs and torso, each consisting of a set joints. Then, the authors used a data mining technique to obtain a spatiotemporal human representation, by capturing spatial configurations of body parts in one frame (by spatial-part-sets) as well as body part movements across a sequence of frames (by temporal-part-sets), as illustrated in Fig. 8. With this human representation, the approach was able to obtain a hierarchical data that can simultaneously model the correlation and motion of human joints and body parts. Nie et al. implemented a spatial-temporal And-Or graph model to represent humans at three levels including poses, spatiotemporal-parts, and parts. The hierarchical structure of this body model captures the geometric and appearance variations of humans at each frame. Du et al. introduced a deep neural network to create a body part model and investigate the correlation of body parts.
Bio-inspired body part methods were also introduced to extract mid-level features for skeleton-based representation construction, based on body kinematics or human anatomy. Chaudhry et al. implemented a bio-inspired mid-level feature to represent people based on 3D skeleton information through leveraging findings in the area of static shape encoding in the neural pathway of the primate cortex . By showing primates various 3D shapes and measuring the neural response when changing different parameters of the shapes, the primates’ internal shape representation can be estimated, which was then applied to extract body parts to construct skeleton-based representations. Zhang and Parker implemented a bio-inspired predictive orientation decomposition (BIPOD) using mid-level features to construct representations of people from 3D skeleton trajectories, which is inspired by biological research in human anatomy. This approach decomposes a human body model into five body parts, and then projects 3D human skeleton trajectories onto three anatomical planes (i.e., coronal, transverse and sagittal planes), as illustrated in Fig. 9. By estimating future skeleton trajectories, the BIPOD representation possesses the ability to predict future human motions.
3 Manifold-Based Representations
A number of methods in the literature transited the skeleton data in 3D Euclidean space to another topological space (i.e., manifold) in order to process skeleton trajectories as curves within the new space. This category of methods typically utilizes a trajectory-based representation.
Vemulapalli et al. introduced a skeletal representation that was created in the Lie group , which is a curved manifold, based on the observation that 3D rigid body motions are members of the space. Using this representation, joint trajectories can be modeled as curves in the Lie group, shown in Fig. 10(a). This manifold-based representation can model 3D geometric relationships between joints using rotations and translations in 3D space. Since analyzing curves in the Lie group is not easy, the approach maps the curves from the Lie group to its Lie algebra, which is a vector space. Gong and Medioni introduced a spatio-temporal manifold and a dynamic manifold warping method, which is an adaptation of dynamic time warping methods for the manifold space. Spatial alignment is also used to deal with variations of viewpoints and body scales. Slama et al. introduced a multi-stage method based on a Grassmann manifold. Body joint trajectories are represented as points on the manifold, and clustered to find a ‘control tangent’ defined as the mean of a cluster. Then a query human joint trajectory is projected against the tangents to form a final representation. This manifold was also applied by Azary and Savakis to build sparse human representations, shown in Fig. 10(b). Anirudh et al. introduced the transport square-root velocity function (TSRVF) to encode humans in 3D space, which provides an elastic metric to model joint trajectories on Riemannian manifolds. Amor et al. proposed to model the evolution of human skeleton shapes as trajectories on Kendall’s shape manifolds, and used a parameterization-invariant metric for aligning, comparing, and modeling skeleton joint trajectories, which can deal with noise caused by large variability of execution rates within and across humans. Devanne et al. introduced a human representation by comparing the similarity between human skeletal joint trajectories in a Riemannian manifold .
4 Summary
Single or multi-layer human representations based on low-level features directly extract features from 3D skeletal data without considering the physical structure of human body. The kinematic body structure is coarsely encoded by human representations based on mid-level features extracted from body part models, which can capture the relationship of not only joints but also body parts. Manifold-based representations map motion joint trajectories into a new topological space, in the hope of finding a more descriptive representation in the new space. Good performance of all these human representations was reported in the literature. However, with the complexity increment of the activities, especially long-time activities, low-level feature structures may not a good choice due to their limited representation capability. In this case, body-part models and manifold-based representations can often improve recognition performance.
Feature Engineering
Feature engineering is one of the most fundamental research problems in computer vision and machine learning research. Early feature engineering techniques for human representation construction are manual; features are hand-crafted and their importance are manually decided. In recent years, we have been witnessing a clear transition from manual feature engineering to automated feature learning and extraction. In this section, we categorize and analyze human representations based on 3D skeleton data from the perspective of feature engineering. The feature engineering approach used by each human representation is summarized in the Feature Engineering column in Tables 3–6.
Hand-crafted features are manually designed and constructed to capture certain geometric, statistical, morphological, or other attributes of 3D human skeleton data, which dominated the early skeleton-based feature extraction methods and are still intensively studied in modern research.
Lv and Nevatia decomposed the high dimensional 3D joint space into a set of feature spaces where each of them corresponds to the motion of a single joint or a combination of related multiple joints. Ofli et al. proposed a human representation called the Sequence of the Most Informative Joints (SMIJ), by selecting a subset of skeletal joints to extract category-dependent features. Zhao et al. described a method of representing humans using the similarity of current and previously seen skeletons in a gesture recognition application. Pons-Moll et al. used qualitative attributes of the 3D skeleton data, called posebits, to estimate human poses, by manually defining features such as joint distance, articulation angle, relative position, etc. Huang et al. proposed to utilize hand-crafted features including skeletal joint positions to locate key frames and track humans from a multi-camera video. In general, the majority of the existing skeleton-based human representations employ hand-crafted features, especially, the methodologies based on histograms and manifolds, as presented by Tables 3–6.
2 Representation Learning
In many vision and reasoning tasks, good performance is all about the right representation. Thus, automated learning of skeleton-based features has become highly active in the task of human representation construction based on 3D skeletal data. These skeleton-based representation learning methods can be broadly divided into three groups: dictionary learning, unsupervised feature learning, and deep learning.
Dictionary learning aims at learning a basis set (dictionary) to encode a feature vector as a sparse linear combination of basis elements, as well as to adapt the dictionary to the data in a specific task. Learning a dictionary is the foundation of the bag-of-words encoding. In the literature of 3D skeleton-based representation creation, the k-means algorithm and sparse coding are the most commonly used techniques for dictionary learning. A number of these methods are reviewed in Section 4.3.
2.2 Unsupervised Feature Learning
The objective of unsupervised feature learning is to discover low-dimensional features that capture the underlying structure of the input data in a higher dimension. For example, the traditional PCA method is applied for dimension reduction to extract low-dimensional features from raw skeleton features . Negin et al. designed a feature selection method to build human representations from 3D skeletal data. This approach describes humans via a collection of time-series feature computed from the skeletal data, and discriminatively optimizes a random decision forest model over this collection to identify the most effective set of features in time and space dimensions.
Very recently, several multi-modal feature learning approaches via sparsity-inducing norms were introduced to integrate different types of features, such as color-depth and skeleton-based features, to produce a compact, informative 3D representation of people. Shahroudy et al. recently developed a multi-modal feature learning method to fuse the RGB-D and skeletal information into an integrated set of discriminative features. This approach uses the group- norm to force features from the same view to be activated or deactivated together, and applies the norm to allow a single feature within a deactivated view to be activated. The authors also introduced a multi-modal multi-part human representation based on a hierarchical mixed norm , which regularizes structured features of each joint subset and applies sparsity between them. Another heterogenous feature learning algorithm was introduced by Hu et al. . The approach casted joint feature learning as a least-square optimization problem that employs the Frobenius matrix norm as the regularization term that provides an efficient, closed-form solution.
2.3 Deep Learning
While unsupervised feature learning allows for assigning a weight to each feature element, this methodology still relies on manually crafted features as the initial set. Deep learning, on the other hand, attempts to automatically learn a multi-level representation directly from raw data, by exploring a hierarchy of factors that may explain the data. Several such approaches were developed to learn human representations from 3D skeletal joint positions directly acquired by sensors in recent several years. For example, Du et al. proposed an end-to-end hierarchical recurrent neural network (RNN) to construct a skeleton-based human representation. In this method, the whole skeleton is divided into five parts according to human physical structure, and separately fed into five bidirectional RNNs. As the number of layers increases, the representations extracted by the subnets are hierarchically fused to build a higher-level representation, as illustrated in Fig. 11. Zhu et al. introduced a method based on RNNs with Long Short-Term Memory (LSTM) to automatically learn human representations and model long-term temporal dependencies. In this method, joint positions are used as the input at each time slot to the LST-RNNs that can model the joint co-occurrences to characterize human motions. Wu and Shao proposed to utilize deep belief networks to model the distribution of skeleton joint locations and extract high-level features to represent humans at each frame in 3D space. Salakhutdinov et al. proposed a compositional learning architecture that integrates deep learning models with structured hierarchical Bayesian models. Specifically, this approach learns a hierarchical Dirichlet process (HDP) prior over top-level features in a deep Boltzmann machine (DBM), which simultaneously learns low-level generic features, high-level features that capture the correlation among the low-level features, and a category hierarchy for sharing priors over the high-level features.
3 Summary
Hand-crafted features still dominate human representations based on 3D skeletal data in the literature. Although several approaches showed great performance various applications, hand-crafting these manual features typically requires significant domain knowledge and careful parameter tuning. Most hand-crafted feature extraction methods are sensitive to their parameters values; poor parameter tuning can dramatically decrease the recognition performance. Also, the requirement of domain knowledge makes hand-crafted features not robust to various situations. Unsupervised dictionary and feature learning approaches can automatically determine which types of skeleton-based features or templates are more representative, although they typically use hand-craft features as the input. Deep learning, on the other hand, can directly work with the raw skeleton information, and automatically discover and create features. However, deep learning methods are typically computationally expensive, which currently might not be suitable for online, real-time applications.
Discussion
In this section, we compare the accuracy and efficiency of different approaches using several most used datasets, including MSR Action3D, CAD-60, MSRC-12, and HDM05, which cover both structured light sensors (Kinect v1) and motion capture sensor systems. The performance is evaluated using the precision metric, since almost all the existing approaches report the precision results. The detailed comparison of different approaches is presented in Table 7.
From Table 7, it is observed that there is no single approach that is able to guarantee the best performance over all datasets. Performance of each approach varies when applied to different benchmark datasets. Generally, methods using multimodal information can have better activity recognition performance in comparison to methods based on single feature modality. We can also observe that the bag-of-words feature encoding is able to improve the performance. For feature structure transition, the reviewed representations obtain similar recognition performance as shown in Table 7. For feature engineering, learning-based methods, including deep learning, unsupervised feature learning and dictionary learning are proved to provide superior activity recognition results in comparison to traditional hand-crafted feature engineering methods.
As a side note, several public software packages are available, which implement 3D skeletal representations of people. The representations with open-source implementations include Ker-RP , lie group manifold , orientation matrix , temporal relational features , node feature map . We provide the web link to these open-source packages in the reference .
2 Future Research Directions
Human representations based on 3D skeleton data can possess several desirable attributes, including the ability to incorporate spatio-temporal information, invariance to variations of viewpoint, human body scale, and motion speed, and real-time, online performance. The characteristics of each reviewed representation are presented in Tables 3–6. While significant progress has been achieved on human representations based on 3D skeletal data, there are still numerous research opportunities. Here we briefly summarize some of the prevalent problems and provide possible future directions.
Fusing skeleton data with human texture and shape models. Although 3D skeleton data can be applied to construct descriptive representations of humans, it is incapable of encoding texture information, and therefore cannot effectively represent human-object interaction. In addition, other human models such as shape-based representation can also increase the description capability of humans. Integrating texture and shape information with skeleton data to build a multisensory representation has the potential to address this problem and improve the descriptive power of the existing space-time human representations.
General representation construction via cross-training. A variety of devices can provide skeleton data but with different kinematic models. It is desirable to develop cross-training methods that can utilize skeleton data from different devices to build a general representation that works with different skeleton models . A method of unifying skeleton data to the same format is also useful to integrate available benchmarks dataset and provide sufficient data to modern data-driven, large-scale representation learning methods such as deep learning.
Protocol for representation evaluation. There is a strong need of a protocol to benchmark skeleton-based human representations, which must be independent of learning and application-level evaluations. Although the representations have been qualitatively assessed based on their characteristics (e.g., scale-invariance, etc.), a beneficial future direction is to design quantitative evaluation metrics to facilitate evaluating and comparing the human representations.
Automated skeleton-based representation learning. Deep learning and multi-modal feature learning have recently shown compelling performance in a variety of computer vision and machine learning tasks, but are not well investigated in skeleton-based representation learning and can be a promising future research direction. Moreover, as human skeletal data contains kinematic structures, an interesting problem is how to integrate this structure as a prior in representation learning.
Real-time, anywhere skeleton estimation of arbitrary poses. Skeleton-based human representations heavily rely on the quality of 3D skeleton tracking. A possible future direction is to extract skeleton information of unconventional human poses (e.g., beyond gaming related poses using a Kinect sensor). Another future direction is to reliably extract skeleton information in an outdoor environment using depth data acquired from other sensors such as stereo vision and LiDAR. Although recent works based on deep learning showed promising skeleton tracking results, real-time processing must be ensured for real-word online applications.
Conclusion
This paper presents a unique and comprehensive survey of the state-of-the-art space-time human representations based 3D skeleton data that is now widely available. We provide a brief overview of existing 3D skeleton acquisition and construction methods, as well as a detailed categorization of the 3D skeleton-based representations from four key perspectives, including information modality, representation encoding, structure and topological transition, and feature engineering. We also compare the pros and cons of the methods in each perspective. We observe that multimodal representations that can integrate multiple feature sources usually lead to better accuracy in comparison to methods based on a single individual feature modality. In addition, learning-based approaches for representation construction, including deep learning, unsupervised feature learning and dictionary learning, have demonstrated promising performance in comparison to traditional hand-crafted feature engineering methods. Given the significant progress in current skeleton-based representations, there exist numerous future research opportunities, such as fusing skeleton data with RGB-D images, cross-training, and real-time, anywhere skeleton estimation of arbitrary poses.