SLOPER4D: A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments
Yudi Dai, Yitai Lin, Xiping Lin, Chenglu Wen, Lan Xu, Hongwei Yi, Siqi Shen, Yuexin Ma, Cheng Wang
Introduction
Urban-level human motion capture is attracting more and more attention, which targets at acquiring consecutive fine-grained human pose representations, such as 3D skeletons and parametric mesh models, with accurate global locations in the physical world. It is essential for human action recognition, social-behavioral analysis, and scene perception and further benefits many downstream applications, including Augmented/Virtual Reality, simulation, autonomous driving, smart city, sociology, etc. However, capturing extra large-scale dynamic scenes and annotating detailed 3D representations for humans with diverse poses is not trivial.
Over the past decades, a large number of datasets and benchmarks have been proposed and have greatly promoted the research in 3D human pose estimation (HPE). They can be divided into two main categories according to the capture environment. The first class usually leverages marker-based systems , cameras , or RGB-D sensors to capture human local poses in constrained environments. However, the optical system is sensitive to light and lacks depth information, making it unstable in outdoor scenes and difficult to provide global translations, and the RGB-D sensor has limited range and could not work outdoors. The second class attempts to take advantage of body-mounted IMUs to capture occlusion-free 3D poses in free environments. However, IMUs suffer from severe drift for long-term capturing, resulting in misalignments with the human body. Then, some methods exploit additional sensors, such as RGB camera , RGB-D camera , or LiDAR to alleviate the problem and make obvious improvement. However, they all focus on HPE without considering the scene constraints, which are limited for reconstructing human-scene integrated digital urban and human-scene natural interactions.
To capture human pose and related static scenes simultaneously, some studies use wearable IMUs and body-mounted camera or LiDAR to register the human in large real scenarios and they are promising for capturing human-involved real-world scenes. However, human pose and scene are decoupled in these works due to the ego view, where auxiliary visual sensors are used for collecting the scene data while IMUs are utilized for obtaining the 3D pose. Different from them, we propose a novel setting for human-scene capture with wearable IMUs and global-view LiDAR and camera, which can provide multi-modal data for more accurate 3D HPE.
In this paper, we propose a huge scene-aware dataset for sequential human pose estimation in urban environments, named SLOPER4D. To our knowledge, it is the first urban-level 3D HPE dataset with multi-modal capture data, including calibrated and synchronized IMU measurements, LiDAR point clouds, and images for each subject. Moreover, the dataset provides rich annotations, including 3D poses, SMPL models and locations in the world coordinate system, 2D poses and bounding boxes in the image coordinate system, and reconstructed 3D scene mesh. In particular, we propose a joint optimization method for obtaining accurate and natural human motion representations by utilizing multi-sensor complementation and scene constraints, which also benefit global localization and camera calibration in the dynamic acquisition process. Furthermore, SLOPER4D consists of over 15 sequences in 10 scenes, including library, commercial street, coastal runway, football field, landscape garden, etc., with 2k13k area size and trajectory length for each sequence. By providing multi-modal capture data and diverse human-scene-related annotations, SLOPER4D opens a new door to benchmark urban-level HPE.
We conduct extensive experiments to show the superiority of our joint optimization approach for acquiring high-quality 3D pose annotations. Additionally, based on our proposed new dataset, we benchmark two critical tasks: camera-based 3D HPE and LiDAR-based 3D HPE, as well as provide benchmarks for GHPE.
Our contributions are summarized as follows:
We propose the first large-scale urban-level human pose dataset with multi-modal capture data and rich human-scene annotations.
We propose an effective joint optimization method for acquiring accurate human motions in both local and global by integrating LiDAR SLAM results, IMU poses, and scene constraints.
We benchmark two HPE tasks as well as a GHPE task on SLOPER4D, demonstrating its potential of promoting urban-level 3D HPE research.
Related Work
Many datasets have been proposed with different sensors and setups to facilitate the research on 3D human pose estimation. The H3.6M is a large-size dataset providing synchronized video with optical-based MoCap in studio environments. To perform markerless capture in different indoor scenes, PROX uses an RGB-D sensor to scan a single person. EgoBody uses multiple RGB-D sensors to pre-scan the room and scan the interacting persons. LiDARHuman26M can capture long-range human motions with static LiDAR and IMUs. However, they are limited to static environments, human activities, and interactions. 3DPW is the first dataset providing 3D annotations in the wild which uses a single hand-held RGB camera to optimize human pose from IMUs for a certain period of frames. It doesn’t provide accurate global translation and 3D scenes. HPS reconstructs the human body pose using IMUs and self-localizes it with a head-mounted camera in large 3D scenes, but it heavily relies on the pre-built map. HSC4D removes the reliance on the pre-built map and achieves global human motion capture in large scenes. However, the camera in HPS and the LiDAR in HSC4D are only used to perceive the environment rather than capture human data. With the scene-aware dataset we proposed for global human pose estimation, we can benchmark the 3D HPE in the wild with the LiDAR or camera modalities.
2 Human Localization and Scene Mapping
Human self-localization aims at estimating the 6-DoF of the human subject in global coordinates. The image-based methods regress locations directly from a single image with a pre-built map. The scene-specific property makes them hard to generalize to unseen scenes. LiDAR is widely used in Simultaneous Localization and Mapping (SLAM) due to its robustness and low drift. To address the drift problem and improve robustness in dynamic motions, RGB cameras , IMU , or both , have been integrated with the mapping task. Most attention has been paid to autonomous driving or robotics from the third-person view and they usually do not focus on humans. To achieve self-localization, LiDAR is designed as backpacked and hand-held. To efficiently capture human motions and reconstruct urban scenes, we utilize LiDAR with a built-in IMU (different from the IMUs for motion capture) and propose a pipeline for constructing multi-modal data. This approach provides accurate information on human motions at both local and global levels, as well as enables mapping in large outdoor environments.
3 Global 3D Human Pose Estimation
Most studies recover human meshes in camera coordinate or root-relative poses. Recovering global human motions in unconstrained scenes is a challenging topic in computer vision and has gained more and more research interest in recent years. IMU sensors are widely used in commercial and research activities , and are attached to body limbs to capture human motions in studio-environments. But it suffers severe drift in the wild. Some methods rely on additional RGB or pre-scan maps , or LiDAR to complement the IMUs in large-scale scenes. Based on human-scene interaction, some work proposed scene-aware solutions using static cameras to obtain accurate and scene-natural human motions. 4DCapture uses a dynamic head-mounted camera to self-localize and reconstruct the scene with the Struct From Motion method. However, it often fails when the illumination changes in the wild. MOVER uses a single camera to optimize the 3D objects in a static scene, resulting in better 3D scene reconstruction and human motions. GLAMR uses global trajectory predictions to constrain both human motions and dynamic camera poses, achieving state-of-the-art results on in-the-wild videos. However, it lacks a benchmark for quantitatively comparing different HPE methods on a global level. To deal with this limitation, we propose SLOPER4D, the first large-scale urban-level human pose dataset with rich 2D/3D annotations.
SLOPER4D Dataset
SLOPER4D collects scene-aware 4D human data with our body-worn capturing system in urban scenes. In this section, we first introduce the data acquisition in Sec. 3.1, second, we detail the data construction and annotation process in Sec. 3.2, then we introduce the global optimization-based Sec. 3.3 method to obtain high quality both 3D/2D data, finally, we compare our dataset in Sec. 3.4 with the existing datasets and highlight our novelty.
Hardware setup. As shown in LABEL:fig:teaser, during the data collection procedure, the scanning person follows the performer (IMUs wearer) and scans him with a LiDAR and a camera on the helmet. Additionally, Fig. 1 shows the hardware details of our capturing system. Regarding the sensor module, the 128-beams Ouster-os1 LiDAR and the DJI-Action2 wide-range camera are rigidly installed on the helmet. To capture raw human motions, we use Noitom’s inertial MoCap product, PN Studio, to attach 17 wireless IMUs to the IMU wearer’s body limbs, torso, and head. The camera’s field of view (FOV) is 116°84° and the LiDAR’s FOV is 360°45°. To make the performer within the LiDAR’s FOV as much as possible, we tilt the LiDAR down around 45°. Regarding the storage module, the scanning person’s backpack places a wireless IMU data receiver, a 24V battery, and an Intel NUC11. The mini-computer NUC11 stores IMU data from the wireless receiver and point clouds from LiDAR in real time. Videos are stored locally in the camera. The LiDAR and NUC11 are both powered by the battery.
Coordinate systems. Let’s define three coordinate systems: 1) IMU coordinate system {}: the origin is at LiDAR wearer’s spine base at the starting time, and the axis is pointing left/upward/forward of the human. 2) LiDAR Coordinate system {}: the origin is at the center of the LiDAR, and the axis is pointing right/forward/upward of the LiDAR. 3) Global/World coordinate system {}: the origin is on the floor of the LiDAR wearer’s starting position, and the axis is pointing right/forward/upward of the LiDAR wearer.
Calibration. Following the setup in , we use a chessboard to calibrate the camera intrinsic and introduces a terrestrial laser scanner (TLS) to obtain accurate camera extrinsic parameter, . Due to the LiDAR point cloud being too sparse, we manually choose the corresponding points both on the 2D image and the TLS map registered to the point cloud, and then we solve the perspective-n-point (PnP) problem to obtain . For every 3D scene, the calibration , which transforms {} to {} is manually set to make the ground’s -axis upward and height to zero for the starting position. By using singular value decomposition, the calibration , which transforms {} to {}, is calculated through the similarity between IMU trajectory and LiDAR trajectory on the XY plane. Synchronization. The synchronization of data from multiple sensors in human subject data is achieved through peak detection. Before and after the capture, the subject is asked to perform jumps. Then the peak height time in IMU is automatically detected and the peak times in the LiDAR and camera data are manually identified. Finally, all modalities are aligned by the peaks and downsampled to match the LiDAR frame rate of 20 Hz.
2 Data Processing
2D pose detection. We use Detectron to detect and Deepsort to track humans in videos. However, the tracking often fails due to the IMUs wearer entering/exiting the field of view or occlusions. To solve this problem, we manually assign the same ID for the tracked person in a video sequence. As for 3D point cloud reference, we project them on images according to the . However, due to the jitter brought by dynamic motions, the camera and the LiDAR are not perfectly rigidly connected. Thus, will be further optimized in Sec. 3.3.
LiDAR-inertial localization and mapping. The LiDAR-only method often fails in mapping because of the dynamic head rotation and crowded urban environments. Incorporating an IMU can compensate for motion distortion in a LiDAR scan and provide an accurate initial pose. Using a LiDAR with an integrated IMU, and by combining Kalman filter-based lidar-inertial odometry with factor graph-based loop closure optimization, we successfully estimate the ego-motion of LiDAR and build the global consistency 3D scene map with n frame point clouds . To provide accurate scene constrain in Sec. 3.3, we utilize the VDB-Fusion to generate a clean scene mesh that excludes moving objects.
3 Data Optimization
To obtain precise and scene-plausible human motion in the world coordinate system, we use scene geometry with several physic-based terms to perform joint optimizations to find the optimal motion that minimize . In a k-frame segment, the optimization is written as:
where is a smoothness term, which consists of a translation loss , an orientation loss , and a joints loss . is a scene-aware-contact term, is a pose prior term, and is a mesh-to-points term. The , , , , , and are loss terms’ coefficients. is minimized with a gradient descent algorithm.
Scene-aware contact term. we compare the movement of every foot vertices in IMU motions and label the foot as stable if its velocity is less than 0.1 . Finally, the Chamfer Distance (CD) between this foot and its closest surface is expressed as the scene contact loss .
Pose prior term. The poses estimated by IMUs are roughly accurate but will likely cause some misalignments to the end of the body limb due to the accumulating error. Hence, is used to constrain the close to the initial value at the beginning of the optimization.
Mesh-to-points term. The point cloud from the moving LiDAR provides strong prior depth information. However, though the SMPL mesh is watertight and complete, the human points are sparse and partial, which makes the registration methods such as ICP, not ideal as expected. To address this issue, we propose a viewpoint-based mesh-to-point loss function . First, we remove the hidden SMPL mesh faces from the LiDAR’s viewpoint. Then we sample points, denoted as , from the remaining faces by LiDAR resolution. The loss is defined as the Chamfer Distance from to .
All loss terms functions are detailed as follows:
Camera extrinsic optimization. We aim to optimize extrinsic parameters for every frame by minimizing the , which comprises of the keypoints loss and the bounding box loss . The measures the mean square error (MSE) between the 2D human keypoints in the image and the 3D human keypoints of the optimized SMPL model projected to the image with ; the computes the Intersection over Union(IoU) loss between the 2D human bounding box in the image and the 3D human bounding box projected to the image with .
where and are constant coefficients.
4 Dataset Comparison
SLOPER4D is the first large-scale urban-level human pose dataset with multi-modal capture data and rich human-scene annotations for GHPE. The head-mounted LiDAR and camera are utilized to simultaneously record the IMU-wearer’s activities, including running outside, playing football, visiting, reading, climbing/descending stairs, discussing, borrowing a book, greeting, etc.
The dataset consists of 15 sequences from 12 human subjects in 10 locations. There are a total of 100k LiDAR frames, 300k video frames, and 500k IMU-based motion frames captured over a total distance of more than 8 and an area of up to 13,000 . The results of our dataset are shown in Fig. 3. For the captured person, we provide the segmentation of 3D points from LiDAR frames and 2D bounding boxes from images synchronized with LiDAR. We also provide 3D pose annotations with SMPL format. Compared to other datasets Tab. 1, it is worth mentioning that SLOPER4D provides the 3D scene reconstructions and accurate global translation annotations, allowing us to quantitatively study the scene-aware global pose estimation from both LiDAR and monocular videos. In addition to the dense 3D point cloud map reconstructed from the LiDAR, SLOPER4D provides the high-precision colorful point cloud map from a Terrestrial Laser Scanner (Trimble TX5) for better visualization and map comparison.
Experiments
In this section, we first evaluate SLOPER4D Dataset qualitatively, indicating that our dataset is solid enough to benchmark new tasks. Then we perform a cross-dataset evaluation to further assess our dataset’s novelty on two tasks: LiDAR-based 3D HPE and camera-based 3D HPE. Finally, we introduce the new benchmark, GHPE, and perform experiments on GLAMR. More quantitative evaluations and experiments are in the supplementary material.
Training/Test splits. We split our data into training and test sets for LiDAR/Camera-based pose estimation. The training set of SLOPER4D contains eleven sequences of data with a total of 80k LiDAR frames and corresponding RGB frames. The test set has four sequences of data with around 20k LiDAR frames and corresponding RGB frames.
For global pose estimation, we select three challenging scenarios for evaluation. The first one is a single-person football training scenario with highly dynamic motions. The second one is running along a coastal runway. The third one is a garden tour involving daily motions.
Evaluation metrics. For 3D HPE, we employ Mean per joint position error (MPJPE) and Procrustes-aligned MPJPE (PA-MPJPE) for evaluation. MPJPE is the mean euclidean distance between the ground-truth and predicted joints. PA-MPJPE first aligns the predicted joints to the ground-truth joints by carrying out rigid transformation based on Procrustes analysis and then calculates MPJPE. For global trajectory evaluation, we utilize Absolute Trajectory Error (ATE) and the Relative Pose (the pose refers to orientation here) Error (RPE) in visual SLAM systems , where the ATE is well-suited for measuring the global localization and, in contrast, the RPE is suitable for measuring the system’s drift, for example, the drift per second. Global MPJPE (G-MPJPE) is MPJPE calculated by placing the SMPL model in the global coordinates.
Qualitative evaluation. For the human pose qualitative evaluation, we project the SMPL to the image and visualize the 3D human with corresponding LiDAR points in 3D space (shown in Fig. 3). The results demonstrate that the 3D human mesh aligns well with 3D environments and 2D images. As a large-scale urban-level human pose dataset, SLOPER4D provides multi-modal capture data and rich human-scene annotations, as well as diverse challenging human activities in large scenes. To evaluate our optimization method, we first compare our method with the results from ICP. As shown in Fig. 4, the scene-aware constraints and human mesh-to-points constraint efficiently optimize the local poses, global translation, and even the orientation error from IMU. To show the effectiveness of the camera extrinsic optimization, we report the results in Fig. 5. The 2D projecting error was visually lowered after optimization.
We evaluate root-relative 3D human pose estimation with different modalities, namely the LiDAR and the camera. 3DPW is an in-the-wild human motion dataset that is most related to us. With VIBE, we cross-evaluated our dataset’s camera modal by using 3DPW. LiDARHuman26M is a lidar-based dataset for long-range human pose estimation. We can cross-evaluate our dataset’s LiDAR modal with it. Tab. 2(a) shows the evaluation results on LiDAR-based 3D pose estimation task and Tab. 2(b) shows the results on camera-based 3D pose estimation. Taking the results from Tab. 2(a), for example, when the model is trained from another dataset only, the errors are the largest. But the error will be further reduced by around 60% when training on LiDARHuman26M and our dataset together. It suggests a domain gap exists between different LiDAR sensors, and both datasets complement each other. The results of another task show that the pre-trained VIBE model generalizes better on 3DPW than on our dataset. But the error on 3DPW increases when finetuned on our dataset, while the error decreases on our dataset. This suggests that the pre-trained model complements SLOPER4D better than the opposite. Comparing the results across different modalities, the error on our dataset from the method trained on mixed LiDAR point cloud datasets is 13% lower than the method trained on the images.
2 Benchmark on Global Human Pose Estimation
In this subsection, we benchmark the GHPE task of GLAMR on SLOPER4D. GLAMR is a global occlusion-aware method for 3D global human mesh recovery from dynamic monocular cameras. For the scale uncertainty of the monocular camera, we compute the affine matrix from the estimated trajectory to the ground truth trajectory and rotate, translate and scale the estimated trajectory before error computation.
Tab. 3 reports the global trajectory error with ATE and RPE, Tab. 4 reports the global human pose metric, and Fig. 6 shows the ATE error mapped on GT trajectory. Comparing the results on the three scenes, the football and Garden001 have a significantly lower RPE in the global scene. In comparison, GLAMR performs the worst on the running scene, with an ATE’s RMSE of 29.48 m. This scene has the largest area size and the highest human pace. GLAMR achieves a low PA-MPJPE of 86.3mm on Garden001, a sequence with daily walking and visiting motions. It’s the first time that we have tested the GPHE on such large outdoor scenes. GLAMR achieves relatively better results on daily human motion while performing worse on high-dynamic activities in the wild. The interesting point is that the trajectory tendency is pretty similar to the reference, even in dynamic football training motions, which demonstrates the ability of GLAMR to be a baseline. It is expected that more research will focus on GHPE in real-world interactive scenarios, and the experiments show our SLOPER4D’s potential to promote urban-level GHPE research.
Discussions
Limitations. Firstly, SLOPER4D is limited to single-person capture though it perceives multiple-person data. Secondly, the camera and LiDAR are not synchronized online, causing tedious offline work if the camera loses frames even with a low time offset (<50 ). Finally, texture information from the camera is not fully exploited for color and texture reconstruction of scenes and humans. In our future work, we will propose an online synchronization algorithm and extend our work to multiple-person capturing.
Conclusions. We propose the first large-scale urban-level human pose dataset with multi-modal capture data and rich human-scene annotations. Based on our proposed new dataset, we benchmark two critical tasks, camera-based 3D HPE and LiDAR-based 3D HPE. SLOPER4D also benchmarks the GHPE task. The results demonstrate the potential of SLOPER4D in boosting the development of these areas.
Our work contributes to extending motion capture to large global scenes based on the current methods and datasets. We hope this work will foster future creation and interaction in urban environments.
Acknowledgements. We thank Zhiyong Wang for helping us incorporate FAST-LIO2 into our mapping system. This work was supported in part by the National Natural Science Foundation of China (No.62171393, No.62206173), the Fundamental Research Funds for the Central Universities (No.20720220064), the open fund of PDL (WDZC20215250113, 2022-KJWPDL-12), and FuXiaQuan National Independent Innovation Demonstration Zone Collaborative Innovation Platform (No.3502ZCQXT2021003). We also acknowledge support from Shanghai Frontiers Science Center of Human-centered Artificial Intelligence (ShangHAI).