HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling
Zhongang Cai, Daxuan Ren, Ailing Zeng, Zhengyu Lin, Tao Yu, Wenjia Wang, Xiangyu Fan, Yang Gao, Yifan Yu, Liang Pan, Fangzhou Hong, Mingyuan Zhang, Chen Change Loy, Lei Yang, Ziwei Liu
Introduction
Sensing and modeling humans are longstanding problems for both computer vision and computer graphics research communities, which serve as the fundamental technology for a myriad of applications such as animation, gaming, augmented, and virtual reality. With the advent of deep learning, significant progress has been made alongside the introduction of large-scale datasets in human-centric sensing and modeling . In this work, we present HuMMan, a comprehensive human dataset consisting of 1000 human subjects, captured in total 400k sequences and 60M frames. More importantly, HuMMan features four main properties listed below.
- Multiple Modalities. HuMMan provides a basket of data formats and annotations in the hope to assist exploration in their potential complementary nature. We build HuMMan with a set of 10 synchronized RGB-D cameras to capture both video and depth sequences. Our toolchain then post-process the raw data into sequences of colored point clouds, 2D/3D keypoints, statistical model (SMPL) parameters, and model-free textured mesh. Note that all data and annotations are temporally synchronized, while 3D data and annotations are spatially aligned. In addition, we provide a high-resolution scan for each of the subjects in a canonical pose.
- Mobile Device. With the development of 3D sensors, it is common to find depth cameras or low-power LiDARs on a mobile device in recent years. In view of the surprising gap between emerging real-life applications and the insufficiency of data collected with mobile devices, we add a mobile phone with built-in LiDAR in the data collection to facilitate the relevant research.
- Action Set. We design HuMMan to empower comprehensive studies on human actions. Instead of empirically selecting daily activities, we propose to take an anatomical point of view and systematically divide body movements by their driving muscles. Specifically, we design 500 movements by categorizing major muscle groups to achieve a more complete and fundamental representation of human actions.
- Multiple Tasks. To facilitate research on HuMMan, we provide a whole suite of baselines and benchmarks for action recognition, 2D and 3D pose estimation, 3D parametric human recovery, and textured mesh reconstruction. Popular methods are implemented and evaluated using standard metrics. Our experiments demonstrate that HuMMan would be useful for multiple fields of study, such as fine-grained action recognition, point cloud-based parametric human recovery, dynamic mesh sequence reconstruction, and transferring knowledge across devices.
In summary, HuMMan is a large-scale multi-modal dataset for 4D (spatio-temporal) human sensing and modeling, with four main features: 1) multi-modal data and annotations; 2) mobile device included in the sensor suite; 3) action set with atomic motions; 4) standard benchmarks for multiple vision tasks. We hope HuMMan would pave the way towards more comprehensive sensing and modeling of humans.
Related Works
Action Recognition. As an important step towards understanding human activities, action recognition is the task to categorize human motions into predefined classes. RGB videos with additional information such as optical flow and estimated poses and 3D skeletons typically obtained from RGB-D sequences are the common input to existing methods. Datasets for RGB video-based action recognition are often collected from the Internet. Some have a human-centric action design whereas others introduce interaction and diversity in the setup . Recently, fine-grained action understanding is drawing more research attention. However, these 2D datasets lack 3D annotations. As for RGB-D datasets, earlier works are small in scale . As a remedy, the latest NTU RGB-D series features 60-120 actions. However, the majority of the actions are focused on the upper body. We develop a larger and more complete action set in HuMMan.
2D and 3D Keypoint Detection. Estimation of a human pose is a vital task in computer vision, and a popular pose representation is human skeletal keypoints. The field is categorized by output format: 2D and 3D keypoint detection, or by the number of views: single-view and multi-view pose estimation . For 2D keypoint detection, single-frame datasets such as MPII and COCO provide diverse images with 2D keypoints annotations, whereas video datasets such as J-HMDB , Penn Action and PoseTrack provide sequences of 2D keypoints. However, they lack 3D ground truths. In contrast, 3D keypoint datasets are typically built indoor data to accommodate sophisticated equipment, such as Human3.6M , CMU Panoptic , MPI-INF-3DHP , TotalCapture , and AIST++ . Compared to these datasets, HuMMan not only supports 2D and 3D keypoint detection but also textured mesh reconstruction assist in more holistic modeling of humans.
3D Parametric Human Recovery. Also known as human pose and shape estimation, 3D parametric human recovery leverages human parametric model representation (such as SMPL , SMPL-X , STAR and GHUM ) that achieves sophisticated mesh reconstruction with a small amount of parameters. Existing methods take keypoints , images , videos , and point clouds as the input to obtain the parameters. Joint limits and contact are also important research topics. Apart from those that provide keypoints, various datasets also provide ground-truth SMPL parameters. MoSh is applied on Human3.6M to generate SMPL annotations. CMU Panoptic and HUMBI leverages keypoints from multiple camera views. 3DPW combines a mobile phone and inertial measurement units (IMUs). Synthetic dataset such as AGORA renders high-quality human scans in virtual environments and fits SMPL to the original mesh. Video games have also become an alternative source of data . In addition to SMPL parameters that do not model clothes or texture, HuMMan also provides textured meshes of clothed subjects.
Textured Mesh Reconstruction. To reconstruct the 3D surface, common methods include multi-view stereo , volumetric fusion , Poisson surface reconstruction , and neural surface reconstruction . To reconstruct texture for the human body, popular approaches include texture mapping or montage , deep neural rendering , deferred neural rendering , and NeRF-like methods . Unfortunately, existing datasets for textured human mesh reconstruction typically provide no sequential data , which is valuable to the reconstruction of animatable avatars . Moreover, many have only a limited number of subjects . In contrast, HuMMan includes diverse subjects with high-resolution body scans and a large amount of dynamic 3D sequences.
Hardware Setup
We customize an octagonal prism-shaped multi-layer framework to accommodate calibrated and synchronized sensors. The system is 1.7 m in height and 3.4 m in side length of its octagonal cross-section as illustrated in Fig. 2.
RGB-D Sensors. Azure Kinect is popular with both academia and the industry with a color resolution of 19201080, and a depth resolution of 640576. We deploy ten Kinects to capture multi-view RGB-D sequences. The Kinects are strategically placed to ensure a uniform spacing, and a wide coverage such that any body part of the subject, even in most expressive poses, is visible to at least two sensors. We develop a program that interfaces with Kinect’s SDK to obtain a data throughput of 74.4 MB per frame and 2.2 GB per second at 30 FPS before data compression.
Mobile Device. An iPhone 12 Pro Max is included in the sensor suite to allow for the study on a mobile device. Besides the regular color images of resolution 19201440, the built-in LiDAR produces depth maps of resolution 256192. We develop an iOS app upon ARKit to retrieve the data.
High-Resolution Scanner. To supplement our sequential data with high-quality body shape information, a professional handheld 3D scanner, Artec Eva, is used to produce a body scan of resolution up to 0.2 mm and accuracy up to 0.1 mm. A typical scan consists of to faces and to vertices, with a 4K (40964096) resolution texture map.
2 Two-Stage Calibration
Image-based Calibration. To obtain a coarse calibration, we first perform image-based calibration following the general steps in Zhang’s method . However, we highlight that Kinect’s active IR depth cameras encounter over-exposure with regular chessboards. Hence, we customize a light absorbent material to cover the black squares of the chessboard pattern. In this way, we acquire reasonably accurate extrinsic calibration for Kinects and iPhones.
Geometry-based Calibration. Image-based calibration is unfortunately not accurate enough to reconstruct good-quality mesh. Hence, we propose to take advantage of the depth information in a geometry-based calibration stage. We empirically verify that image-based calibration serves as a good initialization for geometry-based calibration. Hence, we randomly place stacked cubes inside the framework. After that, we convert captured depth maps to point clouds and apply multi-way ICP registration to refine the calibration.
3 Synchronization
Kinects. As the Azure Kinect implements the Time-of-Flight principle, it actively illuminates the scene multiple times (nine exposures in our system) for depth computation. To avoid interference between individual sensors, we use the synchronization cables to propagate a unified clock in a daisy chain fashion, and reject any image that is 33 ms or above out of synchronization. We highlight that there is only a 1450-us interval between exposures of 160 us; our system of ten Kinects reaches the theoretical maximum number.
Kinect-iPhone. Due to hardware limitations, we cannot apply the synchronization cable to the iPhone. We circumvent this challenge by implementing a TCP-based communication protocol that computes an offset between the Kinect clock and the iPhone ARKit clock. As iPhone is recording at 60 FPS, we then use the offset to map the closest iPhone frames to Kinect frames. Our test shows the synchronization error is constrained below 33 ms.
Toolchain
To handle the large volume of data, we develop an automatic toolchain to provide annotations such as keypoints and SMPL parameters. Moreover, dynamic sequences of textured mesh are also reconstructed. The pipeline is illustrated in Fig. 3. Note that there is a human inspection stage to reject low-quality data with erroneous annotations.
There are two stages of keypoint annotation (I and II) in the toolchain. For stage I, virtual cameras are placed around the minimally clothed body scan to render multi-view images. For stage II, the color images from multi-view RGB-D are used. The core ideas of the keypoint annotation are demonstrated below, with the detailed algorithm in the Supplementary Material.
where contains the indices of connected keypoints and calculates the average bone length of a given 3D keypoint sequence. Note that 3) and 4) are jointly optimized.
Keypoint Quality. We use and as keypoint annotations for 2D Pose Estimation and 3D Pose Estimation, respectively. To gauge the accuracy of the automatic keypoint annotation pipeline, we manually annotate a subset of data. The average Euclidean distance between annotated 2D keypoints and reprojected 2D keypoints is pixels on the resolution of .
2 Human Parametric Model Registration
Keypoint Energy. SMPLify estimates camera parameters to leverage 2D keypoint supervision, which may be prone to depth and scale ambiguity. Hence, we develop the keypoint energy on 3D keypoints. For simplicity, we denote as , the global rigid transformation derived from the SMPL kinematic tree as , the joint regressor as . We formulate the energy term:
Surface Energy. To supplement 3D keypoints that do not provide sufficient constraint for shape parameters, we add an additional surface energy term for registration on the high-resolution minimally clothed scans in stage I only. We use bi-directional Chamfer distance to gauge the difference between two mesh surfaces:
where and are the mesh vertices of the high-resolution scan and SMPL.
Shape Consistency. Unlike existing work that enforces an inter-beta energy term due to the lack of minimally clothed scan of each subject, we obtain accurate shape parameters from the high-resolution scan that allow us to apply constant beta parameters in the registration in stage II.
Full-body Joint Angle Prior. Joint rotation limitations serve as an important constraint to prevent unnaturally twisted poses. We extend existing work that only applies constraints on elbows and knees to all joints in SMPL. The constraint is formulated as a strong penalty outside the plausible rotation range (with more details included in the Supplementary Material):
where and are the upper and lower limit of a rotation angle. Note that each joint rotation is converted to three Euler angles which can be interpreted as a series of individual rotations to decouple the original axis-angle representation.
3 Textured Mesh Reconstruction
Point Cloud Reconstruction and Denoising. We convert depth maps to point clouds and transform them into a world coordinate system with camera extrinsic parameters. However, the depth images captured by Kinect contain noisy pixels, which are prominent at subject boundaries where the depth gradient is large. To solve this issue, we first generate a binary boundary mask through edge finding with Laplacian of Gaussian Filters. Since our cameras have highly overlapped views to supplement points for one another, we apply a more aggressive threshold to remove boundary pixels. After the point cloud is reconstructed from the denoised depth images, we apply Statistical Outlier Removal to further remove sprinkle noises.
Geometry and Depth-aware Texture Reconstruction. With complete and dense point cloud reconstructed, we apply Poisson Surface Reconstruction with envelope constraints to reconstruct the watertight mesh. However, due to inevitable self-occlusion in complicated poses, interpolation artifacts arise from missing depth information, which leads to a shrunk or a dilated geometry. These artifacts are negligible for geometry reconstruction. However, a prominent artifact appears when projecting a texture onto the mesh even if the inconsistency between the true surface and the reconstructed surface is small. Hence, we extend MVS-texturing to be depth-aware in texture reconstruction. We render the reconstructed mesh back into the camera view and compare the rendered depth map with the original depth map to generate the difference mask. We then mask out all the misalignment regions where the depth difference exceeds a threshold . The masked regions do not contribute to texture projection. As shown in Fig. 6(b), the depth-aware texture reconstruction is more accurate and visually pleasing.
Action Set
Understanding human actions is a long-standing computer vision task. In this section, we elaborate on the two principles, following which we design the action set of 500 actions: completeness and unambiguity. More details are included in the Supplementary Material.
Completeness. We build the action set to cover plausible human movements as much as possible. Compared to the popular 3D action recognition dataset NTU-RGBD-120 whose actions are focused on upper body movements, we employ a hierarchical design to first divide possible actions into upper extremity, lower limbs, and whole-body movements. Such design allows us to achieve a balance between various body parts instead of over-emphasizing a specific group of movements. Note that we define whole body movements to be actions that require multiple body parts to collaborate, including different poses of the body trunk (e.g. lying down and sprawling). Fig. 7(c) demonstrates the action hierarchy and examples of interesting actions that are vastly diverse.
Unambiguity. Instead of providing a general description of the motions , we argue that the action classes should be clearly defined and are easy to identify and reproduce. Inspired by the fact that all human actions are the result of muscular contractions, we propose a muscle-driven strategy to systematically design the action set from the perspective of human anatomy. As illustrated in Fig. 7(a)(b), major muscles are identified by professionals in fitness and yoga training, who then put together a list of standard movements associated with these muscles. Moreover, we cross-check with the action definitions from existing datasets to ensure a wide coverage.
Subjects
HuMMan consists of 1000 subjects with a wide coverage of genders, ages, body shapes (heights, weights), and ethnicity. The subjects are instructed to wear their personal daily clothes to achieve a large collection of natural appearances. We demonstrate examples of high-resolution scans of the subjects in Fig. 8. We include statistics in the Supplementary Material.
Experiments
In this section, we evaluate popular methods from various research fields on HuMMan. To constrain the training within a reasonable computation budget, we sample 10% of data and split them into training and testing sets for both Kinects and iPhone. The details are included in the Supplementary Material.
Action Recognition. HuMMan provides action labels and 3D skeletal positions, which can verify its usefulness on 3D action recognition. Specifically, we train popular graph-based methods (STGCN and 2s-AGCN ) on HuMMan. Results are shown in Table 2. Compared to NTU RGB+D, a large-scale 3D action recognition dataset and a standard benchmark that contains 120 actions , HuMMan may be more challenging since 2s-AGCN achieves Top-1 accuracy of 88.9% and 82.9% on NTU RGB+D 60 and 120 respectively, but 74.1% only on HuMMan. The difficulties come from the whole-body coverage design in our action set, instead of over-emphasis on certain body parts (e.g. NTU RGB+D has a large proportion of upper body movements). Moreover, we observe a significant gap between Top-1 and Top-5 accuracy (30%). We attribute this phenomenon to the fact that there are plenty of intra-actions in HuMMan. For example, there are similar variants of push-ups such as quadruped push-ups, kneeling push-ups, and leg push-ups. This challenges the model to pay more attention to the fine-grained differences in these actions. Hence, we find HuMMan would serve as an indicative benchmark for fine-grained action understanding.
3D Keypoint Detection. With the well-annotated 3D keypoints, HuMMan supports 3D keypoint detection. We employ popular 2D-to-3D lifting backbones as single-frame and multi-frame baselines on HuMMan. We experiment with different training and test settings to obtain the baseline results in Table 3. First, in-domain training and testing on HuMMan are provided. The values are slightly higher than the same baselines on Human3.6M (on which FCN obtains MPJPE of 53.4 mm). Second, methods trained on HuMMan tend to generalize better than on Human3.6M. This may be attributed to HuMMan’s diverse collection of subjects and actions.
3D Parametric Human Recovery. HuMMan provides SMPL annotations, RGB and RGB-D sequences. Hence, we evaluate HMR , not only one of the first deep learning approaches towards 3D parametric human recovery but a fundamental component for follow-up works , to represent image-based methods. In addition, we employ VoteHMR , a recent work that takes point clouds as the input. In Table 4, we find that HMR has achieved low MPJPE and PA-MPJPE, which may be attributed to the clearly defined action set and the training set already includes all action classes. However, VoteHMR is not performing well. We argue that existing point cloud-based methods rely heavily on synthetic data for training and evaluation, whereas HuMMan provides genuine point clouds from commercial RGB-D sensors that remain challenging.
Textured Mesh Reconstruction. We gauge mesh geometry reconstruction quality of PIFu, PIFuHD, and Function4D (F4D) in Table 5 with Chamfer distance (CD) as the metric. Note that benefiting from the multi-modality signals, HuMMan supports a wide range of surface reconstruction methods that leverage various input types like PIFu (RGB-only), 3D Self-Portrait (single-view RGBD video), and CON (multi-view depth point cloud).
Mobile Device. It is under-explored that if model trained with the regular device is readily transferable to the mobile device. In Table 6, we study the performance gaps across devices. For the image-based method, we find that there exists a considerable domain gap across devices, despite that they have similar resolutions. Moreover, for the point cloud-based method, the domain gap is much more significant as the mobile device tends to have much sparser point clouds as a result of lower depth map resolution. Hence, it remains a challenging problem to transfer knowledge across devices, especially for point cloud-based methods.
Discussion
We present HuMMan, a large-scale 4D human dataset that features multi-modal data and annotations, inclusion of mobile device, a comprehensive action set, and support for multiple tasks. Our experiments point out interesting directions that await future research, such as fine-grained action recognition, point cloud-based parametric human estimation, dynamic mesh sequence reconstruction, transferring knowledge across devices, and potentially, multi-task joint training. We hope HuMMan would facilitate the development of better algorithms for sensing and modeling humans.
Acknowledgements. This work is supported by NTU NAP, MOE AcRF Tier 2 (T2EP20221-0012), NSFC No.62171255, and under the RIE2020 Industry Alignment Fund - Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).
References
Appendix 0.A Overview
We provide additional details of data collection (Section 0.B), hardware (Section 0.C), toolchain (Section 0.D), action set (Section 0.E), subjects (Section 0.F), experiments (Section 0.G), and a more complete dataset comparison (Section 0.H).
Appendix 0.B Additional Details of Data Collection
The data collection has two stages for each subject. 1) each subject receives two high-resolution scans, one with natural clothes on and the other with a tight-fitting suit on, both captured by the Artex Eva 3D Scanner. To ensure the high quality of the scans, the subjects are instructed to stand in a special pose (the canonical pose) on a turntable, that allows for a 360-degree full-body scanning with minimal self-occlusion. Each high-resolution scan includes an MTL information file, an OBJ mesh file, and a BMP texture file. 2) After that static body scanning, the subject enters the framework and follows instructions to perform 40-60 actions, randomly sampled from the action set that contains 500 actions. Each action that a subject performs is a sequence, that consists of ten Kinect RGB-D sequences and an iPhone RGB-D sequence. We show sample frames collected with our hardware setup in Fig. 9. Each sequence takes 5-15 seconds and 150-450 frames at 30 FPS per view. We compress all sequential data in a custom data format SMC that is developed based on HDF5 format. The SMC file also contains additional information such as camera parameters, subject ID, and action ID.
Appendix 0.C Additional Details of Hardware
We provide more details on the RGB-D sensor (Azure Kinect). We set operating mode to NFOV unbinned for the depth cameras, which results in the largest view overlap with the color camera and the densest point clouds. The depth camera in this mode has an FOV of . The operating range of the depth sensor in this mode is between 0.5 m to 3.86 m. The typical systematic error of the depth sensor is less than 11 mm + 0.1% of distance with a standard deviation of less than 17 mm. In view of the limited FOV and depth error-distance relation, we design our aluminum framework such that the subject is around 2 m away from the Kinects: at that distance, the FOV can accommodate the subject’s whole body, without incurring any extra depth error.
C.2 Synchronization
Our data sampling program runs on a workstation, and it 1) integrates the Kinect SDK, and 2) communicates with the iPhone app developed based on ARKit through TCP. Since there is no existing hardware approach to Kinect-iPhone synchronization, we develop a method to compute the difference between Kinect clock and iPhone ARKit clock . Hence, we first obtain the offset from the workstation to the Kinects as
where is the Kinect clock time and is the workstation’s system time, obtained at the same moment. We also send a message to the iPhone app, which records down the iPhone system clock upon receiving the message and sends back a message to the workstation to complete a round trip. We compute the offset from the iPhone system clock to the workstation system clock as
where is the round trip time taken. Note that there is an additional offset between the ARKit clock and the iPhone system clock , computed as
where is the ARKit clock. Finally, the required clock difference is
C.3 Point Clouds
Both Kinect and iPhone produce depth maps that can be converted to point clouds. However, iPhone’s point cloud is much sparser than Kinect’s. We show unprocessed raw point clouds produced by the two types of sensors in Fig. 10. In addition, iPhone does not report the LiDAR accuracy; we empirically find that iPhone point clouds are noisier, especially at the object boundaries, than Kinect point clouds.
Appendix 0.D Additional Details of Toolchain
The overall pipeline for keypoint annotation is summarized in Algorithm 1.
D.2 Full-body Angle Prior
It is surprisingly difficult to find literature that provides a complete analysis of joint movement ranges, especially rotations in three degrees of freedom (DOF). Hence, we take references from artists’ guidelines on human anatomyhttps://design.tutsplus.com/articles/human-anatomy-fundamentals-flexibility-and-joint-limitations--vector-25401 and 3D modelers’ suggested practiceshttps://wiki.secondlife.com/wiki/Suggested_BVH_Joint_Rotation_Limits, to simplify the constraint such that the three DOF movement range is bounded by the maximum ranges in each of the DOF. Despite that this formulation is not perfect, it provides constraints that are otherwise completely absent. To easily apply the per-axis ranges, we convert the axis-angle representation into Euler angles and define the Z-axis to be aligned with the child bone of the joint in the kinematic tree (for example, forearm is the child bone of the joint elbow). To circumvent gimbal lock as much as possible, we define the joint frame coordinate such that the second rotation axis (Y-axis) always falls on the less flexible axis (for which the rotation is unlikely to reach 90∘). Hence, we define the X-axis as the axis around which the largest rotation is achieved. Y-axis is finally defined with X- and Z-axis fixed. All values undergo manual inspection and are adjusted empirically. Note that the Euler angle rotation is used to generate a loss only; the joint rotation is still in axis-angle representation.
D.3 Annotation Quality of SMPL Parameters.
To evaluate the body shape, we compute the per-vertex error on the high-resolution scan that is the uni-directional Chamfer distance from registered SMPL mesh vertices to the high-resolution scan vertices. Note that high-resolution scans have been scaled to the real height of scanned persons. The mean per-vertex error is 0.16 mm. We also visualize the registration quality in Fig. 11. To evaluate the body pose, we compute the per-joint error as the L2 Euclidean distance between 3D keypoints and 3D joints of registered SMPL on the dynamic sequences. The mean per-joint error is 38.18 mm. Note that the error is largely attributed to the difference in the joint definition of the keypoint detector and the parametric model. As a reference, registration with an accurate optical marker system yields a per-joint error of 29.34 mm.
Appendix 0.E Additional Details of Action Set
Design Process. In HuMMan, we design a hierarchical structure for a systematic coverage of different body parts to collate a complete and unambiguous action set. Specifically, we have body at the center as the first order. The second order consists of whole body, upper extremity and lower limbs that categorize actions by major body parts. After that, we propose a muscle-driven strategy to further split each major body part into main muscle groups according to human anatomy as the third order. Finally, we involve domain experts to design a series of action variants associated with each muscle in the fourth order. The full action hierarchy is demonstrated in Fig. 12.
Motion Diversity. As HuMMan contains a large amount of data, we further conduct a preliminary study on the motion diversity for further research on the motion prior learning. Specifically, We compute the mean standard deviation of joint angles of three datasets: 3DPW (0.159), AMASS (0.208), and HuMMan (0.269). The higher mean standard deviation indicate higher diversity in motions. Although the AMASS dataset with a large-scale MoCap data is wildly used in many recent works to pretrain models, HuMMan has more diversity in joint angles, showing its potential for human motion-related tasks.
Appendix 0.F Additional Details of Subjects
Statistics. HuMMan consists of 1000 subjects. To evaluate the diversity, we include key statistics (gender, age, height and weight) of the subjects in Fig. 13.
Ethics. HuMMan involves a large number of human subjects so that we pay special attention to address ethic concerns. The recruitment process is conducted on an entirely voluntary basis. Actors and actresses who participate in HuMMan are well-informed, with legal agreements signed to acknowledge that the data will be made public for research purposes.
Appendix 0.G Additional Details of Experiments
HuMMan contains a massive scale of subjects (1000), actions (500), sequences (400k) and frames (60M). To constrain training and testing within a reasonable computation budget, we sample only 10% of the data. We then develop three protocols to split iPhone and Kinect data into training and test sets. Protocol 1 (P1): split by subjects, the training and test set are mutually exclusive and contain 70% and 30% of the subjects respectively. P1 is used for all experiments in the main paper. Protocol 2 (P2): split by actions. We split actions into three categories according to major body parts involved: upper extremity, lower limbs, and whole body. Training is conducted on one category whereas the test is conducted on the other two. Protocol 3 (P3): split by views. Model is trained on only one view (the front view, or the view of the iPhone and the Kinect with ID 0) and tested on all views.
G.2 2D Keypoint Detection
We study 2D keypoint detection baselines on HuMMan primarily for 2D-to-3D keypoint lifting. CPN is a cascaded pyramid network to improve hard keypoints detection. HRNet is a novel high-resolution network that obtains high performance on COCO dataset , and LiteHRNet is an efficient version of HRNet. The comparison results are listed in Table 7. Because 2D keypoints are often used as an intermediate representation of 3D keypoints in a two-stage manner , the good performance in this task can be helpful to the estimation of subsequent 3D.
G.3 3D Keypoint Detection
3D keypoint detection benchmarks under P1 setting are presented in the main paper and additional benchmarks under P2 and P3 are provided here. In Table 8, we show results on the cross-action (P2) performance of the FCN method . Compared with Protocol 1, we observe that training with fewer actions and testing on unseen actions degrade the precision significantly, especially for cross-evaluation on the whole body category which seems to have a large action distribution misalignment with the other two categories. Furthermore, we report results of cross-view (P3) in Table 9. When the model is only trained on one view (i.e., View 0), we observe a considerable domain gap across different views as the errors increase as the deviation from the test view from the training view increases. The experiment results indicate that cross-view 3D keypoint detection is challenging.
G.4 3D Parametric Human Recovery
In addition to P1 benchmarks for 3D parametric human recovery presented in the main paper, we also provide more benchmarks under P2 and P3. In Table 10, we evaluate the cross-action (P2) performance of the HMR baseline. We find that testing on unseen poses is challenging (compared to P1 benchmark results). Moreover, whole body actions seem to have a distribution that is further away from lower limbs and upper extremity actions. In Table 11, we study the cross-view setting (P3), which is even worse than the cross-action setting. The HMR baseline is trained on View 0, and gives a clear trend that the greater the viewing angle difference, the larger the errors. View 5 is directly opposite View 0 and yields the largest error.
G.5 Textured Mesh Reconstruction
To fully demonstrate the capacity of HuMMan, we also provide the results of Function4D as a baseline for textured mesh reconstruction since it combines both volumetric fusion and implicit surface reconstruction for volumetric capture in real-time. The results of Function4D, using 4 (ID: 0,3,6,9) views, are shown in Fig. 14.
Appendix 0.H A More Complete Dataset Comparison
In Table 12, we provide a more thorough comparison of HuMMan with similar datasets for 1) action recognition, 2) 2D and 3D keypoint detection, 3) 3D parametric human recovery, and 4) mesh reconstruction.